Skip to content
Worfilo

Documentation

PII masking: keep personal data away from the model

Updated September 28, 2026

Support tickets, forms and notes are full of personal data that an agent rarely needs to do its job. The Mask PII node replaces it with tokens such as <PERSON_1> and <EMAIL_1> before the text reaches a model, so the agent still understands the message without seeing who wrote it.

Add the Mask PII node

Put a Mask PII node, from the Tool group in the palette, between the data and the agent. Its Text field is a template and defaults to {{ input }}, whatever arrives on its input. The masked text leaves on out.

Agent user prompt
Classify this ticket by urgency and topic:

{{ nodes.mask.output }}
What the agent sees
Hi, this is <PERSON_1>. My card <CREDIT_CARD_1> was charged twice.
Call me on <PHONE_1> or write to <EMAIL_1>.

Within one run, the same value always gets the same token, so two Mask PII nodes in a workflow agree on who <PERSON_1> is.

Profiles: what to look for

A profile is a named set of data types. Pick one on the node:

  • General (the default): names, contact details and common identifiers.
  • Healthcare: everything in General, plus medical record numbers, NHS and Medicare numbers, doctors, hospitals, dates, conditions and medications. It runs at the strict setting, because a missed patient identifier costs more than an extra mask.
  • Finance: everything in General, plus IBANs, tax and social security numbers, and bank routing and sort codes.
  • Minimal: email, phone and card numbers only.

To narrow a profile, list the exact Entities you want masked; leave it empty to use the profile's own list.

How detection works

Three detectors run together:

  • Recognizers for structured identifiers. Any format with a checksum, such as card numbers, IBANs and NHS numbers, is validated, so a number of the right shape but the wrong check digit is not masked.
  • A multilingual model (GLiNER) for names, places, organisations, dates and other free-text data, in any language.
  • Your Always mask list, for terms that must never reach the model, such as a project code name.

Sensitivity (relaxed, balanced or strict) sets how confident the model must be before it masks something. Stricter catches more and over-masks more. It does not change the recognizers, since a checksum either passes or it does not.

Countries

Email, phone, card, IBAN and IP address detection always runs. National identifiers only run for the countries you turn on under Countries, using ISO codes:

  • US: social security and Medicare numbers. GB: NHS and National Insurance numbers.
  • IN: Aadhaar and PAN. LK: national identity card numbers, old and new.
  • CA: SIN. AU: Medicare and TFN. DE and FR: tax ID and social security numbers.

Use ALL to turn on every country. Default phone region helps read local phone numbers written without a country code.

Never mask and Always mask

Never mask keeps terms that look personal but are not, such as your product or clinic name. Always mask forces terms to be masked. A term on both lists is masked. Common medical eponyms such as Parkinson and tech terms such as Kafka are already protected from being read as names.

What is stored, and if masking is unavailable

The original values are kept encrypted in a short-lived store for the run, and never written to disk. The run records only how many values of each type were masked, never the text itself.

If the shield is down decides what happens when masking cannot run. The default, pass through, lets the text continue unmasked so the workflow still finishes. Fail sends the error to the node's error port instead. Choose Fail whenever unmasked data must never reach the model.

Build it on the canvas

Create a free account, describe the workflow or wire it yourself, and run it in the browser.