PII masking for training data that fails closed: detect, mask, measure, hold back

PII masking for training data has one property that masking for a live request does not: there is no second chance. A record that reaches a training run with a name or an account number in it can be memorised by the model, emitted later, and cannot be recalled by deleting the file. That asymmetry is the whole argument for failing closed.

This post compares the three detection methods and the three things that can be done with a detected span, explains what fail-closed means mechanically, and says how to measure recall without flattering yourself. Domain differences come last, because that is where general rules stop working.

What PII is, and why training data is the hard case

NIST SP 800-122 describes personally identifiable information as information that can be used to distinguish or trace an individual's identity, such as a name or a date and place of birth, together with any other information that is linked or linkable to that individual. The second half is the part detectors miss: a date of birth, a postcode and a job title are each harmless and together are a person.

Training data makes this harder in three ways. Volume: a corpus is too large to read, so detection is statistical and recall is never complete. Persistence: a model can memorise and reproduce what it saw, and the duplicates that inflate memorisation multiply the exposure. Irreversibility: once trained, the only remedy is to retrain. So the place to act is before the training set exists, and before the record reaches any third-party model used for extraction, one of the hops in the path text takes through an LLM integration.

Three detection methods and what each misses

Production detectors combine all three: Presidio, an open-source example, describes its recognisers as named entity recognition, regular expressions, rule-based logic and checksums with relevant context. What none of them sees is the quasi-identifier, the combination of individually innocent fields that together identify someone, which needs a record-level rule rather than a span-level detector.

Masking, pseudonymisation or dropping

Masking replaces the span with a typed placeholder such as [PERSON_1], numbered so that the same entity gets the same placeholder within a record and references still resolve. Structure survives, the identity does not, and nothing is kept that could reverse it.

Pseudonymisation replaces the span with a consistent surrogate and keeps the mapping somewhere else. The GDPR defines it as processing such that the data can no longer be attributed to a specific person without additional information that is kept separately; the data remains personal data, and the mapping table becomes the most sensitive object in the system. It is right when linkage across records matters for the task and wrong when nobody will ever need to reverse it.

Dropping removes the record. It loses data, which is why teams avoid it, and it is the only correct answer when the personal data is the content rather than an incident in it: a complaint letter, a medical narrative, a transcript of one identifiable person.

Why PII masking for training data fails closed

Fail closed means that when the masking stage cannot clear a record, the record does not continue. "Cannot clear" has three concrete cases: the detector returned a span with confidence below the threshold at which masking is applied; the detector errored or timed out on the record; or a record-level rule, such as a quasi-identifier check, could not be evaluated. In each case the record is held back with a reason code, routed to review, and counted.

The default in most pipelines is the opposite: an exception in a stage is caught, logged and skipped, and the record flows on with nothing masked. That is fail open, and it is invisible until someone finds a name in a model's output. The fix is structural: the masking stage emits a verdict per record, and the only verdict that lets a record continue is "cleared"; everything else, including no verdict at all, holds.

The cost is bounded: held records are a share of the delivery that a person has to look at or the buyer has to accept losing, and both counts belong in the quality report. A record wrongly held costs one record; a record wrongly passed costs a retraining.

Measuring recall honestly

Recall is the share of personal data actually present that the detector found. It cannot be computed from the detector's own output, because the denominator is exactly what the detector did not see. The only honest estimate comes from a labelled sample: records drawn at random from the delivery, stratified by source and document type, annotated by people, and compared with the detector's spans.

The general method holds; the entity lists and the definition of "the content" do not. In finance, the identifiers are account numbers, card numbers, tax identifiers and the counterparties named in transaction narratives, and a masked statement must still reconcile, so amounts and dates stay while the party goes.

In legal text the parties are often the subject matter: a judgment is about named people, and rules on anonymisation vary by court and jurisdiction, so the rule is set per source rather than per corpus. In health, free-text clinical notes are the hard case: dates, small geographic units, institution names and ages above a threshold are treated as identifying on their own, which is why regulatory identifier lists for health data are long. Each domain therefore carries its own rules, as the pages for finance data and legal data set out.

Runix Pipeline, the tooling behind Runix Data, is designed to identify and mask personal data before it reaches a model or a training set, and to fail closed when detection is uncertain, with held records counted in the quality report. Runix Pipeline is in development with design partners; Runix Data, which runs on it, is in early access.

Questions this raises

What does fail closed mean for PII masking?

When the masking stage cannot clear a record, because detection is uncertain, the detector errored, or a record-level rule could not be evaluated, the record is held back rather than passed through. Held records are counted and reviewed instead of silently continuing unmasked.

Is pseudonymisation the same as masking?

No. Masking replaces a span with a placeholder and keeps nothing that could reverse it, while pseudonymisation keeps a separate mapping so the data can be re-linked, which means it is still personal data under the GDPR definition.

How should recall be measured?

From a labelled sample drawn at random from the delivery and annotated by people, reported with the sample size per entity type and per source. A detector cannot measure its own recall, and canaries only measure what was planted.

Related to this post: Runix Pipeline. Tell us what you are building and we reply within one business day.