PII Scrubbing & Data Minimisation
Automated detection and redaction of faces, license plates and names, with manual audit and QA sign-off — data minimisation before delivery.
Data minimisation means a dataset carries only what the use case needs. Cognegica strips personal information — faces, license plates, names and other identifiers — through an automated detection-and-redaction pipeline backed by a manual audit and QA sign-off, so personal data is removed before delivery rather than shipped and regretted. This page documents that pipeline stage by stage.
Why minimise
How is personal information removed from a dataset?
The safest personal data is the data you no longer hold. Before a dataset is delivered, identifiers that are not needed for the task are detected and redacted — and the redaction is checked by a human, because automated detection alone misses edge cases.
Automated detection and redaction
Standard pipeline stages detect and redact faces, license plates, names and other PII across the relevant modalities. Detection thresholds and the exact identifier set are configured per engagement.
Manual audit and QA sign-off
A manual audit reviews a sample for missed or over-redacted content, feeding corrections back before QA sign-off. This is the same multi-layer QA discipline used across our annotation quality frameworks. De-identification is one stage of the wider provenance framework, and keeps sovereign datasets responsible to build.
The scrubbing pipeline
From detection to QA sign-off
An automated pass followed by human audit. The identifier set and thresholds are configured per engagement.
-
1
Automated detection
Standard pipeline stages detect PII candidates — faces, license plates, names and other identifiers — across the relevant modalities.
Detection pass complete
-
2
Redaction (faces / plates / names / PII)
Detected identifiers are redacted or masked so the delivered data carries only what the use case needs.
Identifiers redacted
-
3
Manual audit
A human audit reviews a sample for missed or over-redacted content and routes corrections back.
Manual audit sample passed
-
4
QA sign-off
Multi-layer QA confirms minimisation against the agreed bar before delivery.
QA sign-off recorded
Data type, method, check
What gets scrubbed, and how
Each personal-data type, the automated method that detects and redacts it, and the manual check that verifies it.
| Data type | Automated method | Manual check |
|---|---|---|
| Faces | Automated face detection and blurring / masking in images and video | Audit sample for missed or partial faces |
| License plates | Automated plate detection and redaction in images and video | Audit sample for missed or partial plates |
| Names & named entities | Automated named-entity detection and masking in text and transcripts | Native-linguist review for context-dependent names |
| Other PII (IDs, contacts, locations) | Pattern- and entity-based detection and redaction | Manual audit against the engagement's identifier set |
Standard pipeline stages — no vendor names or accuracy figures are implied. The identifier set, detection thresholds and any reported metric are editable placeholders scoped and confirmed per engagement.
Related work
Minimisation in context
-
Data provenance framework
Where de-identification sits in the end-to-end provenance lifecycle.
See the framework -
Sovereign data
Residency, consent and provenance for regulated programs.
Learn more -
Annotation quality frameworks
The multi-layer QA discipline behind the manual audit.
Read more
About PII scrubbing
Questions about PII and minimisation
- What personal data do you remove?
Faces, license plates, names and named entities, and other identifiers such as IDs, contacts and locations — scoped to the engagement's identifier set.
- Is removal automated or manual?
Both. Automated detection and redaction run first, then a manual audit reviews a sample for missed or over-redacted content before QA sign-off.
- Do you publish accuracy figures?
We don't imply fixed accuracy figures here. Detection thresholds and any reported metric are scoped and confirmed per engagement rather than asserted as a blanket number.
Ship datasets that carry only what they need.
Tell us the modalities and identifier set — we'll scope automated scrubbing with a manual audit.