Agentic Research

De-identification

Also: 去識別化 · 匿名化 · 脫敏 · anonymization

Removing or rewriting the person-pointing parts of data so it stays analysable but (ideally) no longer reveals who it is about.

When you will meet it

It is the mandatory step whenever you test an agent on real data, build a knowledge base, or hand cases to an external model. The cost of skipping it is covered under PII; the more dangerous move is doing it and then relaxing — de-identification is frequently reversible, especially against an attacker holding side information.

An analogy

Like pixelating faces in a photo. Usually enough — but a unique background, a date, a name tag, or a viewer who already knows the subject makes the pixelated photo identifiable anyway. The mosaic blocks the most obvious road, not all roads.

Minimal example

原始:「王小明,男,34 歲,某市某科技公司工程師,
       3 月 12 日因車禍送醫,診斷為左腿骨折」

去識別化:「一名 30-39 歲男性,居住於某地區,
       於某季因車禍就醫,診斷為下肢骨折」

攻擊者若握有旁側資訊(例如知道 3 月 12 日某區發生過
一起工程師車禍),第二條照樣能對回第一條。

Watch the generalisation: specific fields become ranges (34 → 30-39, date → season), keeping most of the analytical value — but every rare combination that survives adds reversibility. De-identification is a trade between utility and exposure, not a switch.

What people get wrong

  • Treating de-identification as an irreversible done state. With side information — another dataset, news coverage, insider knowledge — an attacker can re-identify individuals. Re-identification (linkage) attacks are well documented in both research and practice.
  • Equating anonymisation with de-identification. Strict anonymisation demands irreversibility and is far harder; most teams actually perform pseudonymisation or generalisation, then drop all further protection under the banner of "anonymised".

Related terms

Next