Data classification
Also: 資料分級 · 數據分級 · 敏感資料分類 · data classification
Sorting data by sensitivity first — public, internal, confidential, restricted — then deciding where each class may flow, including whether it may enter a model's context.
When you will meet it
It is step zero before any data meets a model, and the prerequisite for every other control: a sandbox limits what the agent may touch, but whether this particular data may be touched at all needs an answer first. Without classification every downstream control is guessing, and "what must never be fed to a model" has no answer at all.
An analogy
Like clearance labels in an archive: public shelves, internal shelves, a locked cabinet. With labels, who may enter which aisle is decidable — cleaners, visitors, staff. Without them only two extremes remain: everyone sees everything, or nobody sees anything.
Minimal example
一個最小分級表(示意,依組織調整):
公開 官網內容、已發布的研究 → 任何模型皆可
內部 一般程式碼、內部文件 → 限公司核准的模型
機密 客戶名單、財務明細、合約 → 先脫敏或取得授權,否則不進上下文
受限 密碼、金鑰、PII、未公開併購案 → 原則上永不進模型Note classification is a property of data, not of systems: the same agent may touch internal data and must not touch restricted data. Once classes exist, sandboxes and permission gates have something to reference — the rule stops being "be careful" and becomes "confidential or above requires approval".
What people get wrong
- Adopting the AI first and classifying later. Retro-fitting draws lines in an already-leaked pool: once data has entered a third-party model's context you cannot recall it, and usually cannot confirm how it was used.
- Classifying finer than anyone can operate. Thirty tiers with per-document review end in nobody classifying, or everything quietly becoming "internal". A coarse scheme people actually follow beats an elegant one on paper.
Related terms
Next
- Enterprise Data Security Basics: What You Must Never Feed an AI8 minChinese only
- AI Compliance and Risk Boundaries: A Legal Checklist for Enterprise Rollout18 minChinese only