Red teaming
Also: 紅隊 · 紅隊測試 · red teaming · 對抗性測試
Deliberately attacking your own system from the attacker's position, to find the failures before users or real attackers do.
When you will meet it
You need it before shipping any agent that reads external content or takes actions. Checklist review ("we added input filtering ✓") confirms controls exist, not that they work — bypasses live in the combinations you never imagined: another encoding, another language, a slow multi-turn setup. A red team's value is precisely that it does not play from your checklist.
An analogy
Like paying a thief to rob your own bank: a fire drill confirms everyone knows where the exits are; only hiring the thief confirms the vault actually resists. Drills follow scripts; thieves do not.
Minimal example
一個最小的紅隊迴圈(示意):
1. 定義範圍:哪些行為算失敗(洩漏系統提示?執行文件裡的指令?)
2. 設計攻擊向量:直接注入、間接注入、工具回傳夾帶、
編碼混淆、語言切換、多輪鋪陳……
3. 執行並記錄:每個向量記下 輸入/輸出/有沒有被攔/攔在哪一層
4. 修好一個,全部重跑:防止修好 A 卻弄壞 B
同一組攻擊要能重複執行 —— 換了模型或提示之後,
才知道防禦是變強還是變弱。Steps 3 and 4 are the point: a red team's output is not "tested, looks fine" but a re-runnable attack corpus. Model and prompt updates drift behaviour; what was blocked last month may pass this month.
What people get wrong
- Testing only single-turn, English-only, direct injection. What lands most often in practice is indirect injection (hidden in the pages, documents and tool returns the agent reads) and slow multi-turn setups — the combinations outside your checklist are where attackers actually live.
- Using one model as both gatekeeper and subject under test. Its blind spots are systematic: the same misjudgement appears on the attack side and the defence side, hiding the gap perfectly. The referee must not also be a player.
Related terms
Next
- AI Security and Red Teaming: Prompt Injection, Jailbreak Defense7 min
- Enterprise Data Security Basics: What You Must Never Feed an AI8 minChinese only