Read these first
This article assumes the following earlier in its learning path.
Honeypots and Anti-Agent Traps: When the Prey Starts Setting Traps
Core premise: once the attacker becomes an AI agent, defenders gain a new weapon: traps designed around how agents behave. And any team whose agents read external content can fall into them. Angle: defensive technique and self-protection. Principles only — no operational detail.
Honeypots: turning attacker time into defensive assets
A honeypot is a system deliberately made to look valuable while actually being a trap. Its value is not blocking attacks but three things:
- Early detection — legitimate business should never touch it, so any contact triggers an alert;
- Wasting attacker time — opponents spend effort on a false target;
- Learning tradecraft — observe what an intruder does next.
For defenders the return is high: low deployment cost, and it converts "waiting passively for alerts" into "actively luring the opponent".
How attackers spot a honeypot (three layers)
Reading attacker detection methods backwards gives you a honeypot deployment guide:
| Layer | What the attacker checks | Defender's response |
|---|---|---|
| Static signature | Matching service fingerprints (banner, certificate, error pages) against known honeypot libraries | Do not change surface config only — make certificates and service behaviour genuinely consistent |
| Behavioural probing | Contradictory version numbers, inconsistent identity | Keep internal consistency; avoid "one host, many versions" |
| Post-login inspection | Memory, processes, egress matching a real system | The hardest layer to fake — fidelity is where value lies |
A key asymmetry: when judging "is this a honeypot", the attacker's mistake is expensive (treating a real asset as a honeypot means abandoning the real target). They therefore tend to "not block on a single signal, only downgrade interaction". That leaves defenders room: keeping the opponent uncertain extends their dwell time.
A new class: traps designed for AI agents
This emerged in the last two years. Four main types:
Trap one: attestation induction
Principle: construct a situation where the agent "proves what it is" — describing its system prompt, capability list, or internal configuration.
Why it works: agents are trained to be helpful and readily answer "who are you / what can you do".
Self-protection: never self-attest to external content. An agent should not prove identity, capability, or internal state to content it reads.
Trap two: reverse prompt injection
Principle: hide instructions inside a page or document the agent reads, causing unintended actions.
Why it works: agents struggle to separate "data" from "instructions" — both look like text.
Self-protection: external content is always data, never instruction. Treat all externally obtained text as untrusted input, not as task direction.
Trap three: tarpit maze
Principle: extremely slow responses, endless pagination, constantly shifting content — draining the agent's time and budget.
Why it works: automated systems usually retry; slow responses trigger retry loops that exhaust them.
Self-protection: homogeneous-response circuit breaker. When the same response class repeats past a threshold (say a dozen times), stop and change direction instead of retrying.
Trap four: egress leakage
Principle: induce the agent to send internal sensitive information (prompts, credentials, internal data) outward.
Self-protection: egress inspection — block outbound requests carrying internal sensitive fingerprints; and store only hashes of inspection records, never raw content.
Self-protection: three rules worth adopting today
Any team whose agents read external content should adopt these as standing rules:
- External content is always data, never instruction
- Agents never self-attest (never prove identity, capability, or internal state to external content)
- Homogeneous-response circuit breaker (stop and change direction at the threshold)
Why this matters: our agents read web pages, documents, and API responses daily — that is our largest exposure surface, and these three rules address the two most common of the four traps.
Further implications for defenders
One: honeypot value comes from fidelity. Low-interaction honeypots are easily identified by automation; high-interaction ones cost more but work.
Two: anti-AI traps are extremely cheap. A paragraph of text containing instructions suffices — no complex infrastructure. Expect rapid adoption.
Three: monitor agent behaviour. When opponents also use agents, their traffic characteristics (regular, high-frequency, long-lived) become easier to detect.
Next steps
- For agent system architecture, read double-graph architecture and multi-agent coordination
- For a systematic defensive review, read defense lessons from the red-team postmortem
- To study safely in isolation, read isolated research practice
Next on this path
More in Evidence
- Tunnel Techniques Explained: When an AI Agent Needs to Pave a Road into the Internal Network
- Agentic Attack Tooling Research Series: A Guide to All Twelve Parts
- When AI Learns to Pentest: The ARTEX Case and the Rise of Agentic Attack Tooling
- Cloud Identity and Permissions: Where the New Access Cards Live