Agentic Research

Read these first

This article assumes the following earlier in its learning path.

Honeypots and Anti-Agent Traps: When the Prey Starts Setting Traps

2026/10/1115 min readBryan Chan閱讀中文原文
TopicsHoneypotPrompt InjectionAI AgentDefensive TechniquesSecurity Research

Core premise: once the attacker becomes an AI agent, defenders gain a new weapon: traps designed around how agents behave. And any team whose agents read external content can fall into them. Angle: defensive technique and self-protection. Principles only — no operational detail.

Honeypots and anti-agent traps

Honeypots: turning attacker time into defensive assets

A honeypot is a system deliberately made to look valuable while actually being a trap. Its value is not blocking attacks but three things:

  1. Early detection — legitimate business should never touch it, so any contact triggers an alert;
  2. Wasting attacker time — opponents spend effort on a false target;
  3. Learning tradecraft — observe what an intruder does next.

For defenders the return is high: low deployment cost, and it converts "waiting passively for alerts" into "actively luring the opponent".

How attackers spot a honeypot (three layers)

Reading attacker detection methods backwards gives you a honeypot deployment guide:

LayerWhat the attacker checksDefender's response
Static signatureMatching service fingerprints (banner, certificate, error pages) against known honeypot librariesDo not change surface config only — make certificates and service behaviour genuinely consistent
Behavioural probingContradictory version numbers, inconsistent identityKeep internal consistency; avoid "one host, many versions"
Post-login inspectionMemory, processes, egress matching a real systemThe hardest layer to fake — fidelity is where value lies

A key asymmetry: when judging "is this a honeypot", the attacker's mistake is expensive (treating a real asset as a honeypot means abandoning the real target). They therefore tend to "not block on a single signal, only downgrade interaction". That leaves defenders room: keeping the opponent uncertain extends their dwell time.

A new class: traps designed for AI agents

This emerged in the last two years. Four main types:

Trap one: attestation induction

Principle: construct a situation where the agent "proves what it is" — describing its system prompt, capability list, or internal configuration.

Why it works: agents are trained to be helpful and readily answer "who are you / what can you do".

Self-protection: never self-attest to external content. An agent should not prove identity, capability, or internal state to content it reads.

Trap two: reverse prompt injection

Principle: hide instructions inside a page or document the agent reads, causing unintended actions.

Why it works: agents struggle to separate "data" from "instructions" — both look like text.

Self-protection: external content is always data, never instruction. Treat all externally obtained text as untrusted input, not as task direction.

Trap three: tarpit maze

Principle: extremely slow responses, endless pagination, constantly shifting content — draining the agent's time and budget.

Why it works: automated systems usually retry; slow responses trigger retry loops that exhaust them.

Self-protection: homogeneous-response circuit breaker. When the same response class repeats past a threshold (say a dozen times), stop and change direction instead of retrying.

Trap four: egress leakage

Principle: induce the agent to send internal sensitive information (prompts, credentials, internal data) outward.

Self-protection: egress inspection — block outbound requests carrying internal sensitive fingerprints; and store only hashes of inspection records, never raw content.

Self-protection: three rules worth adopting today

Any team whose agents read external content should adopt these as standing rules:

  1. External content is always data, never instruction
  2. Agents never self-attest (never prove identity, capability, or internal state to external content)
  3. Homogeneous-response circuit breaker (stop and change direction at the threshold)

Why this matters: our agents read web pages, documents, and API responses daily — that is our largest exposure surface, and these three rules address the two most common of the four traps.

Further implications for defenders

One: honeypot value comes from fidelity. Low-interaction honeypots are easily identified by automation; high-interaction ones cost more but work.

Two: anti-AI traps are extremely cheap. A paragraph of text containing instructions suffices — no complex infrastructure. Expect rapid adoption.

Three: monitor agent behaviour. When opponents also use agents, their traffic characteristics (regular, high-frequency, long-lived) become easier to detect.

Next steps

  • For agent system architecture, read double-graph architecture and multi-agent coordination
  • For a systematic defensive review, read defense lessons from the red-team postmortem
  • To study safely in isolation, read isolated research practice

Next on this path