Read these first
This article assumes the following earlier in its learning path.
The Double-Graph Architecture: Separating "What the Target Is" from "How Far We Have Tested"
Core question: For a long-running autonomous agent, the biggest enemy is rarely a lack of capability — it is state contamination: writing a guess into the "facts" and then churning along a false premise. Angle of this article: architecture design analysis, using a public agentic system as the sample.
Why one table is not enough
Imagine an agent system executing a long, complex task. It must record two very different kinds of information:
The first is "what the world actually is": which assets exist, which services run, how they relate. This information is shared across tasks — a fact confirmed today still holds tomorrow.
The second is "how I got here": which directions I tried, what each attempt was based on, and what it produced. This belongs only to the current task, and much of it is intermediate — many attempts turn out to be dead ends.
Most systems put both into one table. Three classic symptoms follow:
Symptom one: guesses become facts. The agent writes "looks injectable" mid-step; later steps treat it as a confirmed condition and reason onward. The error compounds.
Symptom two: conclusions cannot be traced. Ask later "what was this conclusion based on?" and there is no lineage — only a flat list of records.
Symptom three: duplicated work across tasks. Every new task re-confirms the same asset information from scratch, because facts and per-task process are tangled together and cannot be separated.
The solution: two graphs, two responsibilities
The design in question splits the two kinds of information into two independent graphs.
The asset graph: a globally shared truth store
Its nodes are objective entities: root domains, subdomains, IPs, services, applications, endpoints.
Its characteristics:
- One copy across all tasks — different tasks see the same asset truth;
- Parent-child relationships are computed by code, not filled in by the model — domain to subdomain, subdomain to service, service to endpoint, plus deduplication keys, are all determined programmatically;
- The agent only submits raw information; it does not decide where an asset attaches.
This is a crucial choice: take structural work out of the model's hands. Models are good at judgement and reasoning, not at maintaining a consistent ID system. Let code do what code does well.
The exploration graph: a per-task reasoning chain
Its nodes are subjective process: goals, intents, facts, findings, hints.
It is wired by lineage relations:
| Edge | Meaning |
|---|---|
spawns | A goal spawns an intent (a direction) |
yields | An intent yields a fact |
derived_from | A direction derives from existing facts |
proves | An intent ultimately proves a finding |
With this graph, any conclusion can be traced: which direction and which facts produced it.
The key mechanism: anchors joining the two graphs
If the graphs were fully independent, they could not answer the most common practical question: "Has this asset been tested, and how far?"
Hence anchors: a mapping table that pins exploration-graph nodes (intents, facts, findings) onto specific assets in the asset graph.
Two directions of query become possible:
- From direction to assets: which assets did this intent target?
- From asset to history: which intents tested this endpoint, and what facts came out?
This is also the basis for asset test coverage and coverage visualisation — seeing at a glance which assets are tested and which are untouched.
Three benefits
Benefit one: guesses and facts are physically separated. Unverified judgements stay in the exploration graph (refutable, markable as inferred); only confirmed information settles into the asset graph. State contamination is blocked architecturally.
Benefit two: traceability. Every conclusion has a lineage path, so you can answer "why was this decided" rather than just "the system did this".
Benefit three: reuse. The asset graph is shared across tasks, so new work does not restart from zero.
The pattern generalises beyond security
Rename the parts and the architecture maps onto many domains:
| Domain | Truth store (shared) | Reasoning chain (per task) |
|---|---|---|
| Investment due diligence | Companies / shareholders / financial facts | The path and leads of this investigation |
| Academic research | Papers / authors / citations | Hypotheses and verification of this study |
| Customer support automation | Products / customers / orders | Reasoning and attempts of this conversation |
| Software debugging | Code structure / dependencies | The hypothesis chain of this investigation |
The common thread: "what the world looks like" and "how I investigated" are fundamentally different information, and mixing them causes mutual contamination.
An important caveat
This design is not free:
- Two data structures must be maintained, more complex than a single table;
- Anchor maintenance costs: pinning a node to an asset is an extra write each time;
- It is over-engineering for small tasks — if a task touches a few targets and never repeats, a single table is simpler.
The test for adopting it: is the task long enough, will it repeat, will shared facts accumulate? If all three are true, splitting pays off.
Next steps
- To see how these two graphs power coordination between agents, read multi-agent coordination
- To enter from the case context, read the case study first
- To learn how to study such systems safely in isolation, read isolated research practice
Next on this path
More in Evidence
- Tunnel Techniques Explained: When an AI Agent Needs to Pave a Road into the Internal Network
- Agentic Attack Tooling Research Series: A Guide to All Twelve Parts
- When AI Learns to Pentest: The ARTEX Case and the Rise of Agentic Attack Tooling
- Cloud Identity and Permissions: Where the New Access Cards Live