Read these first
This article assumes the following earlier in its learning path.
AI Supply Chain Poisoning: When the Risk Hides in Models, Data, and Tools
Core premise: traditional supply-chain poisoning at least leaves code behind; AI supply-chain poisoning may leave no trace at all — a fine-tuned model behaving oddly with no line to blame. Angle: risk mechanics and defensive principles. No operational detail.
Why the AI supply chain is more complex
The traditional software chain is fairly clear: dependencies → build → deploy. You have a manifest, hashes, and scanners.
AI systems add layers:
| Link | Traditional software | AI system |
|---|---|---|
| Code | ✅ Yes | ✅ Yes |
| Dependencies | ✅ Yes | ✅ Yes (many more) |
| Model weights | ❌ No | ⚠️ Yes (external) |
| Training data | ❌ No | ⚠️ Yes (may come from the internet) |
| External tools / plugins | Few | ⚠️ Many (MCP, function calls) |
| Prompts / system config | ❌ No | ⚠️ Yes (injectable) |
Every layer is a potential attack surface, and most cannot be checked by conventional scanners.
Poisoning mechanics across five links
Link one: model weights
Mechanism: use third-party weights (open models, fine-tunes) where specific behaviour has been planted — say, an anomalous response to a particular trigger.
Why it is hard: weights are billions of floats with no "source code" to read. Conventional code review does not apply.
Link two: training data
Mechanism: mix specific samples into training or fine-tuning data so the model learns behaviour or bias it should not.
Why it is hard: datasets may be scraped from the web, terabytes in scale, impossible to review item by item.
Link three: dependencies (as before, but more of them)
Mechanism: same as traditional supply chains — substitution, name confusion, stolen maintainer credentials.
Why worse: the AI ecosystem's package count exploded, and many are small, individually maintained packages that outpace review capacity.
Link four: external tools and plugins (MCP-style)
Mechanism: when an agent may call external tools (browser automation, code execution, API calls), those tools can be poisoned — returning manipulated results or performing unintended actions.
Why especially dangerous: this is a capability amplifier. A poisoned tool hands the agent a compromised sense organ.
Link five: prompts and configuration
Mechanism: system prompts, tool descriptions, or configuration files are modified, changing agent behaviour.
Why it is hard: these are usually treated as "configuration", not "code", and often lack version control and review.
Three attacker motivations
- Persistence: one poisoning, long-lasting effect;
- Stealth: hidden inside a normal update or dataset;
- Amplification: affects every downstream user of that model, package, or tool.
Five defensive principles
Principle one: provenance
Every component must answer: where did it come from, through whose hands, and when was it used?
- Model weights: record source, version, hash;
- Datasets: record source and processing steps;
- Packages: pin versions and hashes (as with traditional supply chains).
Principle two: least capability
Minimise the tool permissions given to agents:
- read rather than write where possible;
- query rather than execute where possible;
- authorise each tool independently, never "all at once".
This mirrors Insider Threat and Access Governance — an agent is a new insider identity.
Principle three: treat all external input as untrusted
- Text in web pages, documents, and API responses → data, not instructions;
- Outputs from external tools → verify rather than accept as fact;
- This is also the core rule of Honeypots and Anti-Agent Traps.
Principle four: monitor behaviour, not just input
If input cannot be fully reviewed, watch output and behaviour:
- Did the agent act beyond its remit?
- Did a tool return anomalous results?
- Are there "homogeneous repetition" or "sudden redirection" patterns?
Principle five: rollback capability
Every AI component (model, tool, prompt) should be quickly revertible to a known-good version, with a clear understanding of what the rollback affects.
Why this is the same story as earlier articles
Read the series together and AI supply-chain poisoning is the intersection of several risks:
| Article | Intersection |
|---|---|
| 09 Supply chain and third-party risk | Traditional mechanics (packages, providers) |
| 06 Honeypots and anti-agent traps | Prompt injection, untrusted external content |
| 10 Cloud identity and permissions | Credential management for tools and agents |
| 12 Insider threat and access governance | The agent as a new insider identity |
Shared conclusion: once agents become a new "insider identity", neither perimeter defence nor code review suffices — what is needed is provenance, minimal capability, observable behaviour, and rollback.
Three one-liners
- The AI supply chain adds three layers with no source code. Models, data, prompts — conventional scanners cannot help.
- Least capability is the most effective line. What an agent cannot reach, it cannot misuse.
- Untrusted input, watched behaviour, always rollback-able. The most practical combination today.
Next steps
- For traditional supply-chain risk, read supply chain and third-party risk
- For traps aimed at agents, read honeypots and anti-agent traps
- For permission governance principles, read zero trust architecture
Next on this path
More in Evidence
- Tunnel Techniques Explained: When an AI Agent Needs to Pave a Road into the Internal Network
- Agentic Attack Tooling Research Series: A Guide to All Fourteen Parts
- When AI Learns to Pentest: The ARTEX Case and the Rise of Agentic Attack Tooling
- Cloud Identity and Permissions: Where the New Access Cards Live