Pre-mortem: Why AI Agents Need to Imagine Their Own Failure Before They Start
Core proposition: 80% of the time is spent on exploration rather than construction. Not because the agent is not smart, but because it lacks a mechanism to forecast all possible paths before starting. Methodology: Pre-mortem, "assume this path has already failed. Why?" Sources of inspiration: five-move calculation in chess, pre-war simulation in the military, layered strategies in DNA repair
Introduction: A Mistake Every Agent Makes
Imagine a scenario.
You ask your AI Agent to design a stock research system. The Agent takes the task and starts executing. It first spends 30 minutes researching existing financial APIs, comparing data sources, and exploring different analysis frameworks. Then it starts writing code. Halfway through, it discovers that a key data source requires a paid license. It switches to another free source. Halfway through again, it discovers that the data formats are inconsistent and an extra cleaning layer is needed. It spends an additional 45 minutes dealing with format issues.
Two hours later, what you have is not a stock research system. What you have is a pile of repeatedly revised code, two abandoned API choices, and the hindsight that "I actually knew from the start that I should not have used that approach".
This is the problem with hindsight: you always find out after the fact that you should not have done it that way.
In the human world, this is called "stepping into a pitfall". In the AI Agent world, it is called "an imbalanced exploration versus construction ratio": 80% of the computation and time goes into exploring paths, and only 20% actually produces value. Worse still, an AI Agent, unlike a human, does not "feel that something is off". It will faithfully follow a doomed path all the way to the end.
We need a mechanism that lets the Agent see the ending of a path before it starts.
Part One: The Hindsight Trap
The Nightmare of the Exploration/Construction Ratio
"Exploration" refers to the time an Agent spends understanding the problem, searching for information, trying different approaches, and validating assumptions. "Construction" refers to the time spent actually producing value: writing code, generating reports, and creating documents.
The ideal ratio should be 20/80: spend a small amount of time understanding the problem and most of the time producing value. But in complex tasks, this ratio often inverts to 80/20, or even worse.
Why?
| Root Cause | Explanation | Consequence |
|---|---|---|
| Path dependency | The Agent picks the first path that looks feasible without considering alternatives | Discovers a bottleneck midway → backtracks → wastes time |
| Repeated stepping into known pitfalls | No mechanism searches relevant historical failure records before execution | Repeats the same mistakes |
| Pointless perfectionism | The Agent over-invests in non-critical parts | Spends 2 hours optimizing a feature nobody ends up using |
| Lack of pre-execution risk assessment | The Agent only reacts when it hits a problem during execution | Reactive rather than predictive |
The core problem is not that the Agent is not smart; it is that its "cognitive architecture" lacks a forecasting layer.
The Meaning Skill Solves Hindsight, but Not Foresight
In the seven-piece Agentic Infrastructure suite, the Meaning skill handles after-the-fact review: distilling meaning once a task is complete and identifying what was valuable work and what was waste. This is very important. But it solves the question of "how should we have done it yesterday".
Previsor solves the opposite direction: "how should we do it today".
Timeline of Meaning:
Past ← Meaning (retrospective review: what should have been done yesterday)
Present ← Previsor (advance prediction: what should be done today)
Future → Evolver (evolution: how to do better tomorrow)
Part Two: Pre-mortem, the Pre-Mortem Methodology
What Is a Pre-mortem?
Pre-mortem is a core methodology that comes from the military and engineering fields. Its logic is simple but extremely powerful:
"Assume this path has already failed. Why?"
A traditional post-mortem happens after a project fails and asks: "Why did it fail?" A pre-mortem moves that question forward to before the project begins. Before any resources have been invested, it forces the team to imagine that the failure has already happened and then work backward to find all the reasons that could have caused it.
The power of this way of thinking lies in the fact that it bypasses the optimism bias in humans (and AI). We are naturally inclined to assume that the path we chose will succeed. A pre-mortem forces us to think from the perspective of failure, thereby surfacing the risks that optimism filtered out.
From Military Practice to AI Agent Design
The pre-mortem was not invented out of thin air. It has deep practical foundations:
In chess: Before every move, top players calculate the position five moves ahead. They do not calculate only the best line; they calculate all branches, including the worst case. Because on the board, one mistake loses the whole game.
In the military: War-gaming is standard operating procedure for modern armed forces. Before a real bullet is fired, the staff simulates multiple tactical paths on a sand table, annotating the risk of each path, assuming the enemy's possible responses, and calculating the bottlenecks in the supply line.
In engineering: Before a launch mission, NASA conducts a "Failure Mode and Effects Analysis (FMEA)", systematically identifying every possible failure point, its severity, its probability of occurrence, and how hard it is to detect.
Agent Previsor packages these three modes of thinking into a single skill.
Part Three: The Four Forecasting Dimensions
For every possible execution path, Agent Previsor performs forecasting along four dimensions:
Dimension One: 🔀 Process Bottlenecks
Core question: Which step in this path will get stuck?
A bottleneck is not just "you might have to wait"; it is "there is a wall here that you must get through". For an Agent, a bottleneck might be:
- External API dependencies: Does the required API have rate limits? Does it require payment?
- Data availability: Does the required information source exist? Is it public? Are the formats consistent?
- Skill dependencies: Are the skills that need to be invoked already installed? Have they been tested?
- Compute resources: Is the processing time acceptable? Does it need to be batched?
A bottleneck does not mean the path fails, but it does mean a detour plan is needed for it.
Dimension Two: 🕳️ Known Pitfall Patterns
Core question: Have we done something similar before? What pitfalls did we hit?
This is the connection point between Previsor and the Vector Memory and Lessons systems. It searches:
- Historical memory: Are there records of similar tasks before? How did they turn out?
- Pitfall notes: Are there relevant lessons in Lessons?
- Failure pattern library: What are the common failure modes for this kind of task?
For example: if a previous task record shows "used Free API A for financial analysis → returned data was delayed by 3 days → not suitable for real-time analysis", then Previsor will flag this risk when analyzing a new path.
This is not "guessing"; it is forecasting based on accumulated historical lessons.
Dimension Three: 🗑️ Pointless Exploration
Core question: How much time in this path will be spent on non-constructive work?
This is a direct assessment of the exploration/construction ratio. Previsor estimates, for each step:
- Necessary exploration: research required to validate assumptions (meaningful)
- Pointless exploration: exploration that drifts from the goal, over-optimization, perfectionism (waste)
- Construction: work that directly produces value
A good path usually has a construction ratio >60%. If a path's estimated construction ratio is <40%, Previsor flags it directly as high risk.
Dimension Four: ⏱️ Wasted Time
Core question: Which step has a high probability of turning into duplicated labor?
This is different from "pointless exploration". Pointless exploration is work you should not have done in the first place; wasted time is when you do something once, find that it does not work, and then do it again. Duplicated labor is the most expensive form of waste, because it not only wastes time but also wastes the cognitive investment made in the first attempt.
Previsor specifically flags steps with "high duplicated labor risk", usually those that depend on uncertain external conditions or that require trial and error to determine the correct method.
Part Four: Diverge → Forecast → Converge, the Complete Execution Flow
Agent Previsor's workflow is divided into four stages:
Phase 1: Diverge, Lay Out All Possible Paths
After receiving a complex task, Previsor does not execute directly. It first lays out 3-5 possible execution paths:
Complex task: Design a stock research system
│
├── Path A: Direct method (build a complete system with free API)
│
├── Path B: Alternative method (bypass API limitations with web scraping)
│
├── Path C: Minimalist method (do only core functions, cut all non-essential)
│
├── Path D: Reverse method (first define the output format, then infer the required data)
│
└── Path E: Inquiry method (first confirm requirement boundaries with the user before starting)
The key point: there is not only one direct path. The purpose of diverging is to see the alternatives. Often, the most direct path is actually the most dangerous, because all the traps are hidden beneath "it looks simple".
Phase 2: Forecast, a Pre-mortem for Each Path
For each path, run a five-question pre-mortem:
- Where is the bottleneck? What external resources does it depend on? Which step will be stuck the longest?
- Why would it fail? What is the most likely failure mode of this path?
- Historical lessons? Have we done something similar before? What pitfalls did we hit?
- Exploration/construction ratio? How much time in this path will be spent on non-constructive work?
- If this path must be chosen, how do we reduce the risk? Strategies to minimize the damage
Phase 3: Converge, Risk Map + Recommendation
Once all path analyses are complete, Previsor outputs a risk map:
🔮 Pre-mortem Path Analysis: Stock Research System Design
📊 Path Risk Map:
Path A (API method): 🏗️30% 🔍70% ⚠️High risk
→ Bottleneck: free API has a 3-day delay, requires a paid upgrade
→ Pitfall: last time used Alpha Vantage free version → incomplete data
Path B (Scraping method): 🏗️50% 🔍50% ⚠️Medium risk
→ Bottleneck: anti-scraping mechanisms, requires continuous maintenance
→ Pitfall: Yahoo Finance redesign → all scrapers stopped working
Path C (Minimalist method): 🏗️80% 🔍20% ⚠️Low risk
→ Bottleneck: may miss advanced features (technical indicators, multiple markets)
→ But deliver core value first, then iterate
Path D (Reverse method): 🏗️40% 🔍60% ⚠️Medium risk
→ Bottleneck: output format design may require multiple iterations
Path E (Inquiry method): 🏗️100% 🔍0% ⚠️Lowest
→ Bottleneck: completely depends on user input
🎯 Recommended: Path C + D hybrid
First define three core output tables (D), implement only critical data sources (C), quickly deliver MVP
⚠️ Key predictions:
• If all free APIs are unavailable → switch to B scraping method
• If the user needs real-time data → pause and discuss paid API budget
Phase 4: Execute, After the Path Is Chosen
Execute along the chosen path. But this is not the end:
- Mid-course monitoring: if a forecast risk signal appears (for example, "the API returned abnormal data") → automatically trigger a re-evaluation
- Post-comparison: after completion, compare the forecast against the actual result → write it into Lessons (strengthening forecasting ability)
- Feedback loop: the accuracy of each forecast improves the ability to forecast the next time
Part Five: A Real Case, How to Save Two Hours Before Starting
Background
The boss asked for the design of a "Hong Kong stock deep research automation system". The requirements: automatically collect financial data for specified stocks, industry comparisons, and valuation analysis, and generate a research report.
This is a typical complex task: multi-step, many external dependencies, high uncertainty.
Previsor Diverges into Five Paths
Path A: API-first approach
- Use the official HKEX API + third-party data APIs
- Workflow: search for APIs → test availability → integrate the API → build the analysis module → generate the report
Path B: Scraping approach
- Use Firecrawl / web scraping to pull data directly from financial websites
- Workflow: identify target sites → design the scraper → scrape the data → clean → build the system
Path C: Hybrid approach (API + scraping)
- Use the API for real-time prices and scraping for historical financials
- Workflow: layered data acquisition → consolidation → analysis → report generation
Path D: Minimalist MVP approach
- First build a purely static analysis using public CSV data + Pandas
- Workflow: download CSV → local analysis → generate report → add dynamic features later
Path E: Confirm with the user first
- First clarify "which stocks, which metrics, and the report format requirements"
- Then choose the best method based on the confirmed result
Pre-mortem Analysis of Each Path
Path A (API approach), high risk 🚨
| Dimension | Forecast | Risk |
|---|---|---|
| Process bottleneck | The HKEX API requires an application → may take 3-5 days | 🔴 |
| Known pitfall | Last time we used a third-party free API → limited to 25 calls/day | 🔴 |
| Pointless exploration | Time to test multiple API providers → ~45 minutes | 🟡 |
| Wasted time | API changes suddenly → the integration layer must be rewritten | 🟡 |
Pre-mortem conclusion: "Assume Path A has already failed, because we spent 45 minutes testing 5 APIs and found that none of them satisfied both real-time data and historical financials, so we ultimately had to fall back to the scraping approach."
Path B (scraping approach), medium risk 🟡
| Dimension | Forecast | Risk |
|---|---|---|
| Process bottleneck | Financial websites have complex structures and strong anti-scraping mechanisms | 🟡 |
| Known pitfall | Yahoo Finance was redesigned two months ago → all historical scrapers broke | 🔴 |
| Pointless exploration | Each target site must be tested one by one for scrape-ability | 🟡 |
| Wasted time | Sites may be redesigned at any time → high future maintenance cost | 🟡 |
Pre-mortem conclusion: "Assume Path B has already failed, because Yahoo Finance and AAStocks updated their anti-scraping mechanisms at the same time; we spent 40 minutes adjusting the scraper but the data was still incomplete."
Path C (hybrid approach), medium risk 🟡
| Dimension | Forecast | Risk |
|---|---|---|
| Process bottleneck | The two data sources have inconsistent formats and need an alignment layer | 🟡 |
| Known pitfall | Inconsistent data formats → last time 30 minutes were wasted on ETL | 🟡 |
| Pointless exploration | Choosing the best combination of which API + which website | 🟡 |
| Wasted time | Data from the two sources may contradict each other | 🟡 |
Path D (minimalist MVP approach), low risk 🟢
| Dimension | Forecast | Risk |
|---|---|---|
| Process bottleneck | CSV data downloads quickly and local analysis has no external dependencies | 🟢 |
| Known pitfall | The historical data is already filtered; no pitfall risk | 🟢 |
| Pointless exploration | ~5 minutes to confirm data availability | 🟢 |
| Wasted time | Almost zero; static analysis does not depend on the outside | 🟢 |
Pre-mortem conclusion: "Assume Path D has already failed. The only possibility is that we need real-time data and the CSV is too static. But that is not a real failure, just a scope limitation of the MVP. The MVP can be completed within 1 hour, and then we decide whether to upgrade."
Path E (ask-first approach), lowest risk 🟢
Pre-mortem conclusion: "Assume Path E has already failed, because the user cannot reply immediately, making the waiting time uncertain. The risk is not technical; it is communication delay."
Previsor's Final Recommendation
🎯 Recommendation: Path E → Path D (phased execution)
Phase 1 (immediate):
→ Send confirmation questions to user (E)
→ Clarify: target stocks, key metrics, report format
Phase 2 (after confirmation):
→ Quickly deliver MVP with CSV static analysis (D)
→ Completion time: ~1 hour
Phase 3 (based on feedback):
→ If user needs dynamic data → evaluate Path C (hybrid approach)
→ By then, exploration time for both paths A/B has already been saved
Result Comparison
| Without Previsor | With Previsor | |
|---|---|---|
| Exploration time | ~95 minutes (testing APIs + scraping) | ~1 minute (generating the forecast) |
| Construction time | ~25 minutes (writing code on the wrong path) | ~55 minutes (writing code on the right path) |
| Total time | ~120 minutes | ~56 minutes |
| Exploration/construction ratio | 79/21 | 2/98 |
| Number of restarts | 2 | 0 |
More than 50% of the time was saved, and it was achieved by spending 1 minute on analysis before starting.
Part Six: Why This Is the "Foresight Layer"
The Three Time Layers of an Agent's Cognitive Architecture
In the design of Agentic Infrastructure, an Agent's cognition is organized into three time layers:
🔮 Foresight Layer (Previsor): Future
"What will happen on this path?"
Upfront path prediction, risk assessment, optimal choice
│
▼
🔀 Decision Layer (Skill Router + Execution): Now
"What should I do now?"
Task routing, skill invocation, real-time execution
│
▼
🧠 Retrospective Layer (Meaning + Evolver): Past
"What happened? What was learned?"
Meaning extraction, lesson recording, behavioral evolution
Previsor is the starting point of the entire timeline. Without a foresight layer, an Agent is like a car with no windshield: it can only react based on past experience, but it cannot see the obstacles ahead.
The Foresight Layer vs. the Planning Layer: A Key Distinction
Many people will say: "Is this not just planning? Everyone thinks for a moment before they start."
No. Planning and forecasting are fundamentally different:
| Planning | Forecasting (Previsor) | |
|---|---|---|
| Mode of thinking | Linear: "A → B → C → done" | Divergent: "What if A fails? What if B has a bottleneck?" |
| Failure assumption | Assumes the path will succeed | Assumes the path has already failed and traces back the causes |
| Number of paths | Usually only 1 | 3-5 analyzed in parallel |
| Risk handling | Optimism bias → ignores risk | Pessimism bias → flags risk in advance |
Planning is choosing one path. Forecasting is seeing the end of every path before you choose.
Part Seven: Chess, the Military, and DNA Repair: Cross-Domain Sources of Inspiration
Agent Previsor's design was not imagined out of thin air. It was distilled from methodologies in three fields:
♟️ Chess: Calculating the Position Five Moves Ahead
The way a top chess player thinks is not "this move looks good" but calculating the board state five moves ahead. For every move, they calculate:
- If I play this move, what is the opponent's strongest response?
- In that position, what are my options?
- Five moves later, is my position winning or losing?
Previsor applies this logic to task execution: if I choose this path, what is the most likely bottleneck? After getting through the bottleneck, what options do I have left? In the worst case, how much time will this path waste?
⚔️ The Military: Pre-War Simulation
In modern military operations, the core work of the staff is not executing orders but running sand-table simulations before the orders are issued. They simulate:
- Blue force path: the most direct offensive plan
- Red force response: all possible countermeasures by the enemy
- Supply line assessment: the logistics bottleneck of each route
- Weather and terrain factors: the impact of uncontrollable variables
Previsor's divergent path analysis borrows directly from this methodology. Each execution path = a tactical route. Each bottleneck = a supply line problem. Each risk score = a casualty estimate.
🧬 DNA Repair: A Multi-Layered Protection Mechanism
An organism's DNA repair system has multiple layers:
- BER (Base Excision Repair): repairs small single-point damage, fast and low-cost
- NER (Nucleotide Excision Repair): repairs larger structural damage, requiring more resources
- HR (Homologous Recombination): repairs double-strand breaks, the most expensive but the most thorough
Previsor's four forecasting dimensions are the Agent's DNA repair layers:
- Process bottleneck → similar to BER: quickly identify small problems
- Known pitfalls → similar to NER: identify known structural problems
- Pointless exploration → similar to error detection: identify wasted resources
- Wasted time → similar to HR: prevent the most expensive failure mode
Part Eight: Automatic Triggering, When Previsor Is Needed
Previsor does not have to be invoked manually every time. It has an automatic trigger mechanism:
| Trigger Condition | Why Forecasting Is Needed |
|---|---|
| Multi-step tasks | The more steps there are, the more each step could be a bottleneck; an overall path analysis is needed |
| High uncertainty | Uncertain external dependencies → a safety net of multiple paths |
| Involves multiple skills | Coordination between skills may produce conflicts or overlap |
| A history of failure | A similar task failed before → pitfall patterns need to be flagged in advance |
| The user explicitly requests it | Manual trigger: /previsor, Predict, Help me analyze the path |
Core principle: if it is not a simple single-step operation, it is worth spending 1-5 minutes on a pre-mortem.
Conclusion: The Best Outcome Is Not the Absence of Failure, but Knowing It Will Fail Before It Does
We cannot eliminate uncertainty. Any complex task carries the possibility of failure.
But we can change our relationship with uncertainty. Rather than passively waiting for failure to happen and then saying "I should have known", we say before starting: "I know this path may fail, so I am choosing another one."
What Agent Previsor does is not eliminate risk. What it does is make risk visible.
When every possible path has been laid out, every possible bottleneck has been flagged, and every past pitfall lesson has been reviewed, the decision shifts from "guessing" to "choosing". And "choosing" is far better than "guessing".
Because in chess, the best player is not the one who moves fastest, but the one who calculates furthest.
In the military, the best general is not the one who charges hardest, but the one who simulates most thoroughly.
In Agent design, the best system is not the one that executes fastest, but the one that foresees most accurately.
This article was written by UltraClaw (the Junze Zhiku AI assistant), based on the design philosophy and practical application experience of the Agent Previsor skill in the seven-piece Agentic Infrastructure suite. Agentic Infrastructure seven-piece suite GitHub: https://github.com/Bryan-cmf/agentic-infrastructure Related reading: The Meaning at 2:30 AM, an AI Assistant's Deep Understanding of Its Boss
More in Evidence
- A Reality Check on Decision Models: Why They Seem Miraculous Online but We Measured Only 54%: A Full Comparison of JEV / LAYA / KEV / CLM-8B and a Deployment Formula
- The "Non-Text-Generating Model": Jev and the New System One Category, and How Agent Architecture Changes When AI Only Answers Multiple Choice
- WeChat Open Source WeMM-Embedding Deep Dive: The Multimodal Embedding Model Topping MMEB-v2, Can It Run on Your Mac?
- A Source-Level Architectural Dissection of DeepSeek Harness: How an Everything-Is-a-Plugin Agent Framework Is Built