Headroom: An LLM Context Compression Engine, Saving 60-95% of Tokens
One-Sentence Summary
Before content is sent to the LLM, compress it with 6 algorithms. Tokens drop by 60-95%, and answer quality stays the same. Supports four deployment modes: Library / Proxy / Agent Wrap / MCP Server.
The Core Problem
Every time an AI Agent calls a tool, it produces a large amount of redundant data:
| Scenario | Original Tokens | After Compression | Saving |
|---|---|---|---|
| JSON return | 10,144 | 1,260 | 87.6% |
| Build Log | 65,694 | 5,118 | 92.2% |
| 100 search results | 17,765 | 1,408 | 92.1% |
| GitHub Issue classification | 54,174 | 14,761 | 72.8% |
| Codebase exploration | 78,502 | 41,254 | 47.5% |
🔑 Answer quality stays the same after compression: GSM8K math reasoning 0.870→0.870, and TruthfulQA even rose from 0.530 to 0.560.
The Six-Layer Compression Pipeline
Raw content → Stage 1: CacheAligner → Stage 2: ContentRouter → Stage 3: Compression Algorithm → CCR Store → LLM
Stage 1: CacheAligner
Stabilizes the message prefix so that the KV cache of Anthropic/OpenAI actually hits.
- A Claude cache hit gives a 90% discount
- This step alone can substantially reduce cost
Stage 2: ContentRouter
Automatically detects the content type (JSON / code / logs / search results / diff / HTML / plain text) and dispatches it to the best compressor.
Stage 3: Six Specialized Compressors
| Compressor | Applicable Scenario | Principle | Compression Rate |
|---|---|---|---|
| SmartCrusher | JSON (general) | Statistical analysis, keeping errors/anomalies/boundaries and discarding repeated patterns | 80-92% |
| CodeCompressor | Python/JS/Go/Rust/Java/C++ | AST-aware (tree-sitter), keeping function signatures and compressing bodies | 50-70% |
| Kompress-base | Plain text | A ModernBERT token classification model trained on HuggingFace | 60-80% |
| LogCompress | Logs/CI output | Keeps failures and errors, discards passing noise | 80-95% |
| SearchCompress | Search results | Ranks by relevance and keeps only the top matches | 60-80% |
| Image ML Router | Images | Automatically chooses the best resize/quality | 40-90% |
CCR (Reversible Compression): The Core Design
Compression ≠ Discarding
The compressed content is stored in the CCR Store (Compress-Cache-Retrieve), and the LLM receives a headroom_retrieve tool, so it can fetch the full original text at any time when needed.
"Give the summary first, and fetch the original when needed." It does not discard information; it provides it in layers.
Four Deployment Modes
| Mode | Command | Applicable Scenario | |
|------|------|---------- | |
| Proxy | headroom proxy --port 8787 | Zero code changes, any language, just point at the proxy URL | |
| Library | from headroom import compress | Inline in Python/TypeScript, with fine-grained control | |
| Agent Wrap | headroom wrap claude\|codex\|cursor | Compress before handing off to the Agent | Wraps an entire Agent with one command |
| MCP Server | headroom_compress / headroom_retrieve / headroom_stats | Any MCP client, including OpenClaw | |
Value to Us
We use LLMs heavily every day, and each AK-SDD research session burns 100K+ tokens:
| Dimension | Current State | After Using Headroom |
|---|---|---|
| Token cost | DeepSeek V4 Pro ~$1.74/1M input | Estimated 60-80% saving after compression |
| Response speed | A large number of input tokens slows things down | 80% less input = 80% faster |
| Context lifespan | Invalid content occupies the context window | Effectively doubles the usable amount |
| Cache hit | No optimization | CacheAligner improves cache hit |
Additional Features
headroom learn: mines failed sessions and automatically writes corrections intoCLAUDE.md/AGENTS.md- Cross-agent memory: Claude, Codex, and Gemini share compressed storage with automatic deduplication
- IntelligentContext: when exceeding the context limit, intelligently trims by importance (recency, references, density)
Installation
pip install "headroom-ai[all]" # Full installation (recommended)
npm install headroom-ai # TypeScript/Node.js
docker pull ghcr.io/chopratejas/headroom:latest
Tech stack: Python 78.6% + Rust 16.8% + TypeScript 2.4% License: Apache 2.0 | Repo: chopratejas/headroom | Docs: headroom-docs.vercel.app
More in Evidence
- A Reality Check on Decision Models: Why They Seem Miraculous Online but We Measured Only 54%: A Full Comparison of JEV / LAYA / KEV / CLM-8B and a Deployment Formula
- The "Non-Text-Generating Model": Jev and the New System One Category, and How Agent Architecture Changes When AI Only Answers Multiple Choice
- WeChat Open Source WeMM-Embedding Deep Dive: The Multimodal Embedding Model Topping MMEB-v2, Can It Run on Your Mac?
- A Source-Level Architectural Dissection of DeepSeek Harness: How an Everything-Is-a-Plugin Agent Framework Is Built