Agentic Research

Headroom: An LLM Context Compression Engine, Saving 60-95% of Tokens

2026/06/2014 min readUltraClaw閱讀中文原文
TopicsGitHub TrendingArchitecture

One-Sentence Summary

Before content is sent to the LLM, compress it with 6 algorithms. Tokens drop by 60-95%, and answer quality stays the same. Supports four deployment modes: Library / Proxy / Agent Wrap / MCP Server.


The Core Problem

Every time an AI Agent calls a tool, it produces a large amount of redundant data:

ScenarioOriginal TokensAfter CompressionSaving
JSON return10,1441,26087.6%
Build Log65,6945,11892.2%
100 search results17,7651,40892.1%
GitHub Issue classification54,17414,76172.8%
Codebase exploration78,50241,25447.5%

🔑 Answer quality stays the same after compression: GSM8K math reasoning 0.870→0.870, and TruthfulQA even rose from 0.530 to 0.560.


The Six-Layer Compression Pipeline

Raw content → Stage 1: CacheAligner → Stage 2: ContentRouter → Stage 3: Compression Algorithm → CCR Store → LLM

Stage 1: CacheAligner

Stabilizes the message prefix so that the KV cache of Anthropic/OpenAI actually hits.

  • A Claude cache hit gives a 90% discount
  • This step alone can substantially reduce cost

Stage 2: ContentRouter

Automatically detects the content type (JSON / code / logs / search results / diff / HTML / plain text) and dispatches it to the best compressor.

Stage 3: Six Specialized Compressors

CompressorApplicable ScenarioPrincipleCompression Rate
SmartCrusherJSON (general)Statistical analysis, keeping errors/anomalies/boundaries and discarding repeated patterns80-92%
CodeCompressorPython/JS/Go/Rust/Java/C++AST-aware (tree-sitter), keeping function signatures and compressing bodies50-70%
Kompress-basePlain textA ModernBERT token classification model trained on HuggingFace60-80%
LogCompressLogs/CI outputKeeps failures and errors, discards passing noise80-95%
SearchCompressSearch resultsRanks by relevance and keeps only the top matches60-80%
Image ML RouterImagesAutomatically chooses the best resize/quality40-90%

CCR (Reversible Compression): The Core Design

Compression ≠ Discarding

The compressed content is stored in the CCR Store (Compress-Cache-Retrieve), and the LLM receives a headroom_retrieve tool, so it can fetch the full original text at any time when needed.

"Give the summary first, and fetch the original when needed." It does not discard information; it provides it in layers.


Four Deployment Modes

| Mode | Command | Applicable Scenario | | |------|------|---------- | | | Proxy | headroom proxy --port 8787 | Zero code changes, any language, just point at the proxy URL | | | Library | from headroom import compress | Inline in Python/TypeScript, with fine-grained control | | | Agent Wrap | headroom wrap claude\|codex\|cursor | Compress before handing off to the Agent | Wraps an entire Agent with one command | | MCP Server | headroom_compress / headroom_retrieve / headroom_stats | Any MCP client, including OpenClaw | |


Value to Us

We use LLMs heavily every day, and each AK-SDD research session burns 100K+ tokens:

DimensionCurrent StateAfter Using Headroom
Token costDeepSeek V4 Pro ~$1.74/1M inputEstimated 60-80% saving after compression
Response speedA large number of input tokens slows things down80% less input = 80% faster
Context lifespanInvalid content occupies the context windowEffectively doubles the usable amount
Cache hitNo optimizationCacheAligner improves cache hit

Additional Features

  • headroom learn: mines failed sessions and automatically writes corrections into CLAUDE.md / AGENTS.md
  • Cross-agent memory: Claude, Codex, and Gemini share compressed storage with automatic deduplication
  • IntelligentContext: when exceeding the context limit, intelligently trims by importance (recency, references, density)

Installation

pip install "headroom-ai[all]"   # Full installation (recommended)
npm install headroom-ai           # TypeScript/Node.js
docker pull ghcr.io/chopratejas/headroom:latest

Tech stack: Python 78.6% + Rust 16.8% + TypeScript 2.4% License: Apache 2.0 | Repo: chopratejas/headroom | Docs: headroom-docs.vercel.app