Reduce coding agent token usage with a local compression layer
If you want to cut the heavy token costs that coding agents incur when reading tool outputs or logs, this tool compresses data in between. It acts as an intermediate step that compresses tool results or file contents locally before an AI agent sends them to an LLM. The original data remains safely stored locally and can be retrieved when needed, maintaining response quality while reducing costs.
The local compression process
This tool performs compression on the user's machine without sending prompts or file contents to external servers. It detects the content type and selects a dedicated compressor for JSON, source code, or plain text. JSON arrays and repetitive log lines often achieve over 90% compression, while already dense text sees lower rates. The compressed data is passed to the LLM, and if the model determines it needs the full original, it is immediately restored from the local cache.
Agent integration methods
Users can choose from three methods that fit their development environment. They can use a library approach to insert code directly into Python or TypeScript apps, or run a proxy server to route all requests without modifying code. Additionally, the headroom wrap command configures major coding agents like Claude Code, Cursor, and Codex in one go. This method starts a local proxy and automatically changes agent settings, and the headroom unwrap command can revert to the original state at any time. For MCP clients, a dedicated server allows access to the compression features.
Reducing output tokens
The tool reduces not only input data but also the response tokens generated by the model. It appends an instruction to answer concisely at the end of the system prompt to prevent the model from adding unnecessary introductions or repeating already provided context. It also automatically lowers reasoning effort after simple tool results, such as file reads or passing tests, to save costs. This feature supports both Anthropic and OpenAI compatible APIs, but it is disabled by default and requires setting an environment variable to activate. Output token savings are reported as estimates, and leaving some conversations as a control group allows for verification with measured figures.
Actual savings and accuracy
Benchmarks based on real MCP server output formats showed 21% token savings for code search and 57% for SRE incident debugging. Savings of 42% and 30% were also confirmed for codebase exploration and GitHub issue classification, respectively. In terms of accuracy, it recorded a score of 0.870 on the GSM8K math benchmark, identical to the baseline, and maintained 97% accuracy on SQuAD v2 and BFCL tool use benchmarks. The savings effect is larger when data is repetitive, while it is minimal for short conversations or already compressed text. To check the savings rate with your own actual traffic, run the headroom savings command.
Things to check before installing
This tool requires Python 3.10 or higher, with native wheels provided for macOS Apple Silicon and Linux. Intel macOS users must install via Docker, and in x86 environments, some features are disabled if the AVX2 instruction set is missing. If you use an SSL inspection proxy on a corporate network, certificate-related errors may occur, so you should configure the root certificate in advance. If you only use native compression features from a single provider and do not need cross-agent memory, this tool may not be suitable. It cannot be used in sandboxed environments where local process execution is restricted.