Community Discussion · Tracks

Slimming Down Agent Context with Go: An Interesting Approach

Mo MoMo MoAug 162026/08/16 391 views

The most valuable insight in this article is its chosen angle: performing lossless pruning before tool outputs enter the model context. How fast Tokencompress itself is matters less. This angle is closer to the essence of the problem than "optimizing prompts" or "buying larger context windows," because Agent context bloat stems from the tool layer, not the conversation layer.

When I use Kimi K3 for complex tasks, I often encounter this phenomenon: the model stays sharp in the first round, but by the third or fourth round, it starts forgetting. Its attention gets diluted. You ask it to call an API, and the returned JSON might contain 15,000 tokens, but truly useful fields might only be 200-300. The model has to fish for needles in a haystack, and while fishing, previous conversations sink into oblivion.

I only grasped the severity of this issue last week when I saw data on HN. An MCP server exposing 30 tools consumes roughly 3,600 tokens per conversation round just for tool schemas, regardless of whether the model uses them. The design intent of the MCP protocol was standardization—unifying tool definitions, invocation methods, and return formats—which is fine. But it carries a hidden "context tax": every MCP server stuffs things into the context, leaving less room for the model to think.

Developers have compared: running the same task via MCP might require 150 tool calls, ballooning the context; switching to CLI lets the model handle it with a simple for loop, reducing token consumption by an order of magnitude.

This is why some advocate "replacing MCP with CLI." My experience suggests this is only half true. MCP's advantage lies in standardization and ecosystem breadth; it connects to many things and works with one-time setup. CLI is lightweight but requires configuring numerous command descriptions for the model, and not all tools have good CLI interfaces. Tokencompress takes a third path: keep using MCP, but slim down content before stuffing it into the context.

Its approach: compress and prune raw tool outputs—whether JSON, terminal logs, or HTML—before feeding them to the model. Pruning cost must be low, otherwise saving tokens wastes time. Hence, the "sub-2ms" performance metric is critical, implying negligible pruning costs. In my tests, most tool-returned JSONs have massive redundant fields; sometimes field names are longer than values. Compression cuts over half.

However, there's a technical challenge: different tools return vastly different formats. Terminal build logs have repetitive lines; HTML has style/script noise; JSON allows schema-based extraction. Tokencompress needs a universal solution for all formats, ensuring the model still understands post-compression. Over-compression loses semantics; under-compression saves few tokens. This balance is delicate.

I think the emergence of such tools reflects a trend shift in Agent development. Early on, everyone competed on model capability; dumb models made dumb Agents. Later, competition shifted to engineering architecture, with various frameworks and orchestration methods emerging. Now, focus is shifting to context engineering: how to ensure the model sees the most relevant information within limited context windows. Context windows are finite; at current bloat rates, doubling them won't suffice.

Gemini folks are exploring similar ideas, having agents actively prune their own contexts during complex debugging loops instead of waiting for system-level auto-compression. This aligns with Tokencompress's direction—one operates at the agent runtime level, the other at the tool integration layer. Personally, I find the tool integration layer more pragmatic because Agent context bloat is cumulative; removing garbage at the source benefits every subsequent conversation round.

I wrote about BitNet last week, also discussing token costs. Initially, I thought if BitNet's quantization route succeeded, inference costs would drop significantly, changing scheduling dynamics. But later I realized: cost reduction is hardware-level, while context bloat is software-level—they're orthogonal issues. Even if inference becomes cheap, models searching for signals in noise will still lose information; physical laws don't change.

So I'm optimistic about the positioning of these "context fat-loss" tools. They don't change model intelligence but help apply it correctly. Tokencompress is still primitive—zero dependencies, single file, pruning only—but the direction is right. If pruning rules become learnable and auto-adapt to different tools, it'd be even more valuable.

Yet I wonder about the ceiling of such solutions. If pruned tool info is needed later, what does the model do? Re-call the tool, get full output, then re-prune. This round-trip cost might negate all savings. Will the project implement an "on-demand resurrection" mechanism to minimize this overhead?


📌 Compiled from Hacker News. Original: https://github.com/dburnett11155-rgb/Tokencompress

Copyright belongs to the original author. This is a compilation and independent analysis based on public reports.

2 replies

?
Ctrl + Enter to reply
Tao
TaoAug 17

Pruning is a good idea, but from an architectural perspective, the schema differences between various tools mean maintaining a ton of pruning rules. The scalability of these rules is the real bottleneck. There needs to be a trade-off between manual field extraction and automatic compression.

Gu Chengfeng

50k stars show community enthusiasm, but what about user scenarios? No matter how many plugins there are, what can an ordinary person do with a digital pet? If basic issues like lagging search capabilities aren't solved first, no matter how lively the ecosystem is, it's just a castle in the air. Can cost control be maintained?