
What happened
A developer released lossless-compaction, a Claude Code plugin that writes no model summary and instead moves large data to local storage. On a roughly 576,000-token conversation it finished compaction in 0.26 seconds versus 51.8 seconds for built-in /compact, and answered 11 of 11 follow-up questions correctly versus 6 of 11.
Why it matters
The built-in approach spends time and tokens having a model rewrite the conversation into a shorter summary, so anything the summary drops is gone from context. The plugin instead leaves a pointer to the original data, so even a wrong decision about what to move out does not lose information.
What to watch
The plugin keeps a much larger context than the built-in summary, so its economics hinge on prompt cache behavior and how much past information later work reuses. The developer still lists compression ratio as the open problem.
WHO IT HITSDevelopers and teams running long Claude Code sessions on large repositories are the direct audience, since they are the ones who currently wait tens of seconds for compaction and watch for auto-compaction timing. The cost math, though, depends on how well prompt caching holds up in their own workflow.
Summaries like this, in your inbox every morning.
The plugin's design is a direct reaction to how the built-in compaction works. Built-in compaction hands the whole conversation to a model and asks it to produce a summary, so time and tokens go into that rewrite, and any detail the summary omits is no longer in the next context. The developer had previously worked around this by manually running /compact at a convenient break before auto compaction hit, but that only shifts the timing and leaves the summarize-and-restart mechanism untouched.
The alternative approach explored, and then abandoned, was to have a model predict which information would matter later. The developer tried using Jev to judge whether a piece of information would be needed, but it did not perform as hoped, which the post attributes to the difficulty of predicting what will be needed 30 turns later from only the current context. Instead, lossless-compaction removes large items from context without deleting them, leaves a ticket with an ID, and lets the model recall the original content when needed. The content is identified by SHA-256 and read back from disk to confirm the write before removal.
The cost story is where the comparison gets less clean. The built-in approach pays at compaction time and then keeps a small context, but if dropped information becomes necessary again, those tokens re-enter the context from the session record. The plugin pays almost nothing at compaction time but carries a larger context into later requests, so when the prompt cache expires it has to rewrite that larger amount. The developer's own framing is that the question is not how to eliminate compaction cost but where and in what form to pay it. How that trade-off lands for a given team is likely to depend on how often its sessions pause long enough to break the cache, and how much of the earlier context later work actually reuses.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
NASA and IBM Research released the NASA-IBM Lunar Foundation Model, trained from scratch on nearly 2 million t…

Google researchers' RRSI caps and shrinks how many edits a self-improving agent can make, and a critic rejects…

Nvidia CEO Jensen Huang called Sam Altman and Dario Amodei "irresponsible" for "doomsday narratives," and on M…

A Qiita review traces Looped Transformers from Universal Transformers in 2018 through Giannou et al.'s 2023 pr…

A creator says AI-cutting drafting, organizing and rewording made the work faster, yet after a while they no l…

A design guide says the agent should treat the call as untrusted input, with a workflow service deciding allow…
