AIToday
Large Language ModelsAI Coding AssistantsOpen-Source AIZenn AI/MLPublished: Oct 4, 2026, 22:00 JST

lossless-compaction cuts Claude Code compaction from 51.8秒 to 0.26秒

lossless-compaction cuts Claude Code compaction from 51.8秒 to 0.26秒

3 Key Points

  1. What happened

    A developer released lossless-compaction, a Claude Code plugin that writes no model summary and instead moves large data to local storage. On a roughly 576,000-token conversation it finished compaction in 0.26 seconds versus 51.8 seconds for built-in /compact, and answered 11 of 11 follow-up questions correctly versus 6 of 11.

  2. Why it matters

    The built-in approach spends time and tokens having a model rewrite the conversation into a shorter summary, so anything the summary drops is gone from context. The plugin instead leaves a pointer to the original data, so even a wrong decision about what to move out does not lose information.

  3. What to watch

    The plugin keeps a much larger context than the built-in summary, so its economics hinge on prompt cache behavior and how much past information later work reuses. The developer still lists compression ratio as the open problem.

WHO IT HITSDevelopers and teams running long Claude Code sessions on large repositories are the direct audience, since they are the ones who currently wait tens of seconds for compaction and watch for auto-compaction timing. The cost math, though, depends on how well prompt caching holds up in their own workflow.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

The plugin's design is a direct reaction to how the built-in compaction works. Built-in compaction hands the whole conversation to a model and asks it to produce a summary, so time and tokens go into that rewrite, and any detail the summary omits is no longer in the next context. The developer had previously worked around this by manually running /compact at a convenient break before auto compaction hit, but that only shifts the timing and leaves the summarize-and-restart mechanism untouched.

The alternative approach explored, and then abandoned, was to have a model predict which information would matter later. The developer tried using Jev to judge whether a piece of information would be needed, but it did not perform as hoped, which the post attributes to the difficulty of predicting what will be needed 30 turns later from only the current context. Instead, lossless-compaction removes large items from context without deleting them, leaves a ticket with an ID, and lets the model recall the original content when needed. The content is identified by SHA-256 and read back from disk to confirm the write before removal.

The cost story is where the comparison gets less clean. The built-in approach pays at compaction time and then keeps a small context, but if dropped information becomes necessary again, those tokens re-enter the context from the session record. The plugin pays almost nothing at compaction time but carries a larger context into later requests, so when the prompt cache expires it has to rewrite that larger amount. The developer's own framing is that the question is not how to eliminate compaction cost but where and in what form to pay it. How that trade-off lands for a given team is likely to depend on how often its sessions pause long enough to break the cache, and how much of the earlier context later work actually reuses.

FAQ
How do I install lossless-compaction?
In a Claude Code session on version 2.1.287 or later, type: /plugin install lossless-compaction --marketplace yottayoshida/lossless-compaction. No configuration is required.
Does lossless-compaction lose information when it compacts?
No. It moves big items such as tool results out of the context but keeps a ticket with an ID, and the saved content is verified back from disk with SHA-256 before being removed. If saving fails, the information stays in context.
Is lossless-compaction cheaper than the built-in /compact?
In the first measurement it was cheaper in 5 of 6 conversation types, and more expensive in the one type that filled the context window. The developer says the total cost difference narrows once follow-up work is included, and it depends on prompt cache.

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleTSMC, $2.4 trillion chipmaker behind AI, reports $143 billion revenue