
What happened
A developer moved Codex work to Pi Coding Agent, running gpt-6-sol at high thinking. Implementation-phase use of the 5-hour limit fell from 15–20% to a stable 10% per change, and the review phase moved to Opus 5.5.
Why it matters
The same coding work now consumes roughly half the limit per change, so sessions stretch further before the cap bites. That is what let the developer keep shipping changes rather than stalling mid-phase.
What to watch
Review still consumes around 10%, so one change costs about 20% of the 5-hour limit in total — better, but far from continuous development. Whether review can be trimmed without moving it to a different model is the open question.
WHO IT HITSDevelopers running coding agents against fixed usage caps are the audience here — they can cut limit burn by swapping the harness rather than the model. Teams relying on Codex for implementation work in particular may find a lighter harness stretches their quota further.
Summaries like this, in your inbox every morning.
The switch did not start as a search for a better tool. It started as a quota problem. Codex with gpt-5.6-sol at high thinking once used about 10% of the 5-hour limit for a single implementation pass; that figure later rose to 15–20% even at medium, and gpt-6-sol occasionally spiked to roughly 40% during a single task. Review, once 5–10%, settled above 10% as well.
The developer suspected the harness itself was part of the cause — that the large system prompts Claude Code and Codex add to advertise their many features were feeding the model irrelevant information and burning tokens. Pi Coding Agent was adopted with subagents deliberately left out, since they were one of the suspects. A short appended system prompt and a minimal web extension replaced the heavier defaults.
The result was less about raw performance than about predictability: gpt-6-sol at high thinking now holds near 10% for implementation, and the developer reports no felt drop in quality. Review still costs about 10%, so the review step was handed to Opus 5.5, whose low limit consumption made that possible. Whether this holds once the workflow grows more complex — more subagents, longer tasks — is the part still unproven.
For example, today's edition would include:
AI-summarized, only the topics you pick: one digest a day via Email, LINE, or Slack.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
AWS made Amazon CloudWatch Omni generally available last week, with 17 built-in evaluators that score coherenc…
Google said Gemini Omni 1.1 Flash is now directly available in Google Vids, letting users extend scenes while…

In a September 2025 paper titled "Why Language Models Hallucinate," OpenAI researchers said low-frequency fact…

The developer moved from IDE-centric work in WebStorm and PHPStorm to terminal-centric work with Claude Code…

In a preliminary code-review trial (internal log EXP-001-R2), Claude and Codex each independently found one va…

While building an accounting app called Books tied to マネーフォワード, a developer built Jev Bookmarks
