
What happened
On August 7, 2026, a developer tested GPT-5.6 Sol Ultra (running in Codex Desktop with aggressive sub-agent use) on a raccoon heist game premise originally generated with GPT-3 four years ago. The resulting game significantly improved on Claude Fable 5's version—moving from a single raccoon collecting coins in a backyard to a museum heist where three raccoons stack to steal a golden sardine.
Why it matters
The test shows GPT-5.6 Sol Ultra's ability to produce more complex, feature-rich game output from the same prompt. Though the first version had a bug (oversized eyeballs on each raccoon), the AI failed to spot it during development—requiring manual fixes via follow-up prompts. This suggests that while reasoning power improved, systematic review and independent problem detection remain gaps.
What to watch
Codex spent 52 minutes on the full project; the estimated cost at full API pricing (versus a monthly subscription used here) is available in the AgentsView cost estimate shared in the GitHub repository. The complete transcript and game code, including textures and image-generation prompts, are available on GitHub.
Summaries like this, in your inbox every morning.
The test represents a direct head-to-head comparison of two AI systems on an identical creative task: building a playable game from a single natural-language prompt. The original 2022 premise—generated by GPT-3—was deliberately reused four years later to measure progress. GPT-5.6 Sol's output demonstrated ambition and structural improvement over Claude Fable 5, preserving the heist concept and adding the emergent mechanic of stacking characters to achieve a goal, whereas Claude's version simplified the premise into a single-character coin-collection game. However, the raccoon eyeball bug reveals a practical limitation: despite the AI reviewing screenshots during development, it failed to detect and self-correct an obvious visual glitch. The developer's manual fix—which required explicit prompts naming the problem—suggests that even aggressive sub-agent use in Codex does not automatically confer visual verification or self-critique capabilities. The cost estimate and 52-minute session length provide benchmarks for computational load, though the true API expense is masked by the developer's monthly subscription model.
For example, today's edition would include:
AI-summarized, only the topics you pick: one digest a day via Email, LINE, or Slack.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Instinct, officially Spear Street Technology Inc., announced a $1 billion Series C joined by Sequoia Capital…
At Okta's Oktane event, Charlotte Wylie, Okta's senior vice president and deputy chief security officer, said…
On theCUBE Pod, Dave Vellante said CoreWeave disclosed that 70% of its revenue came from its top three custome…
Meta is launching the Meta Enterprise Platform, a new business unit selling the Muse agent, Meta Business Agen…

Nvidia combined OpenShell, its March open-source sandbox software, with Sentry, a hardware watchdog for its Bl…

In an open letter on Quillette to Scott Alexander, Steven Pinker declined a public debate on AI's existential…
