
What happened
Developer cozyblaze reported that OpenAI's GPT-6 Astra played through the entire game Portal without human help after the initial goal was set, reaching the credits in about 23 hours and 43 minutes. The model controlled the game via MCP and a modified SourcePauseTool, pausing to decide its next moves.
Why it matters
This run cost at least $570 in tokens at Astra's list price, though cozyblaze used a $200 Codex subscription. Cozyblaze noted that OpenAI's 2016 goal was to solve many games with a single agent, and watching this solo playthrough offers a glimpse of that vision. He added that GPT-6 Astra is "the worst model we'll ever get."
What to watch
The test is whether such autonomous game-playing can extend to more complex, unmodified games. The 23-hour runtime and the token cost highlight the current limits of AI agents in real-world tasks.
Ask the AI about this article →
Summaries like this, in your inbox every morning.
The feat is a concrete step toward OpenAI's long-stated ambition of a single agent mastering many games, a vision the developer cozyblaze explicitly referenced. By finishing a full, unmodified puzzle game like Portal, the model demonstrates planning and decision-making that goes beyond narrow single-task benchmarks, even though it relied on a tool to pause the game while it reasoned.
At the same time, the cost and complexity underscore how far general game-playing AI is from being practical. The run burned at least $570 worth of tokens at list price, and required a custom pause tool to give the model time to think. That suggests that while autonomy is advancing, efficiency and the ability to act in real time remain hurdles. The developer's comment that GPT-6 Astra is "the worst model we'll ever get" is a reminder that this capability will likely look primitive compared with what comes next, and that progress is expected to continue.
The significance hinges on whether such autonomous play can be generalized to other, less structured tasks. If this approach scales and becomes cheaper, it could point toward AI agents that handle long-horizon goals in software environments with minimal human input. For now, the achievement is a proof of concept, with the 23-hour runtime and token cost serving as markers of how far there is to go.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Mastercard launched Agent Pay for Machines, a service that lets AI agents and software systems make high-speed…

ServiceNow (NOW) has been named a core partner in a multi-agent AI workflow initiative

Chugai Pharmaceutical has made Claude Code the core of its internal system development, after first adopting G…

ZEALS has started providing its multimodal AI agent platform "Omakase AI" to MUFG Bank

Mitsubishi Electric announced on June 25, 2026, at AWS Summit Japan that over 1,000 people are trialing an AI…

ChillStack and Mitsui Bussan Secure Direction will launch an e-learning version of their hands-on LLM security…
