AIToday
Large Language ModelsTHE DECODERPublished: Sep 18, 2026, 01:00 JST

GPT-6 Astra beats Pokemon FireRed in 18 hours and 12 minutes

GPT-6 Astra beats Pokemon FireRed in 18 hours and 12 minutes

3 Key Points

  1. What happened

    According to operator Clad3815, GPT-6 Astra earned the champion title in Pokemon FireRed in 18 hours and 12 minutes, while GPT-5.6 Sol needed 96 hours and 35 minutes. In Minecraft, Vals AI called off its run after 141 hours.

  2. Why it matters

    Astra's jump over older models is stark, and the same trait that helps it work out unfamiliar mechanics from a few tries can also lead it to overcorrect and lose sight of its actual goal.

  3. What to watch

    The Minecraft potato farming shows the risk — a bad experience becomes a permanent rule that sticks even when it no longer helps, so the test is whether Astra can override such rules when the goal changes.

WHO IT HITSTeams evaluating AI agents for long, multi-step tasks — and researchers building benchmarks for them — now have a concrete case where a model's self-written rules both drove a record-fast completion and caused hours of wasted effort.

Not sure about something? Ask the AI

Summaries like this, in your inbox every morning.

Context & Analysis

The Pokemon and Minecraft runs come from community and company tests that sit alongside the ARC-AGI-3 benchmark, which drops models into unfamiliar, abstract game environments where they have to figure out the rules through their own actions. The pattern ARC Prize described is that Astra translates unfamiliar game mechanics into compact symbolic descriptions and tracks objects, coordinates, rules, and planned actions in a shorthand it develops itself — observations become rules, and rules become plans. In the Fallout 3 run, that approach showed up as cycles of pausing, observing, and acting; at a bed that was supposed to heal the character, it let the game run briefly and tried again; at a keypad, it waited until the game confirmed the button press before moving on. Looking back, Astra boiled these episodes down to a single principle: check whether an input worked before repeating it.

That same trait cuts both ways. The idea of turning experience into reusable rules isn't new — the 2023 research project Voyager had GPT-4 propose Minecraft tasks, write JavaScript code for them, and revise it based on error messages, but Voyager never saw the game and worked through the Mineflayer interface, while Vals AI's run uses general computer use with screen, mouse, and keyboard. In Minecraft, Astra's self-written rule about unguarded chests stuck even when it no longer helped, and it stayed in the potato field for several hours. The outcome likely hinges on whether rule-writing can be paired with the ability to abandon a rule once the situation changes — a question that matters most to teams deploying agents on long, open-ended tasks.

FAQ
How did GPT-6 Astra perform on ARC-AGI-3?
Astra hit 62.7 percent over the standard interface and about 99.9 percent with OpenAI's own harness. GPT-5.6 Sol scored 7.78 percent, and Claude Opus 5 a little over 30 percent.
What caused Astra to farm potatoes in Minecraft?
After a Creeper explosion wiped out its chest and bed, Astra wrote itself a rule to never store items in an unguarded chest again. It then spent several hours farming almost nothing but potatoes.
How long did Vals AI run its Minecraft test with Astra?
Vals AI called off the run after 141 hours. Before that, GPT-6 Astra had spent well over 100 hours in the game.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • OpenAI: Astra model slipped 27 prompt-injection summariesTHE DECODER · 1h ago
  • Claude Cowork folds into Claude ChatBen's Bites · 1h ago
  • Emerald AI alliance targets 100 gigawatts for data centersTechCrunch AI · 1h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleHBR flags 'Four patterns' on AI and psychological safety