
What happened
Good Start Labs trained a 30B model inside 1830: The Game of Railroads and Robber Barons, then tested it on financial research. Both training designs improved in-game, but only the multi-turn terminal agent improved on the Finance-Agent benchmark.
Why it matters
The result shows the training design, not just the game, decides which skills transfer to unfamiliar work. The earlier finding was that Diplomacy training improved customer support and industrial operations benchmarks; this experiment isolates the design variable.
What to watch
Whether broader real-world transfer holds beyond structurally similar tasks remains an open question, per CEO Alex Duffy. Watch whether the general game intelligence model built from its expert models can carry those skills further.
WHO IT HITSThis is directly relevant to teams at frontier AI labs that buy reinforcement learning data and learning environments, and to companies evaluating whether game-trained AI can handle tasks like financial research.
Ask the AI about this article →
Summaries like this, in your inbox every morning.
Good Start Labs grew out of a 2025 Twitch stream where Alex Duffy watched frontier models play Diplomacy and noticed each model behaved differently: OpenAI's o3 won by planning a future betrayal, while Claude's Opus 4 refused to lie and got destroyed. That observation led Duffy to the idea that games with verifiable outcomes could teach AI models skills like strategic thinking. He later wrote on Every that fine-tuning a model on Diplomacy improved its performance on customer support and industrial operations benchmarks. The company was spun out of Every last October with $3.6 million from General Catalyst, Inovia, Every, and angel investors.
The 1830 experiment is the latest step in that line of work. The game, a nineteenth-century railroad strategy title, includes a stock market mechanic and mirrors a finance workflow: models find information in a database, put it into an Excel file, reason over it, create functions, and calculate an answer. The published results compared single-turn question answering with a multi-turn terminal agent that uses tools to explore, plan, and adapt in real time. Both improved in-game, but only the terminal-agent design improved performance on the Finance-Agent benchmark. Meanwhile, Duffy says newer models still diverge on personality axes like betrayal and collaboration, and that for treating the environment as curriculum the harness matters more, not less.
The stake for Good Start Labs is whether it can sell more than custom data and environments — its main customers are frontier labs buying reinforcement learning data. The company is also building a general model from its game-specific expert models, aiming for what Duffy calls a general game intelligence. Whether that broader real-world transfer materializes is the open question Duffy himself acknowledges, and it is likely to determine how far this approach reaches beyond structurally similar tasks.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Keri Tracy, VP and chief audit executive at Newell Brands, told theCUBE that audit has a seat at the table for…
Profound raised $180 million in a Series D led jointly by Sequoia Capital and Kleiner Perkins, at a $1.8 billi…
AIUC raised a $40 million Series A led by Ribbit Capital to start auditing frontier AI models, after previousl…
Factory announced it raised $200 million, backed by Blackstone, Khosla Ventures, Sequoia Capital, NEA and othe…
Salesforce launched Koa, a reasoning model built on Nvidia's Nemotron open-weights platform, and Claudeforce…
Google launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, its most advanced voice models yet, wit…