
What happened
Poolside released Laguna S 2.1, a compact open-weight coding model that scores 70.2 percent on Terminal-Bench 2.1 with thinking enabled—ranking just behind Tencent's Hy3 and ahead of much larger models like DeepSeek-V4-Pro-Max. The model was trained in fewer than nine weeks on 4,096 Nvidia H200 GPUs, using post-training across 409,000 environments.
Why it matters
Poolside treats persistence, verification, and iterative problem-solving as an alternative path to raw scale. The company demonstrated this by having Laguna S 2.1 independently rediscover Erdős Problem #397—a math problem open since 1975—in 40 steps with a reasoning cost of just $0.088. This suggests smaller models optimized for agentic behavior can match or exceed much larger systems on complex coding and reasoning tasks.
What to watch
Laguna S 2.1 is available now on Hugging Face under the OpenMDW 1.1 license, with free hosted access through OpenRouter (256K context window) and a paid tier supporting the full one-million-token window. A free demo chat runs at chat.poolside.ai. The model can also run locally on a single Nvidia DGX Spark.
Summaries like this, in your inbox every morning.
Poolside's approach to Laguna S 2.1 reflects a strategic shift away from the industry's default emphasis on model scale. Rather than pursuing raw parameter count or pre-training data volume, the company focused on post-training and agentic behavior—particularly persistence, verification, and iterative problem-solving. This framing is grounded in three documented trial runs: the model built a working browser engine in 50 minutes, and more remarkably, independently discovered a proof for Erdős Problem #397, a mathematics problem that had remained unsolved since 1975. The cost of that discovery was just $0.088 in reasoning tokens, a concrete demonstration of efficiency that extends beyond typical academic benchmarks.
The technical choices underscore this philosophy. Training across 409,000 agentic environments—including 83,000 for terminal tasks and 168,000 for software engineering workflows—was paired with innovations like multi-harness rollouts (which run the same prompts across several agent environments to reduce overfitting) and a new sandbox system that can selectively block network access to prevent reward hacking. The speed of training, completed in fewer than nine weeks on 4,096 Nvidia H200 GPUs starting May 22, 2026, suggests that agentic post-training can yield substantial performance gains without the extended timelines typical of large-scale pre-training efforts. At the same time, Poolside acknowledges limitations: the model remains closely tuned to its training harness in some cases and tends to produce overly long thinking sequences on competitive math problems—a sign that the current approach is neither a complete replacement for scale nor a finished product.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Anthropic published a blog post on September 8, 2026 (US time), outlining six common prompting anti-patterns t…

Thomas Ptacek published advice on writing with LLMs, stating "Rule Number One: You may not use a single word a…

PrismML released Bonsai 2 27B on Thursday, compressing Alibaba's Qwen3.8 27B to 5.9 GB — a 9x to 10x memory cu…

JPMorgan launched $2,000 monthly spending limits for some engineers using Claude Code, with internal error cod…

MCP now defines a standard way to discover and load Agent Skills from MCP servers, exposed as resources

The UN announced Thursday it is working with Google to build the UN System Data Commons on Google's open sourc…
