AIToday
AI Coding AssistantsOpen-Source AITHE DECODERPublished: Jul 23, 2026, 22:01 JST

Poolside's Laguna S 2.1 rivals larger open models on coding tasks

Poolside's Laguna S 2.1 rivals larger open models on coding tasks

3 Key Points

  1. What happened

    Poolside released Laguna S 2.1, a compact open-weight coding model that scores 70.2 percent on Terminal-Bench 2.1 with thinking enabled—ranking just behind Tencent's Hy3 and ahead of much larger models like DeepSeek-V4-Pro-Max. The model was trained in fewer than nine weeks on 4,096 Nvidia H200 GPUs, using post-training across 409,000 environments.

  2. Why it matters

    Poolside treats persistence, verification, and iterative problem-solving as an alternative path to raw scale. The company demonstrated this by having Laguna S 2.1 independently rediscover Erdős Problem #397—a math problem open since 1975—in 40 steps with a reasoning cost of just $0.088. This suggests smaller models optimized for agentic behavior can match or exceed much larger systems on complex coding and reasoning tasks.

  3. What to watch

    Laguna S 2.1 is available now on Hugging Face under the OpenMDW 1.1 license, with free hosted access through OpenRouter (256K context window) and a paid tier supporting the full one-million-token window. A free demo chat runs at chat.poolside.ai. The model can also run locally on a single Nvidia DGX Spark.

Not sure about something? Ask the AI

Summaries like this, in your inbox every morning.

Context & Analysis

Poolside's approach to Laguna S 2.1 reflects a strategic shift away from the industry's default emphasis on model scale. Rather than pursuing raw parameter count or pre-training data volume, the company focused on post-training and agentic behavior—particularly persistence, verification, and iterative problem-solving. This framing is grounded in three documented trial runs: the model built a working browser engine in 50 minutes, and more remarkably, independently discovered a proof for Erdős Problem #397, a mathematics problem that had remained unsolved since 1975. The cost of that discovery was just $0.088 in reasoning tokens, a concrete demonstration of efficiency that extends beyond typical academic benchmarks.

The technical choices underscore this philosophy. Training across 409,000 agentic environments—including 83,000 for terminal tasks and 168,000 for software engineering workflows—was paired with innovations like multi-harness rollouts (which run the same prompts across several agent environments to reduce overfitting) and a new sandbox system that can selectively block network access to prevent reward hacking. The speed of training, completed in fewer than nine weeks on 4,096 Nvidia H200 GPUs starting May 22, 2026, suggests that agentic post-training can yield substantial performance gains without the extended timelines typical of large-scale pre-training efforts. At the same time, Poolside acknowledges limitations: the model remains closely tuned to its training harness in some cases and tends to produce overly long thinking sequences on competitive math problems—a sign that the current approach is neither a complete replacement for scale nor a finished product.

FAQ
How does Laguna S 2.1 compare to other open models on benchmarks?
With thinking enabled, Laguna S 2.1 scores 70.2 percent on Terminal-Bench 2.1, ranking just behind Tencent's Hy3 and ahead of much larger open models including DeepSeek-V4-Pro-Max and Nemotron 3 Ultra. On DeepSWE, it scores 40.4 percent while some open-weight models with more than one trillion parameters remain below 10 percent.
What is the impact of thinking mode on the model's performance?
Thinking mode has a major impact: without it, Laguna S 2.1's Terminal-Bench score drops to 60.4 percent and its DeepSWE score falls to 16.5 percent. Poolside says no previous Laguna model has shown a larger performance gap between the two modes.
How can I access Laguna S 2.1?
Laguna S 2.1 is available on Hugging Face under the OpenMDW 1.1 license. Hosted access is provided by Baseten, Vercel AI Gateway, and OpenRouter; OpenRouter offers a free endpoint with a 256K context window and a paid endpoint supporting the full one-million-token window. The model can also run locally on a single Nvidia DGX Spark, and a free demo chat is available at chat.poolside.ai.

Get the latest AI Coding Assistants news every morning

For example, today's edition would include:

  • Anthropic says 'just-in-case' prompts waste Claude tokensITmedia AI+ · 4h ago
  • Thomas Ptacek: Never use a word an LLM suggestsSimon Willison's Weblog · 4h ago
  • JPMorgan caps Claude at $2,000 a month for some engineersTop Companies AI · 7h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleCursor launches CFO Council as $60B SpaceX deal nears close