
Poolside released Laguna S 2.1, a 118-billion-parameter coding model that activates only 8 billion parameters per token and scores 70.2% on Terminal-Bench 2.1, placing it 11th on Poolside's leaderboard. Despite being smaller than many rivals, the model matches or beats open models several times its size on agentic coding tasks, demonstrating that a smaller lab can compete at the frontier through transparency and efficiency rather than scale alone. The model is available immediately on Hugging Face.
Summaries like this, in your inbox every morning.
Sign up free →What happened
Poolside, a San Francisco AI lab, released Laguna S 2.1 on Tuesday—a 118-billion-parameter Mixture-of-Experts coding model that activates only 8 billion parameters per token and supports a context window of up to 1 million tokens. The model scores 70.2% on Terminal-Bench 2.1, placing it 11th on Poolside's leaderboard. Weights are available immediately on Hugging Face under the permissive OpenMDW-1.1 license.
Why it matters
The company is making a deliberate bet that radical transparency rather than raw scale is how a smaller lab competes at the frontier. Laguna S 2.1 matches or beats open models several times its size on agentic coding tasks—meaning Poolside can deliver competitive performance without building the largest model. This model builds on the lab's three-year track record quietly selling coding models to governments and defense agencies.
What to watch
The weights are available immediately on Hugging Face, making the model accessible to developers and researchers who want to build on or evaluate Poolside's approach to efficient, transparent model development.
Poolside, a three-year-old AI lab based in San Francisco, has carved out a niche selling coding models to governments and defense agencies—work that has kept the company largely out of the public eye. On Tuesday, it released Laguna S 2.1, marking a significant shift in strategy by making its most capable model publicly available on Hugging Face.
Laguna S 2.1 is a 118-billion-parameter Mixture-of-Experts (MoE) system—a type of neural network where only a fraction of the model's parameters activate for any given input, reducing computational cost. The model activates only 8 billion parameters per token and supports a context window of up to 1 million tokens, meaning it can process very long coding problems or sequences of instructions. According to Poolside's benchmarks, the model scores 70.2% on Terminal-Bench 2.1, a benchmark designed to measure performance on long-horizon terminal tasks (extended sequences of commands or operations a coder might need to perform). On Poolside's compiled leaderboard, Laguna S 2.1 ranks 11th and is positioned ahead of larger models, demonstrating that it matches or beats open models several times its size on agentic coding tasks.
Poolside's release strategy underscores a deliberate bet on transparency and efficiency over raw scale. By publishing the weights immediately on Hugging Face under the permissive OpenMDW-1.1 license, the company is inviting external scrutiny and use, a notable contrast to the growing trend of closed models from larger competitors. For a smaller lab without the capital to build and train models an order of magnitude larger, this approach suggests an alternative path to competitiveness: focus on efficiency, publish results clearly, and let the community validate the claims. Whether other researchers and practitioners confirm Laguna S 2.1's performance gains on real-world tasks will determine whether Poolside's bet on transparency and efficiency reshapes expectations for competitive coding models.
Poolside has built its business over three years by operating outside the spotlight, selling coding models to governments and defense agencies rather than competing in the headline-grabbing race for scale. With Laguna S 2.1, the company is making a strategic pivot: publishing its most capable model publicly on Hugging Face and claiming that a 118-billion-parameter system can outperform much larger rivals on coding benchmarks. The key to this claim is efficiency—the model activates only 8 billion of its 118 billion parameters per token, which means inference is far cheaper than running a fully-dense system of comparable stated size.
The company's decision to emphasize "radical transparency" over scale is notable in an industry increasingly dominated by closed, massive models from well-funded incumbents. By releasing under a permissive open license immediately, Poolside signals confidence in the model's design and invites external validation. The 70.2% score on Terminal-Bench 2.1 and 11th-place ranking on its own leaderboard are concrete benchmarks, though Poolside's proprietary leaderboard may not yet carry the industry weight of independent benchmarks. What matters for the field is whether external researchers and developers confirm that this small, efficient model can indeed match larger alternatives on real-world coding tasks.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
No comments yet. Be the first to share your thoughts!
Log in to join the discussion





Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.
Get Started FreeFree · takes 30 seconds · unsubscribe anytime
1 minute a day. The AI essentials.
200+ sources · Email / LINE / Slack