
What happened
Prism ML Inc. launched Bonsai 2 27B, which compresses a Qwen3.8 27B-based model from about 56 gigabytes to about 5.9 gigabytes while keeping around 98.2% of its capabilities.
Why it matters
The model claims to hold onto most of its original abilities at a fraction of the size. That could let users run capable AI locally instead of sending data to a cloud service.
What to watch
The gains hinge on whether benchmark parity holds in real-world tasks, not just test scores. Watch the agentic and tool calling gap of three points against the original model.
WHO IT HITSEnterprise IT teams evaluating on-device AI for privacy-sensitive tasks, and developers building local apps, may find a smaller option that avoids sending data to third-party cloud services.
Summaries like this, in your inbox every morning.
Prism ML Inc.'s Bonsai 2 27B arrives as a second-generation attempt to squeeze a capable model onto consumer hardware. The company used ternary compression, simplifying the model's weights from 16 bits down to three bits represented by +1, 0, and -1, which lets it store information in a much smaller memory footprint. The article notes that other compression methods called quantization usually strip away accuracy, knowledge and other capabilities, and that Qwen3.8's minimal memory footprint is 9.4 gigabytes.
On benchmarks, Bonsai 2 stayed close to Qwen3.8 within three points on agentic and tool calling, at 77.6 and 79.8 respectively. It scored 81.6 and 82.2 for coding and 82.7 and 81.3 for knowledge and reasoning. The company said the model consumes extremely low power per token at 0.714 megawatt-hours, making it 40% more energy-efficient than other 8B models running at full precision.
The appeal of ultra-small models is that they let users run AI locally without sending inference to the cloud, which can introduce delays or expose sensitive information to third parties. Whether Bonsai 2 27B delivers on that promise for everyday users and enterprises may depend on how well its benchmark parity translates to real tasks, and on whether the local hardware it targets can handle the workloads users actually care about.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
CoreWeave holds its inaugural user conference in San Francisco on Sept
HarnessRouter, an open-source implementation of the Unified Harness Protocol, runs Codex and Claude Code throu…

Gartner predicts over 40% of agentic AI projects will be canceled by end of 2027, citing rising costs, unclear…

Disney appointed Karandeep Anand, former CEO of Character.AI, as its first-ever chief technology officer; Vari…

Google introduced an updated CC AI agent that shifts from a work productivity tool to household coordination…

Moonshot AI's Kimi K3 is now on Amazon Bedrock, with a 1-million-token context window, native vision, and an a…
