
What happened
PrismML released Bonsai 2 27B on Thursday, compressing Alibaba's Qwen3.8 27B to 5.9 GB — a 9x to 10x memory cut versus the original. It matches 98% of Qwen's aggregate benchmark scores, up from 95% for the first Bonsai in March.
Why it matters
Hassibi says PrismML's compression loses virtually no performance compared with the originals, and the gain from 95% to 98% across two releases suggests the technique is improving. That is the bet behind fitting capable reasoning models onto hardware people already own.
What to watch
Whether PrismML can ever reach 100% benchmark parity remains to be seen; Hassibi says compression will likely always have some impact. Watch for the next models, expected in the next couple of months, in the several-hundred-billion-parameter range.
WHO IT HITSAnyone running AI locally — developers and IT teams who want reasoning models on machines they already own rather than rented cloud servers — could eventually benefit, though the 5.9 GB size may still need a high-end phone.
Summaries like this, in your inbox every morning.
PrismML was founded by a group of Caltech researchers and is led by Babak Hassibi, a Caltech professor and compression expert. It has raised a $22.25 million seed round from Khosla Ventures, Cerberus Capital, and Caltech, and counts Ion Stoica, a Databricks co-founder and director of Berkeley's Sky Computing Lab, as an adviser.
This is not the only effort in LLM compression — Multiverse Computing, founded by a professor from Spain's Donostia International Physics Center, works on it too and has raised gobs of cash. Hassibi argues PrismML's tech stands out because its models have lost virtually no performance compared with the originals, and the benchmark match has improved from 95% to 98% across two releases. The original Bonsai has already been downloaded over 11 million times, and PrismML's even smaller models another 2.6 million times, the company says.
The next test is scale. Hassibi expects to release models in the several-hundred-billion-parameter range in the next couple of months, and predicts it will be easier to retain intelligence as models get bigger. Whether the compression holds up at that size — and whether it translates into everyday device use — is the open question.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
ByteDance's second AI-agent phone replaces forced automation with a permission-based approach, but the first m…

OpenAI launched Astra for Law, wrapping GPT-6 Astra in a legal search index covering US case law, statutes, re…
Anthropic detailed three metrics — AI-led R&D, oversight of autonomous AI agents, and compute allocation — dis…
Google Labs opened its experimental AI agent CC to households of up to six people
Shiseido Japan's AI agent for ingredient discovery cut search time by 95% and increased proposed ingredient ca…

Anthropic published a blog post on September 8, 2026 (US time), outlining six common prompting anti-patterns t…
