AIToday
Large Language ModelsOpen-Source AITechCrunch AIPublished: Sep 18, 2026, 10:00 JST

PrismML's Bonsai 2 27B hits 98% of Qwen at 5.9 GB

PrismML's Bonsai 2 27B hits 98% of Qwen at 5.9 GB

3 Key Points

  1. What happened

    PrismML released Bonsai 2 27B on Thursday, compressing Alibaba's Qwen3.8 27B to 5.9 GB — a 9x to 10x memory cut versus the original. It matches 98% of Qwen's aggregate benchmark scores, up from 95% for the first Bonsai in March.

  2. Why it matters

    Hassibi says PrismML's compression loses virtually no performance compared with the originals, and the gain from 95% to 98% across two releases suggests the technique is improving. That is the bet behind fitting capable reasoning models onto hardware people already own.

  3. What to watch

    Whether PrismML can ever reach 100% benchmark parity remains to be seen; Hassibi says compression will likely always have some impact. Watch for the next models, expected in the next couple of months, in the several-hundred-billion-parameter range.

WHO IT HITSAnyone running AI locally — developers and IT teams who want reasoning models on machines they already own rather than rented cloud servers — could eventually benefit, though the 5.9 GB size may still need a high-end phone.

Not sure about something? Ask the AI

Summaries like this, in your inbox every morning.

Context & Analysis

PrismML was founded by a group of Caltech researchers and is led by Babak Hassibi, a Caltech professor and compression expert. It has raised a $22.25 million seed round from Khosla Ventures, Cerberus Capital, and Caltech, and counts Ion Stoica, a Databricks co-founder and director of Berkeley's Sky Computing Lab, as an adviser.

This is not the only effort in LLM compression — Multiverse Computing, founded by a professor from Spain's Donostia International Physics Center, works on it too and has raised gobs of cash. Hassibi argues PrismML's tech stands out because its models have lost virtually no performance compared with the originals, and the benchmark match has improved from 95% to 98% across two releases. The original Bonsai has already been downloaded over 11 million times, and PrismML's even smaller models another 2.6 million times, the company says.

The next test is scale. Hassibi expects to release models in the several-hundred-billion-parameter range in the next couple of months, and predicts it will be easier to retain intelligence as models get bigger. Whether the compression holds up at that size — and whether it translates into everyday device use — is the open question.

FAQ
How small is Bonsai 2 27B?
It compresses Alibaba's Qwen3.8 27B from its original size down to 5.9 GB, a 9x to 10x reduction in memory. That is small enough to fit on a PC and possibly a high-end smartphone.
How much performance does it lose?
Bonsai 2 matches 98% of Qwen's aggregate benchmark scores, up from 95% for the first Bonsai released in March. PrismML CEO Babak Hassibi says compression will likely always have some impact.
How does PrismML's compression work?
It uses "ternary" weights, simplifying each of the model's stored weights from the usual 16 bits down to three values: +1, -1, or 0. With far smaller values to store, the model takes up dramatically less space.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • ByteDance's AI agent phone hits app wallDIGITIMES Asia · 1h ago
  • OpenAI launches Astra for Law, a GPT-6 setup for legal researchSiliconANGLE AI · 4h ago
  • Google opens CC to families of six as shared AI agentSiliconANGLE AI · 4h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleCooley builds GO Public on ChatGPT Work to speed IPO prep