AIToday
Large Language ModelsOpen-Source AISiliconANGLE AIPublished: Sep 19, 2026, 04:00 JST

Prism ML shrinks Qwen3.8 into 5.9GB Bonsai 2 27B

Prism ML shrinks Qwen3.8 into 5.9GB Bonsai 2 27B

3 Key Points

  1. What happened

    Prism ML Inc. launched Bonsai 2 27B, which compresses a Qwen3.8 27B-based model from about 56 gigabytes to about 5.9 gigabytes while keeping around 98.2% of its capabilities.

  2. Why it matters

    The model claims to hold onto most of its original abilities at a fraction of the size. That could let users run capable AI locally instead of sending data to a cloud service.

  3. What to watch

    The gains hinge on whether benchmark parity holds in real-world tasks, not just test scores. Watch the agentic and tool calling gap of three points against the original model.

WHO IT HITSEnterprise IT teams evaluating on-device AI for privacy-sensitive tasks, and developers building local apps, may find a smaller option that avoids sending data to third-party cloud services.

Not sure about something? Ask the AI

Summaries like this, in your inbox every morning.

Context & Analysis

Prism ML Inc.'s Bonsai 2 27B arrives as a second-generation attempt to squeeze a capable model onto consumer hardware. The company used ternary compression, simplifying the model's weights from 16 bits down to three bits represented by +1, 0, and -1, which lets it store information in a much smaller memory footprint. The article notes that other compression methods called quantization usually strip away accuracy, knowledge and other capabilities, and that Qwen3.8's minimal memory footprint is 9.4 gigabytes.

On benchmarks, Bonsai 2 stayed close to Qwen3.8 within three points on agentic and tool calling, at 77.6 and 79.8 respectively. It scored 81.6 and 82.2 for coding and 82.7 and 81.3 for knowledge and reasoning. The company said the model consumes extremely low power per token at 0.714 megawatt-hours, making it 40% more energy-efficient than other 8B models running at full precision.

The appeal of ultra-small models is that they let users run AI locally without sending inference to the cloud, which can introduce delays or expose sensitive information to third parties. Whether Bonsai 2 27B delivers on that promise for everyday users and enterprises may depend on how well its benchmark parity translates to real tasks, and on whether the local hardware it targets can handle the workloads users actually care about.

FAQ
How much smaller is Bonsai 2 27B than the original model?
Prism ML Inc. reduced a Qwen3.8 27B-based model from about 56 gigabytes at full 16-bit size to about 5.9 gigabytes, while keeping around 98.2% of its capabilities.
Can Bonsai 2 27B run without special hardware?
It runs on Nvidia GPUs via CUDA and on Apple devices including Mac, iPhone and iPad via MLX. It can run on an Nvidia GeForce GTX 5090 without quantization, reaching 143 tokens per second.
What license are the model weights under?
The model weights are available today under Apache 2.0 licenses.
SiliconANGLE AIRead Original Article

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • CoreWeave's first user conference set for Sept. 30-Oct. 1SiliconANGLE AI · 1h ago
  • HarnessRouter standardizes agent runs via one protocolDaily Dose of Data Science · 1h ago
  • Gartner: over 40% of agentic AI projects to be canned by 2027AINOW · 1h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleKing Charles III to Nvidia, DeepMind: keep AI under human control