Bonsai 27B
Open-source 27B multimodal model with 1-bit and ternary variants that run locally on a phone or laptop for on-device agentic workflows.
Updated 2026-07-15
Yes. Bonsai 27B is free to use.
Listed pricing: Free — open-source weights under Apache 2.0.
Overview
Bonsai 27B is an open-source, 27-billion-parameter multimodal model family from PrismML, released July 14, 2026, that ships in extreme low-bit variants — including ternary and 1-bit weights — designed to run natively on phones and laptops rather than in a datacenter. The pitch is that you get a model in the 27B class, historically a server-side weight class, small enough to load and run on an iPhone while keeping performance close to a full-precision version.
It's aimed at developers and researchers building on-device and agentic workflows: local assistants, offline multimodal apps, and privacy-sensitive pipelines where sending data to a hosted API is a non-starter. Because the weights are Apache 2.0, teams can fine-tune, quantize further, and ship Bonsai inside a commercial product without licensing friction — the same freedom that made Llama-class and Mistral open weights popular, but at a footprint that targets edge hardware instead of a GPU box.
What sets Bonsai apart is the bet on ternary and 1-bit quantization as a first-class release rather than an afterthought. Most open models ship at full precision and leave community members to produce aggressive quants that lose quality; PrismML is claiming the low-bit variants are the point, positioning Bonsai as the first major 27B-class model engineered to actually fit and run on a phone. Whether the near-full-precision claim holds up under independent benchmarking is the open question — verify the specific quality-vs-size tradeoff against your own eval before betting a product on it.
Is Bonsai 27B free?
Yes. Bonsai 27B is free to use.
What the free tier covers: Entirely free and open-source — weights released under Apache 2.0 with no paid tier from PrismML.
Listed pricing: Free — open-source weights under Apache 2.0.
Pricing on this page has not been re-verified. The entry was last edited on , and no separate pricing check has been run since. Treat the figures as a record of what was published then and confirm on the official site.
Bonsai 27B pricing
| Plan | Price | What's included |
|---|---|---|
| Open Source | Free | Full model weights (ternary, 1-bit, and higher-precision variants) under Apache 2.0. Self-host on-device or on your own hardware; no usage caps or license fees. You supply the compute. |
Full model weights (ternary, 1-bit, and higher-precision variants) under Apache 2.0. Self-host on-device or on your own hardware; no usage caps or license fees. You supply the compute.
Pricing on this page has not been re-verified. The entry was last edited on , and no separate pricing check has been run since. Treat the figures as a record of what was published then and confirm on the official site.
Is Bonsai 27B worth it?
Worth it for Developers and researchers building offline or privacy-sensitive agentic apps who need a multimodal model small enough to run locally on a phone or laptop.
You can test that on the free tier before paying anything. The recorded trade-offs are listed below, and any one of them can settle the question on its own.
The 8.5/10 AI Score is an editorial read of published capability, price and shipping pace. Nobody here has hands-on hours with Bonsai 27B. How we verify.
Worth it if
The strengths recorded against this entry.
- 27B-class model small enough to run on a phone via ternary/1-bit weights — a genuinely new footprint for this size class
- Apache 2.0 license permits commercial use, fine-tuning, and redistribution with no strings attached
- Fully local inference keeps data on-device, a real advantage for privacy-sensitive and offline apps
- Multimodal rather than text-only, broadening what on-device agentic workflows can do
- Free — you only pay for your own compute
Not worth it if
Any one of these blocks your use case.
- 'Near full-precision' claims at 1-bit/ternary are the vendor's own and need independent benchmarking before you trust them
- Running a 27B model on a phone still taxes RAM, thermals, and battery — real-world throughput on older devices is unproven
- Self-hosting means you handle deployment, quantization tooling, and updates; there's no managed API to fall back on
- Brand-new (launched July 14, 2026), so tooling, community quants, and long-term support are still immature
What sets Bonsai 27B apart
- Ternary and 1-bit weight variants, an unusual first-class release for a 27B-class model
- Runs natively on-device on phones and laptops instead of requiring server GPUs
- Apache 2.0 license permits commercial use, fine-tuning, and redistribution with no restrictions
- Multimodal rather than text-only at a size class previously reserved for the cloud
Key features
Low-bit variants
Ships in ternary and 1-bit weight formats alongside higher-precision options, so the same 27B model can be run at a fraction of the usual memory footprint. This is what lets it target phone and laptop RAM budgets instead of server GPUs.
On-device inference
Engineered to run locally on consumer hardware including iPhones, keeping data on the device. That matters for offline use, latency, and privacy-sensitive applications where a hosted API isn't acceptable.
Multimodal 27B
A 27-billion-parameter multimodal family rather than a text-only model, so it can handle mixed inputs within a size class previously reserved for cloud deployment.
Apache 2.0 open weights
The weights are released under a permissive Apache 2.0 license, allowing free commercial use, fine-tuning, and redistribution without the usage restrictions attached to many 'open' model licenses.
How it compares
| Tool | Best for | Pricing | Score |
|---|---|---|---|
| Bonsai 27B | Developers and researchers building offline or privacy-sensitive agentic apps who need a multimodal model small enough to run locally on a phone or laptop. | Free — open-source weights under Apache 2.0 | 8.5/10 |
| Perplexity AI vs Perplexity AI → | Knowledge workers and researchers who want cited, synthesized answers instead of a list of links to click through. | Freemium | 9.4/10 |
| Gemini 3.8 Flash | — | $0.75 per million input tokens, $3.75 per million output tokens | 9.2/10 |
| NotebookLM vs NotebookLM → | Students, researchers, and professionals who need answers grounded strictly in the specific documents they upload. | Free | 9.1/10 |
Compare head-to-head
Related reading
DeepSeek V4.1-Flash Ships With New Architecture
DeepSeek added V4.1-Flash to its API on the date recorded in the September 10 changelog entry with 552B MoE parameters, native vision, and lower API rates.
Mistral €3bn Round Meets OpenAI Wiki Disclosure
Mistral closed a €3bn round above €21bn valuation while OpenAI outlined next steps on agent misalignment reporting from its recent incident.
Mistral Closes €3B Series D at €21B Valuation
Mistral raised €3 billion in Europe's largest tech equity round, led by Samsung at a post-money valuation above €21 billion, to expand sovereign open-weight
Ready to try Bonsai 27B?
Head to the official site to start with Bonsai 27B — pricing and plans are listed above.
Visit Bonsai 27B

