PrismML has released Bonsai 2 27B, a new model designed to run AI locally while requiring far less memory than its full-precision counterpart. With 27.8 billion parameters, the model supports reasoning, coding, image processing, tool use, and multi-step tasks.

In its ternary format, Bonsai 2 27B requires just 5.9 GB of memory, more than nine times less than the full-precision version. At the same time, it retains 98.2% of the baseline model’s aggregate performance across 20 benchmarks. By comparison, the previous Ternary Bonsai 27B retained 95% of the original model’s performance.

In testing, Bonsai 2 27B scored 83.9 points compared with 85.4 for Qwen3.8 27B, the model it is based on. The benchmark suite covered reasoning, mathematics, coding, instruction following, image understanding, and tool use. On an NVIDIA GeForce RTX 5090, Bonsai 2 27B reached speeds of up to 143 tokens per second.

Bonsai 2 27B was trained on Google v5 TPUs and optimized to run on consumer CPUs and edge GPUs. PrismML targets the model at local AI assistants and multimodal agents, private knowledge workflows, computer-use applications, and other long-running multi-step tasks. The ternary version is available to download for free under the Apache 2.0 license.