PrismML Releases Bonsai Image 4B for Local Image Generation

May 27, 2026
PrismML has introduced Bonsai Image 4B, a compressed diffusion model family that enables high-quality image generation directly on local devices such as iPhones and Macs. The models are available under the Apache 2.0 license with open weights and code.

PrismML has introduced Bonsai Image 4B, a family of compressed image generation models designed to run on local devices such as laptops and phones, announced in a press release. The models make high quality diffusion inference practical on hardware ranging from iPhones to Apple Silicon Macs.

Bonsai Image 4B is available in two variants: a 1-bit model and a ternary model. The 1-bit version uses binary transformer weights with group-wise FP16 scaling to achieve maximum compression, shrinking a 4 billion parameter diffusion transformer to 0.93 GB. The ternary variant, which adds a zero state for more representational flexibility, compresses to 1.21 GB. Both retain up to 95 percent of the quality of the full precision model.

On an iPhone 17 Pro Max, Bonsai Image 4B generates a 512 by 512 image in about 9.4 seconds. On a Mac M4 Pro, the same resolution takes about 6 seconds, up to 5.6 times faster than the full precision pipeline. The models are built for local inference across iPhone, Apple Silicon Macs, CUDA GPUs, and small scale serving environments.

Both Bonsai Image 4B variants are released with open weights and code under the Apache 2.0 license. PrismML is also launching Bonsai Studio, an iOS app that allows users to try Bonsai Image 4B directly on Apple Inc. devices.

We hope you enjoyed this article.

Consider subscribing to one of our newsletters like Daily AI Brief.

Also, consider following us on social media:

Subscribe to Daily AI Brief

Daily report covering major AI developments and industry news, with both top stories and complete market updates

Whitepaper

Tensordyne Napier: What If One Rack Could Do the Work of Nine?

Tensordyne

This Tensordyne whitepaper presents Napier, an inference-focused AI processor and rack-scale system based on the company’s TDN Math logarithmic number system. It examines infrastructure requirements for large mixture-of-experts and agentic models, compares major inference architecture approaches, and details the TDN AIP processor, TDN72 pod, TDN Link fabric, and Napier Ultra configuration. The paper reports simulation-based performance, cost, and accuracy-validation results, including Tensordyne’s projected comparison of one Napier rack with a nine-rack Nvidia Rubin plus Groq deployment; the chip is reported as taped out and in fabrication.

Read more