1-bit Bonsai 8B Runs On-Device on iPhone at 40+ tok/s
Melvin Vivas · X video post · 2026-04-01 · Open on X
Topics: LLM Fundamentals, Industry Trends & Job Market · Level: intermediate
Summary
A demo shows PrismML's 1-bit Bonsai 8B running locally on an iPhone 17 Pro at more than 40 tokens per second, which the original poster calls a first for a dense 8B model on iPhone. It runs on Apple MLX and is available in the Locally AI app.
Key points
- Bonsai 8B is a 1-bit quantized dense 8B model from PrismML.
- It runs at over 40 tokens/s on an iPhone 17 Pro.
- Inference uses Apple MLX, which is optimized for Apple Silicon.
- You can try it in the Locally AI app.
- Extreme quantization makes 8B-class models practical on phones.
Resources mentioned
- Bonsai 8B (PrismML) · tool · prismml.com · free
1-bit quantized dense 8B language model built for on-device inference. - MLX · repo · github.com · free
Apple's open-source array and ML framework for efficient inference on Apple Silicon.
Also in: LM Studio MLX v1.8.1: Vision Model Batching and Better Caching (Melvin Vivas on X · notes), Gemma 4 Gets Up to 3x Faster with MTP Drafters (Melvin Vivas on X · notes), Qwen 3.5 2B Runs On-Device on iPhone with MLX (Melvin Vivas on X · notes) - Locally AI · tool · locallyai.app · free
An iPhone app for downloading open models and running them on the device, offline.
Also in: Run Gemma 4 Offline on an iPhone with the Locally AI App (Melvin Vivas on X · notes)
Try this
- Try Bonsai 8B on-device using the Locally AI app.
More in LLM Fundamentals
- Gemma 4 26B for OCR in LM Studio
- Gemma 4: Google's Apache 2.0 Open-Weight Models for Local Hardware
- Qwen 3.6 Plus Preview Is Free for a Limited Time on OpenRouter
- MiniMax 2.7 One-Shots a Linear Clone at 95% Lower Cost
- Qwen3.5 0.8B Does Real-Time Local Video Captioning
- Qwen3.5 0.8B Real-Time Video Captioning on Mac Studio