Qwen 3.5 2B Runs On-Device on iPhone with MLX
Melvin Vivas · X video post · 2026-03-04 · Open on X
Topics: LLM Fundamentals, Industry Trends & Job Market · Level: beginner
Summary
Alibaba's Qwen 3.5 runs locally on an iPhone 17 Pro as a 2B model at 6-bit quantization, using MLX optimized for Apple Silicon. The original poster says it beats models four times its size, has strong visual understanding, and can toggle reasoning on or off. Melvin pitches it as a subscription-free alternative to ChatGPT.
Key points
- The Qwen 3.5 2B model at 6-bit quantization runs on an iPhone 17 Pro.
- It uses MLX optimized for Apple Silicon.
- It is reported to beat models four times its size.
- It has strong visual understanding.
- Reasoning can be toggled on or off.
- Local models avoid subscription costs.
Resources mentioned
- Qwen 3.5 · tool · qwen.ai · free
Alibaba Qwen open model family with small on-device variants and toggleable reasoning.
Also in: Workshop AI: Building Apps with Cloud and Local Agents (GLM 5, Qwen 3.5) (Melvin Vivas on X · notes), Qwen3.5 0.8B Does Real-Time Local Video Captioning (Melvin Vivas on X · notes), Qwen3.5 0.8B Real-Time Video Captioning on Mac Studio (Melvin Vivas on X · notes), Qwen 3.5 Is the Local Model That Holds Up in Claude Code (Melvin Vivas on X · notes) and 7 more - MLX · repo · github.com · free
Apple's open-source array and ML framework for efficient inference on Apple Silicon.
Also in: LM Studio MLX v1.8.1: Vision Model Batching and Better Caching (Melvin Vivas on X · notes), Gemma 4 Gets Up to 3x Faster with MTP Drafters (Melvin Vivas on X · notes), 1-bit Bonsai 8B Runs On-Device on iPhone at 40+ tok/s (Melvin Vivas on X · notes)
Try this
- Run a small quantized Qwen 3.5 model on-device with MLX.
More in LLM Fundamentals
- GLM-5-Turbo: A Fast GLM-5 Variant for Agents Like OpenClaw
- Real-Time Speech Transcription in the Browser with Voxtral and WebGPU
- NVIDIA build.nvidia.com: Try Hosted NIM Model APIs
- Qwen 3.5 One-Shots a Python Stock Analyzer App with yfinance
- GPT-5.3-Codex Now Available in OpenAI's Responses API
- Local Qwen3.5-35B-A3B on a 24GB GPU Builds a Full Game From One Spec