Running Qwen3.8-27B Locally on an M5 Max MacBook with Inco Splash
Melvin Vivas · X video post · 2026-09-19 · 0:26 · 4,827 views · Open on X
Topics: LLMOps, Deployment & Monitoring, LLM Fundamentals · Level: intermediate
Summary
This short post reacts to a quoted announcement of Inco Splash, an open-source inference engine built for Qwen models on Apple silicon. The announcement says it runs Qwen3.8-27B at 144 tokens/s on an M5 Max MacBook Pro. It also says Inco Splash decodes faster than Ollama and oMLX, especially when an agent splits work across sub-agents. The point for learners is that the inference engine you choose can change local LLM speed a lot on the same hardware.
Key points
- Claimed result: Qwen3.8-27B runs at 144 tokens/s on an M5 Max MacBook Pro.
- Inco Splash is an open-source inference engine tuned for one model family (Qwen) and for Apple silicon.
- Claimed decode speed is up to 3x Ollama's and 2x oMLX's.
- The speedup is said to reach almost 4x when an agent fans out into sub-agents, so parallel agent workloads gain the most.
- Lesson: on the same hardware, the inference engine can matter as much as the model. Benchmark engines yourself before you commit to a local stack.
- The numbers are vendor claims from the quoted post, not independent benchmarks.
Resources mentioned
- Inco Splash · tool · github.com · free
Open-source LLM inference engine built for Qwen models on Apple silicon, aimed at fast local decoding. - Qwen3.8-27B · tool · huggingface.co · free
A 27B-parameter open-weight model from Alibaba's Qwen family that you can run locally or call through hosted APIs.
Also in: Deploying Qwen3.8 27B on Hugging Face Inference Endpoints (Melvin Vivas on X · notes), Using Hugging Face credits: Jobs, Inference Endpoints and Open Models (Melvin Vivas on X · notes), Ternary Bonsai 2 27B: 9x smaller model keeping 98.2% of benchmark scores (Melvin Vivas on X · notes), Free Qwen-3.8 27B Model via Infron (Melvin Vivas on X · notes) and 24 more - Ollama · tool · x.com · free · recommended by both Bashiri Smith & Melvin Vivas
Open-source tool for downloading and running LLMs on your own machine with minimal setup.
Also in: Running LLMs Locally Without an Expensive Rig (Melvin Vivas on X · notes), Adding Vercel AI Gateway as a provider in AIBackends with Devin (Melvin Vivas on X · notes), AIBackends: An API Layer Between Your App and AI Models (Now with Jev) (Melvin Vivas on X · notes), Coworker: An Open-Source Desktop Agent App for Local and Cloud Models (Melvin Vivas on X · notes) and 12 more - oMLX · tool · github.com · free
An MLX-based tool for running local models on Apple silicon.
Also in: Claude support for Apple's Foundation Models framework (Melvin Vivas on X · notes) - M5 Max MacBook Pro · other · apple.com · paid
Apple laptop with M5 Max chip, used as the hardware for the local inference benchmark.
More in LLMOps, Deployment & Monitoring
- Adding Vercel AI Gateway as a provider in AIBackends with Devin
- Jev was free on Vercel AI Gateway until Sept 25
- Benchmarking LLM Endpoints with NVIDIA Dynamo AIPerf
- Hugging Face Cache Deduplication with Xet in huggingface_hub v1.32
- Agent Monitor: see traces, tokens and costs of your coding agents
- Agent Monitor: Per-Model Usage Stats for Codex and Claude Code Agents