AI Engineer Study Library

Running Qwen3.8-27B Locally on an M5 Max MacBook with Inco Splash

Melvin Vivas · X video post · 2026-09-19 · 0:26 · 4,827 views · Open on X

Topics: LLMOps, Deployment & Monitoring, LLM Fundamentals · Level: intermediate

Summary

This short post reacts to a quoted announcement of Inco Splash, an open-source inference engine built for Qwen models on Apple silicon. The announcement says it runs Qwen3.8-27B at 144 tokens/s on an M5 Max MacBook Pro. It also says Inco Splash decodes faster than Ollama and oMLX, especially when an agent splits work across sub-agents. The point for learners is that the inference engine you choose can change local LLM speed a lot on the same hardware.

Key points

Resources mentioned

More in LLMOps, Deployment & Monitoring

All of LLMOps, Deployment & Monitoring