LM Studio MLX v1.8.1: Vision Model Batching and Better Caching
Melvin Vivas · X video post · 2026-05-15 · Open on X
Topics: LLMOps, Deployment & Monitoring, AI Dev Tools & Productivity · Level: intermediate
Summary
LM Studio's latest MLX engine update adds batching for vision models in beta and improves caching for faster inference. The post explains how to turn on the beta runtime.
Key points
- Batching for vision models is now available in beta.
- Caching improvements make inference faster overall.
- To enable: turn on Developer Mode, pick the beta runtime channel, and select LM Studio MLX v1.8.1.
- MLX is Apple's machine-learning framework for Apple Silicon Macs.
Resources mentioned
- LM Studio · tool · x.com · free
Desktop app for downloading and running LLMs locally, with a developer mode that serves models through an API.
Also in: Running LLMs Locally Without an Expensive Rig (Melvin Vivas on X · notes), Adding Vercel AI Gateway as a provider in AIBackends with Devin (Melvin Vivas on X · notes), LoRA Fine-Tune Qwen3.5-2B on Your Tweets with Unsloth Studio (Melvin Vivas on X · notes), AIBackends: An API Layer Between Your App and AI Models (Now with Jev) (Melvin Vivas on X · notes) and 31 more - MLX · repo · github.com · free
Apple's open-source array and ML framework for efficient inference on Apple Silicon.
Also in: Gemma 4 Gets Up to 3x Faster with MTP Drafters (Melvin Vivas on X · notes), 1-bit Bonsai 8B Runs On-Device on iPhone at 40+ tok/s (Melvin Vivas on X · notes), Qwen 3.5 2B Runs On-Device on iPhone with MLX (Melvin Vivas on X · notes)
Try this
- In LM Studio, turn on Developer Mode, choose the beta runtime channel and select MLX v1.8.1.
More in LLMOps, Deployment & Monitoring
- Running Qwen3.6-27B Fully in the Browser With WebGPU and wllama
- Unsloth MTP GGUFs make Qwen3.6 run 1.4x faster locally
- Cheap TTS serving: Qwen3-TTS on vLLM-Omni at $3 per 1M characters
- OpenRouter Pareto Code: cost-optimized coding router
- Speeding Up Gemma 4 Inference: MTP (3x) vs. DFlash Speculative Decoding (6x)
- Self-Host a Coding Model on QuickPod with llama-swap and Use It in Claude Code