Running Qwen3.6-27B Fully in the Browser With WebGPU and wllama
Melvin Vivas · X video post · 2026-05-19 · Open on X
Topics: LLMOps, Deployment & Monitoring, LLM Fundamentals, Industry Trends & Job Market · Level: advanced
Summary
Xuan Son Nguyen (@ngxson) of Hugging Face showed Qwen3.6-27B running entirely in a web browser on WebGPU, with no cloud. The model uses a Q2_K_XL GGUF quantization (about 11.2 GB) and is loaded through wllama.
Key points
- Qwen3.6-27B (27B parameters) runs entirely inside the browser.
- It uses Q2_K_XL GGUF quantization, about 11.2 GB.
- It is loaded with wllama, a WebAssembly/WebGPU build of llama.cpp for browsers.
- 100% local on WebGPU, with no cloud calls.
- Heavy quantization is what makes a 27B model fit in the browser.
Resources mentioned
- Qwen3.6-27B · tool · huggingface.co · free
Dense open-weight Qwen model. The MTP GGUF version runs at about 140 tokens/s.
Also in: Meta Muse Glimmer-30B: Open Weights Model That Beats Qwen3.6 37B on Agentic Tasks (Melvin Vivas on X · notes), Unsloth MTP GGUFs make Qwen3.6 run 1.4x faster locally (Melvin Vivas on X · notes), Qwen 3.6 27B: An Open Local Model with Benchmarks Near Claude Opus 4.5 (Melvin Vivas on X · notes), Run Qwen3.6-27B Locally in 18GB RAM with Unsloth GGUFs (Melvin Vivas on X · notes) and 1 more - wllama · repo · github.com · free
A WebAssembly/WebGPU binding of llama.cpp for running GGUF models in the browser. - Xuan Son Nguyen (@ngxson) · person · x.com · free
Hugging Face engineer who builds wllama and works on llama.cpp. - WebGPU · tool · github.com · free
A browser API for GPU compute that lets machine-learning models run fast on the user's device.
Also in: Real-Time Speech Transcription in the Browser with Voxtral and WebGPU (Melvin Vivas on X · notes)
Try this
- Build a fully local in-browser chat app with wllama and a quantized GGUF model.
More in LLMOps, Deployment & Monitoring
- Running GLM-5.2 locally on a 256GB Mac with Unsloth
- Tracking LLM Spend, Tokens & Guardrails with OpenRouter's Activity Explorer
- Bonsai Image Model Running In-Browser with WebGPU
- Unsloth MTP GGUFs make Qwen3.6 run 1.4x faster locally
- Cheap TTS serving: Qwen3-TTS on vLLM-Omni at $3 per 1M characters
- LM Studio MLX v1.8.1: Vision Model Batching and Better Caching