Unsloth releases Qwen3.8-27B GGUF quantizations
Melvin Vivas · X post · 2026-08-14 · Open on X
Topics: LLMOps, Deployment & Monitoring, Industry Trends & Job Market · Level: beginner
Summary
Announces that Unsloth has published GGUF builds of Qwen3.8-27B, so you can run it locally with llama.cpp-compatible tools. The post thanks Unsloth AI.
Key points
- Unsloth publishes GGUF quantizations of new open models soon after they come out
- GGUF files work with llama.cpp-based local runtimes
Resources mentioned
- Unsloth · tool · unsloth.ai · free
Open-source library for fast, memory-efficient fine-tuning of open LLMs (LoRA/QLoRA) on a single GPU or in Colab.
Also in: Run Laya Decision models locally with Unsloth on 4GB RAM (Melvin Vivas on X · notes), Unsloth passes 500M model downloads on Hugging Face (Melvin Vivas on X · notes), Run Qwen-Image-2.1 locally on 12GB VRAM with Unsloth GGUFs (Melvin Vivas on X · notes), Base vs fine-tuned Gemma 4 E2B as a model router (Melvin Vivas on X · notes) and 29 more - Qwen3.8-27B-GGUF · tool · huggingface.co · free
Unsloth's GGUF quantized versions of the Qwen3.8 27B model, ready for llama.cpp and other GGUF runtimes.
Also in: Unsloth passes 500M model downloads on Hugging Face (Melvin Vivas on X · notes), Running Qwen3.8 27B Locally on an RTX 3090 with llama.cpp and the Pi Harness (Melvin Vivas on X · notes), Run Qwen3.8-27B locally with llama.cpp (llama-server) (Melvin Vivas on X · notes)
More in LLMOps, Deployment & Monitoring
- AIBackends Adds Support for LFM2.5-VL-3B
- Running LFM2.5-2.6B Q4_K_M Locally with llama.cpp and Pi
- Run Qwen3.8-27B locally with llama.cpp (llama-server)
- Run Muse Glimmer 30B Locally on an RTX 3090 with llama.cpp
- Running Muse Glimmer 30B Locally with llama.cpp and the Hermes Agent
- Run Muse Glimmer 30B Locally with llama.cpp and Connect It to Hermes Agent