Run Qwen3.8-27B locally with llama.cpp (llama-server)
Melvin Vivas · X post · 2026-08-14 · Open on X
Topics: LLMOps, Deployment & Monitoring, LLM Fundamentals · Level: intermediate
Summary
A complete llama-server command for running Unsloth's Q4_K_M GGUF build of Qwen3.8-27B locally from Hugging Face. It uses a 256K context, flash attention, a 4-bit KV cache, and Qwen's recommended sampling settings. The creator has only tested it for chat so far, not yet with agent harnesses.
Key points
- Load straight from Hugging Face: llama-server -hf unsloth/Qwen3.8-27B-GGUF:Q4_K_M
- -ngl 99 puts all layers on the GPU; -np 1 uses a single parallel slot
- -c 262144 sets a 256K context window; -fa on turns on flash attention
- --cache-type-k q4_0 and --cache-type-v q4_0 shrink the KV cache to 4-bit so the long context uses less VRAM
- Sampling: temp 0.6, top-p 0.95, top-k 20, min-p 0.00, presence_penalty 0.0, repeat_penalty 1.0
- --alias "qwen38-27b" gives the model a short name for OpenAI-compatible clients
- Tested for chat only so far, not yet with agent harnesses
Resources mentioned
- llama.cpp · repo · github.com · free
An open-source C/C++ engine for running GGUF models locally. Its llama-server command provides an OpenAI-compatible HTTP server.
Also in: Running LLMs Locally Without an Expensive Rig (Melvin Vivas on X · notes), llama.cpp / Llama-macOS v0.5.0 release (Melvin Vivas on X · notes), Run llama.cpp GGUF Checkpoints in Hugging Face Transformers (Melvin Vivas on X · notes), llama.cpp v0.4.1 release announcement (Melvin Vivas on X · notes) and 28 more - Qwen3.8-27B-GGUF · tool · huggingface.co · free
Unsloth's GGUF quantized versions of the Qwen3.8 27B model, ready for llama.cpp and other GGUF runtimes.
Also in: Unsloth passes 500M model downloads on Hugging Face (Melvin Vivas on X · notes), Running Qwen3.8 27B Locally on an RTX 3090 with llama.cpp and the Pi Harness (Melvin Vivas on X · notes), Unsloth releases Qwen3.8-27B GGUF quantizations (Melvin Vivas on X · notes) - Qwen3.8-27B · tool · huggingface.co · free
A 27B-parameter open-weight model from Alibaba's Qwen family that you can run locally or call through hosted APIs.
Also in: Deploying Qwen3.8 27B on Hugging Face Inference Endpoints (Melvin Vivas on X · notes), Using Hugging Face credits: Jobs, Inference Endpoints and Open Models (Melvin Vivas on X · notes), Running Qwen3.8-27B Locally on an M5 Max MacBook with Inco Splash (Melvin Vivas on X · notes), Ternary Bonsai 2 27B: 9x smaller model keeping 98.2% of benchmark scores (Melvin Vivas on X · notes) and 24 more
Try this
- Run Qwen3.8-27B locally with the llama-server command shared in the post
- Use the recommended sampling settings (temp 0.6, top-p 0.95, top-k 20)
- Connect the local llama-server endpoint to an agent harness and test tool use
More in LLMOps, Deployment & Monitoring
- AIBackends 0.4.0: Liquid AI LFM2.5 Models on llama.cpp and Transformers
- AIBackends Adds Support for LFM2.5-VL-3B
- Running LFM2.5-2.6B Q4_K_M Locally with llama.cpp and Pi
- Unsloth releases Qwen3.8-27B GGUF quantizations
- Run Muse Glimmer 30B Locally on an RTX 3090 with llama.cpp
- Running Muse Glimmer 30B Locally with llama.cpp and the Hermes Agent