Running Qwen3.8 27B Locally on an RTX 3090 with llama.cpp and the Pi Harness
Melvin Vivas · X video post · 2026-08-31 · 0:23 · 6,224 views · Open on X
Topics: LLMOps, Deployment & Monitoring, AI Dev Tools & Productivity, LLM Fundamentals · Level: intermediate
Summary
A 23-second demo with no speech. It shows Unsloth's Qwen3.8 27B model, quantized to Q4_K_M in GGUF format, running locally through llama.cpp on one RTX 3090. The creator says Pi (@pidotdev) is the best harness he has used with this model so far. His point is that the whole setup is local and free.
Key points
- Model: Qwen3.8 27B in Unsloth's GGUF release, using the Q4_K_M quantization.
- Inference engine: llama.cpp, which runs GGUF models on local hardware.
- Hardware: one NVIDIA RTX 3090 (24 GB VRAM) runs the 27B model at 4-bit Q4_K_M quantization.
- Harness: the creator found Pi (@pidotdev) the best harness to pair with this local model so far.
- Q4_K_M is a 4-bit k-quant. It cuts memory use enough to fit a 27B model on a single consumer GPU, and is usually a good trade-off between quality and size.
- The setup costs nothing to run: it is local, with no API fees.
Resources mentioned
- Qwen3.8-27B-GGUF · tool · huggingface.co · free
Unsloth's GGUF quantized versions of the Qwen3.8 27B model, ready for llama.cpp and other GGUF runtimes.
Also in: Unsloth passes 500M model downloads on Hugging Face (Melvin Vivas on X · notes), Run Qwen3.8-27B locally with llama.cpp (llama-server) (Melvin Vivas on X · notes), Unsloth releases Qwen3.8-27B GGUF quantizations (Melvin Vivas on X · notes) - Unsloth · tool · unsloth.ai · free
Open-source library for fast, memory-efficient fine-tuning of open LLMs (LoRA/QLoRA) on a single GPU or in Colab.
Also in: Run Laya Decision models locally with Unsloth on 4GB RAM (Melvin Vivas on X · notes), Unsloth passes 500M model downloads on Hugging Face (Melvin Vivas on X · notes), Run Qwen-Image-2.1 locally on 12GB VRAM with Unsloth GGUFs (Melvin Vivas on X · notes), Base vs fine-tuned Gemma 4 E2B as a model router (Melvin Vivas on X · notes) and 29 more - Pi · tool · x.com · free
A customizable coding agent that can be extended through its Extensions API and connected to several model providers.
Also in: Sign in with ChatGPT: Setting Usage Limits for Each App (Melvin Vivas on X · notes), Pi reaches v1.0 (Melvin Vivas on X · notes), Claude Code mods: customize behavior and UI with plugins (Melvin Vivas on X · notes), Customizing Your Coding Setup with Pi Coding Agent Extensions (Melvin Vivas on X · notes) and 39 more - llama.cpp · repo · github.com · free
An open-source C/C++ engine for running GGUF models locally. Its llama-server command provides an OpenAI-compatible HTTP server.
Also in: Running LLMs Locally Without an Expensive Rig (Melvin Vivas on X · notes), llama.cpp / Llama-macOS v0.5.0 release (Melvin Vivas on X · notes), Run llama.cpp GGUF Checkpoints in Hugging Face Transformers (Melvin Vivas on X · notes), llama.cpp v0.4.1 release announcement (Melvin Vivas on X · notes) and 28 more
Try this
- Download the Q4_K_M file from unsloth/Qwen3.8-27B-GGUF on Hugging Face.
- Run it with llama.cpp on a GPU with about 24 GB of VRAM, such as an RTX 3090.
- Connect the local model to the Pi harness to use it as a free local coding assistant.
- Build a fully local, free AI coding assistant: Qwen3.8 27B (Unsloth Q4_K_M GGUF) served by llama.cpp on a single RTX 3090 and driven by the Pi harness.
More in LLMOps, Deployment & Monitoring
- Running Qwen3.8-27B EXL3 Locally on an RTX 3090 with 220K+ Context
- Serving Qwen3.8-27B (EXL3) with 262K Context on a 24GB RTX 3090
- Ollama vs vLLM: From Local AI Demo to Production Inference Serving
- Low-Cost Agent Run: DeepSeek V4 Flash via OpenRouter in ohmypi
- Serving Qwen3.8 27B FP8 on an H100 with Baseten dedicated inference
- Running GLM 5.3 Flash on Baseten with the Pi Coding Agent