AI Engineer Study Library

Running Qwen3.8 27B Locally on an RTX 3090 with llama.cpp and the Pi Harness

Melvin Vivas · X video post · 2026-08-31 · 0:23 · 6,224 views · Open on X

Topics: LLMOps, Deployment & Monitoring, AI Dev Tools & Productivity, LLM Fundamentals · Level: intermediate

Summary

A 23-second demo with no speech. It shows Unsloth's Qwen3.8 27B model, quantized to Q4_K_M in GGUF format, running locally through llama.cpp on one RTX 3090. The creator says Pi (@pidotdev) is the best harness he has used with this model so far. His point is that the whole setup is local and free.

Key points

Resources mentioned

Try this

More in LLMOps, Deployment & Monitoring

All of LLMOps, Deployment & Monitoring