Run Qwen3.5-27B GGUF with llama-server for Claude Code on an RTX 3090
Melvin Vivas · X post · 2026-03-08 · Open on X
Topics: AI Dev Tools & Productivity, LLMOps, Deployment & Monitoring, LLM Fundamentals · Level: advanced
Summary
The creator gives the exact llama.cpp command for serving Unsloth's Qwen3.5-27B GGUF (Q4_K_M) on one RTX 3090. He then uses it as the model behind Claude Code with the frontend skill to build UI, and calls local coding on a consumer GPU promising.
Key points
- Command: llama-server -hf unsloth/Qwen3.5-27B-GGUF:Q4_K_M -ngl 99 -c 262144 -fa on
- -hf downloads the model straight from the Hugging Face repo unsloth/Qwen3.5-27B-GGUF.
- Q4_K_M is a 4-bit quantization that lets a 27B model fit on a 24 GB RTX 3090.
- -ngl 99 offloads all layers to the GPU.
- -c 262144 sets a 256K-token context window.
- -fa on turns on flash attention.
- The local server was used with Claude Code plus the frontend skill to generate frontend output.
Resources mentioned
- Qwen3.5-27B-GGUF (Unsloth) · tool · huggingface.co · free
Unsloth's GGUF quantizations of Qwen3.5-27B, made for llama.cpp. - Unsloth · tool · unsloth.ai · free
Open-source library for fast, memory-efficient fine-tuning of open LLMs (LoRA/QLoRA) on a single GPU or in Colab.
Also in: Run Laya Decision models locally with Unsloth on 4GB RAM (Melvin Vivas on X · notes), Unsloth passes 500M model downloads on Hugging Face (Melvin Vivas on X · notes), Run Qwen-Image-2.1 locally on 12GB VRAM with Unsloth GGUFs (Melvin Vivas on X · notes), Base vs fine-tuned Gemma 4 E2B as a model router (Melvin Vivas on X · notes) and 29 more - llama.cpp · repo · github.com · free
An open-source C/C++ engine for running GGUF models locally. Its llama-server command provides an OpenAI-compatible HTTP server.
Also in: Running LLMs Locally Without an Expensive Rig (Melvin Vivas on X · notes), llama.cpp / Llama-macOS v0.5.0 release (Melvin Vivas on X · notes), Run llama.cpp GGUF Checkpoints in Hugging Face Transformers (Melvin Vivas on X · notes), llama.cpp v0.4.1 release announcement (Melvin Vivas on X · notes) and 28 more - Claude Code · tool · code.claude.com · paid · recommended by both Bashiri Smith & Melvin Vivas
Build agents and pipelines from the terminal; the guide's main agentic coding tool.
Also in: Create Claude Code Plugins with /plugin-authoring (Melvin Vivas on X · notes), Claude Code mods: customize behavior and UI with plugins (Melvin Vivas on X · notes), AI Engineer Roadmap Overview: From ML Foundations to RAG, Agents & Ops (Bashiri Smith on Facebook · notes), SkillsBento: Free Plugin Marketplace for Codex and Claude Code (Melvin Vivas on X · notes) and 101 more - Claude Code frontend skill · tool · github.com · free
A Claude Code skill that guides frontend and UI generation.
Try this
- Run: llama-server -hf unsloth/Qwen3.5-27B-GGUF:Q4_K_M -ngl 99 -c 262144 -fa on
- Point Claude Code at the local server and try the frontend skill
- Set up a fully local coding assistant: llama.cpp with Qwen3.5-27B as the model behind Claude Code on a 24 GB GPU
More in AI Dev Tools & Productivity
- Qwen 3.5 Is the Local Model That Holds Up in Claude Code
- Local Qwen 3.5 Adds an Express API and SQLite to Make an App Full-Stack
- Unlimited Free AI Coding with Unsloth's Qwen 3.5 on an RTX 3090
- Running the LTX 2.3 Video Model Locally on an RTX 3090 with ComfyUI
- GPT-5.4 Now Available in Windsurf (Plus Arena Mode)
- Turning Flat Maps Into 3D Street-Style Art With Lovart