Running Qwen3.8-27B EXL3 Locally on an RTX 3090 with 220K+ Context
Melvin Vivas · X post · 2026-09-01 · Open on X
Topics: LLMOps, Deployment & Monitoring, LLM Fundamentals · Level: advanced
Summary
The creator shares the .env settings he used to run Qwen3.8-27B in EXL3 quantization on a single 24GB RTX 3090 with a 220K-token context, using a DFlash2 draft model for speculative decoding. His usual smoke test is building a CRM app with SQLite. The output was decent but needed a few iterations, and the kanban tasks were draggable. The quoted post (from @MiaAI_lab) reports about 64.5 tok/s at 262K context using MTP drafting.
Key points
- Creator's config: GPU_MEM_GB=22, CONTEXT_SIZE=220000, CACHE_QUANT=4, DRAFT=dflash2, DRAFT_DIR=models/Qwen3.8-27B-DFlash2-EXL3-5.0bpw
- The draft model is a 5.0bpw EXL3 DFlash2 model, used for speculative decoding to speed up generation
- Quantizing the KV cache (CACHE_QUANT=4) makes very long contexts fit in 24GB of VRAM
- Quoted config: DRAFT=mtp, CONTEXT_SIZE=262144, CACHE_QUANT=8,4, GPU_MEM_GB=22 on an RTX 3090
- Reported speed at 262K context: 65.98, 63.23 and 64.41 tok/s, averaging about 64.5 tok/s
- Smoke test: ask the model to build a CRM app backed by SQLite with a draggable kanban board; it took several iterations
Resources mentioned
- Qwen3.8-27B · tool · huggingface.co · free
A 27B-parameter open-weight model from Alibaba's Qwen family that you can run locally or call through hosted APIs.
Also in: Deploying Qwen3.8 27B on Hugging Face Inference Endpoints (Melvin Vivas on X · notes), Using Hugging Face credits: Jobs, Inference Endpoints and Open Models (Melvin Vivas on X · notes), Running Qwen3.8-27B Locally on an M5 Max MacBook with Inco Splash (Melvin Vivas on X · notes), Ternary Bonsai 2 27B: 9x smaller model keeping 98.2% of benchmark scores (Melvin Vivas on X · notes) and 24 more - ExLlamaV3 (EXL3) · repo · github.com · free
turboderp's inference library and the EXL3 quantization format it uses to run quantized LLMs on consumer NVIDIA GPUs.
Also in: Running Qwen3.8-27B EXL3 on an RTX 3090 with 262K Context at ~64 tok/s (Melvin Vivas on X · notes), Serving Qwen3.8-27B (EXL3) with 262K Context on a 24GB RTX 3090 (Melvin Vivas on X · notes) - Qwen3.8-27B-DFlash2-EXL3-5.0bpw · tool · huggingface.co · free
A DFlash2 draft model for Qwen3.8-27B, used for speculative decoding. - NVIDIA RTX 3090 · tool · nvidia.com · paid
A consumer GPU with 24GB of VRAM, often used to run LLMs locally.
Also in: Running Hermes Agent with Local Models on an RTX 3090 (Melvin Vivas on X · notes) - SQLite · tool · sqlite.org · free
A lightweight embedded SQL database stored in a file.
Also in: Why Plan Mode Still Matters in AI Coding Agents (Codex, Claude, Cursor) (Melvin Vivas on X · notes), One-Shotting a Slack Clone with Fable 5 (Melvin Vivas on X · notes), Google Antigravity 2.0: Multi-Agent Desktop App Walkthrough (Subagents, Scheduled Tasks) (Melvin Vivas on X · notes), Local Qwen 3.5 Adds an Express API and SQLite to Make an App Full-Stack (Melvin Vivas on X · notes) - MiaAI_lab (@MiaAI_lab on X) · person · x.com · free
X account that published the experimental Qwen3.8-27B EXL3 release for serving long contexts on 24GB GPUs.
Also in: Running Qwen3.8-27B EXL3 on an RTX 3090 with 262K Context at ~64 tok/s (Melvin Vivas on X · notes), Long-Context Local LLM on an RTX 3090: .env Config and ~64 tok/s Benchmark (Melvin Vivas on X · notes), Serving Qwen3.8-27B (EXL3) with 262K Context on a 24GB RTX 3090 (Melvin Vivas on X · notes)
Try this
- Try the shared .env configs to run Qwen3.8-27B locally on a 24GB GPU
- Use a fixed smoke test (e.g., a CRM app with SQLite) to compare local models
- Build a CRM app with SQLite and a draggable kanban board as a smoke test for coding models
More in LLMOps, Deployment & Monitoring
- Long-Context Local LLM on an RTX 3090: .env Config and ~64 tok/s Benchmark
- Docker Template Bundling Coding Agents for GPU Cloud Hosting
- Why Claude Fable 5.1 Costs Less: Cheaper Cache Reads
- Serving Qwen3.8-27B (EXL3) with 262K Context on a 24GB RTX 3090
- Ollama vs vLLM: From Local AI Demo to Production Inference Serving
- Running Qwen3.8 27B Locally on an RTX 3090 with llama.cpp and the Pi Harness