AI Engineer Study Library

Running Qwen3.8-27B EXL3 Locally on an RTX 3090 with 220K+ Context

Melvin Vivas · X post · 2026-09-01 · Open on X

Topics: LLMOps, Deployment & Monitoring, LLM Fundamentals · Level: advanced

Summary

The creator shares the .env settings he used to run Qwen3.8-27B in EXL3 quantization on a single 24GB RTX 3090 with a 220K-token context, using a DFlash2 draft model for speculative decoding. His usual smoke test is building a CRM app with SQLite. The output was decent but needed a few iterations, and the kanban tasks were draggable. The quoted post (from @MiaAI_lab) reports about 64.5 tok/s at 262K context using MTP drafting.

Key points

Resources mentioned

Try this

More in LLMOps, Deployment & Monitoring

All of LLMOps, Deployment & Monitoring