Serving Qwen3.8-27B (EXL3) with 262K Context on a 24GB RTX 3090
Melvin Vivas · X video post · 2026-09-01 · 0:12 · 102,962 views · Open on X
Topics: LLMOps, Deployment & Monitoring, LLM Fundamentals · Level: advanced
Summary
Melvin Vivas shares his results running Mia's (@MiaAI_lab) experimental Qwen3.8-27B EXL3 release on one 24GB RTX 3090. He posts the .env settings he used: MTP draft decoding, a 262K context window, 8/4-bit KV cache quantization and a 22GB GPU memory budget. With those settings he got about 64.5 tokens/s on average. The quoted post says this setup is a way to serve 200K+ context with dflash2 on 24GB-VRAM cards (RTX 3090/4090/5090).
Key points
- Hardware: a single RTX 3090 (24GB VRAM). The quoted post says RTX 4090 and 5090 owners can run it too.
- Model: Qwen3.8-27B in EXL3 format (the ExLlamaV3 quantization format). It is an experimental release, so expect bugs.
- .env config: DRAFT=mtp, CONTEXT_SIZE=262144, CACHE_QUANT=8,4, GPU_MEM_GB=22.
- DRAFT=mtp appears to turn on multi-token-prediction draft decoding (speculative decoding) to speed up generation.
- CACHE_QUANT=8,4 seems to set the KV cache quantization (probably 8-bit keys and 4-bit values), which is what lets a 262K context fit in 24GB.
- GPU_MEM_GB=22 caps VRAM use below the full 24GB to leave some headroom.
- Measured speed at 262K context: 65.98, 63.23 and 64.41 tok/s over three runs, about 64.5 tok/s on average.
- The quoted post says this is the only known way to serve Qwen3.8-27B with 200K+ context using dflash2 on 24GB VRAM.
Resources mentioned
- MiaAI_lab (@MiaAI_lab on X) · person · x.com · free
X account that published the experimental Qwen3.8-27B EXL3 release for serving long contexts on 24GB GPUs.
Also in: Running Qwen3.8-27B EXL3 on an RTX 3090 with 262K Context at ~64 tok/s (Melvin Vivas on X · notes), Long-Context Local LLM on an RTX 3090: .env Config and ~64 tok/s Benchmark (Melvin Vivas on X · notes), Running Qwen3.8-27B EXL3 Locally on an RTX 3090 with 220K+ Context (Melvin Vivas on X · notes) - Qwen3.8-27B EXL3 · tool · huggingface.co · free
Experimental EXL3-quantized build of the Qwen3.8-27B model that can run with 200K+ context on a single 24GB GPU.
Also in: Running Qwen3.8-27B EXL3 on an RTX 3090 with 262K Context at ~64 tok/s (Melvin Vivas on X · notes) - ExLlamaV3 (EXL3) · repo · github.com · free
turboderp's inference library and the EXL3 quantization format it uses to run quantized LLMs on consumer NVIDIA GPUs.
Also in: Running Qwen3.8-27B EXL3 on an RTX 3090 with 262K Context at ~64 tok/s (Melvin Vivas on X · notes), Running Qwen3.8-27B EXL3 Locally on an RTX 3090 with 220K+ Context (Melvin Vivas on X · notes) - dflash2 · repo · github.com · free
Open-source speculative decoding method that uses a block-diffusion drafter to speed up LLM inference, with support for Gemma 4.
Also in: Speeding Up Gemma 4 Inference: MTP (3x) vs. DFlash Speculative Decoding (6x) (Melvin Vivas on X · notes)
Try this
- Use this .env config on a 24GB GPU: DRAFT=mtp, CONTEXT_SIZE=262144, CACHE_QUANT=8,4, GPU_MEM_GB=22.
- Benchmark tokens/s over several runs and average them, as Melvin did.
- Self-host a long-context (262K) 27B model on one consumer GPU and benchmark tokens/s with different KV cache quantization and speculative decoding settings.
More in LLMOps, Deployment & Monitoring
- Docker Template Bundling Coding Agents for GPU Cloud Hosting
- Why Claude Fable 5.1 Costs Less: Cheaper Cache Reads
- Running Qwen3.8-27B EXL3 Locally on an RTX 3090 with 220K+ Context
- Ollama vs vLLM: From Local AI Demo to Production Inference Serving
- Running Qwen3.8 27B Locally on an RTX 3090 with llama.cpp and the Pi Harness
- Low-Cost Agent Run: DeepSeek V4 Flash via OpenRouter in ohmypi