AI Engineer Study Library

Serving Qwen3.8-27B (EXL3) with 262K Context on a 24GB RTX 3090

Melvin Vivas · X video post · 2026-09-01 · 0:12 · 102,962 views · Open on X

Topics: LLMOps, Deployment & Monitoring, LLM Fundamentals · Level: advanced

Summary

Melvin Vivas shares his results running Mia's (@MiaAI_lab) experimental Qwen3.8-27B EXL3 release on one 24GB RTX 3090. He posts the .env settings he used: MTP draft decoding, a 262K context window, 8/4-bit KV cache quantization and a 22GB GPU memory budget. With those settings he got about 64.5 tokens/s on average. The quoted post says this setup is a way to serve 200K+ context with dflash2 on 24GB-VRAM cards (RTX 3090/4090/5090).

Key points

Resources mentioned

Try this

More in LLMOps, Deployment & Monitoring

All of LLMOps, Deployment & Monitoring