AI Engineer Study Library

Running Qwen3.8-27B EXL3 on an RTX 3090 with 262K Context at ~64 tok/s

Melvin Vivas · X video post · 2026-09-15 · 0:12 · 356 views · Open on X

Topics: LLMOps, Deployment & Monitoring, LLM Fundamentals · Level: advanced

Summary

A short post sharing a working local-inference setup for Qwen3.8-27B quantized in the EXL3 format on one 24 GB RTX 3090. Credit for the setup goes to @MiaAI_lab. The post gives the exact .env settings: MTP draft decoding, a 262,144-token context, a quantized KV cache and a 22 GB GPU memory budget. It reports about 64.5 tokens/s on average and says RTX 4090 or newer GPUs could run faster because they support NVFP4.

Key points

Resources mentioned

Try this

More in LLMOps, Deployment & Monitoring

All of LLMOps, Deployment & Monitoring