AI Engineer Study Library

Long-Context Local LLM on an RTX 3090: .env Config and ~64 tok/s Benchmark

Melvin Vivas · X video post · 2026-09-03 · 0:12 · 841 views · Open on X

Topics: LLMOps, Deployment & Monitoring · Level: advanced

Summary

Melvin Vivas shares a short post saying a local LLM setup from @MiaAI_lab runs on a single 24 GB RTX 3090. He gives the .env settings he used: an MTP draft for speculative decoding, a 262K-token context, a quantized KV cache and a 22 GB GPU memory limit. He measured about 64.5 tokens/s on average over three runs. The post doesn't name the model or the inference engine, so you'd need to open the original @MiaAI_lab post to get those.

Key points

Resources mentioned

Try this

More in LLMOps, Deployment & Monitoring

All of LLMOps, Deployment & Monitoring