AI Engineer Study Library

Running Ornith-1.5-35B-A3B (Q4_K_M) at 128k Context on an RTX 3090 with llama.cpp

Melvin Vivas · X video post · 2026-08-20 · 1:51 · 439 views · Open on X

Topics: LLMOps, Deployment & Monitoring, LLM Fundamentals · Level: intermediate

Summary

This is a short local-inference benchmark. It shows how to serve the Ornith-1.5-35B-A3B model, quantized to Q4_K_M GGUF, with llama.cpp's llama-server on a single 24GB RTX 3090. All layers run on the GPU with a 128k context window. The post gives the exact command (flash attention and a q8_0-quantized KV cache) and reports generation speeds by output length and by context size.

Key points

Resources mentioned

Try this

More in LLMOps, Deployment & Monitoring

All of LLMOps, Deployment & Monitoring