AI Engineer Study Library

Serving Ornith-1.5-35B-A3B at 128k Context on an RTX 3090 with llama.cpp

Melvin Vivas · X post · 2026-08-20 · Open on X

Topics: LLMOps, Deployment & Monitoring, LLM Fundamentals · Level: intermediate

Summary

Gives an exact llama-server command for running the Q4_K_M quantization of Ornith-1.5-35B-A3B with a 128k context window, with every layer on a single RTX 3090. It also reports measured speeds, which makes it a practical reference for serving a local LLM.

Key points

Resources mentioned

Try this

More in LLMOps, Deployment & Monitoring

All of LLMOps, Deployment & Monitoring