How inference engines work: the full life of an LLM request
Melvin Vivas · X post · 2026-09-12 · Open on X
Topics: LLMOps, Deployment & Monitoring, LLM Fundamentals · Level: advanced
Summary
The creator bookmarks a quoted post that shares slides from a talk on how LLM inference engines work. The talk follows a request from start to finish: the engine itself, KV and prefix caching, continuous batching, paged attention, chunked prefill, sampling, and agentic loops seen from inside the engine.
Key points
- Inference engines manage each request from the moment it arrives until the last token is sent back.
- KV caching reuses attention keys and values that were already computed; prefix caching reuses them across requests that start with the same prompt.
- Continuous batching adds and removes requests from the batch at every step to keep GPU usage high.
- Paged attention stores the KV cache in fixed-size blocks, like memory pages, to cut down on fragmentation.
- Chunked prefill splits long prompts into chunks so prefill work can run alongside decoding.
- Sampling turns the model's output probabilities (logits) into the next token, and agentic loops can be optimized from inside the engine.
Resources mentioned
- How inference engines actually work (talk slides) · article · x.com · free
Talk slides that follow an LLM request through an inference engine: caching, batching, paged attention, chunked prefill, sampling and agentic loops.
Try this
- Study how inference engines work: KV/prefix caching, continuous batching, paged attention, chunked prefill and sampling.
More in LLMOps, Deployment & Monitoring
- Running Qwen3.8-27B EXL3 on an RTX 3090 with 262K Context at ~64 tok/s
- llama.cpp v0.4.1 release announcement
- DeepSeek V4.1 Flash Off-Peak Pricing as a Cheap Fallback Model
- SGLang v0.5.19 release: new models and beam search
- llama.cpp's built-in web UI for testing local models
- Local Speech-to-Text with LFM2.5-Audio-1.5B and llama.cpp