Skip to content
Ishaan Reddy

Cookbook · Inference

Serving

3 minservinginferencevllmkv-cache

Why serving is a different problem than training

Training cares about throughput across a fixed, known dataset. Serving cares about latency and throughput across requests arriving unpredictably, from different users, at different times, wanting different amounts of output. A serving system's job is to keep a GPU busy and responsive under that unpredictability, which needs different techniques than the training loop does.

Batching requests together

Running one request at a time wastes most of a GPU's parallel compute capacity, but naively batching requests together runs into a problem: different requests finish generating at different lengths, so a simple fixed batch has to wait for its slowest member before starting the next batch.

Continuous batching (also called in-flight batching) solves this by not waiting: as soon as one request in a batch finishes, a new request from the queue takes its slot immediately, without waiting for every other request in the batch to finish first. This keeps GPU utilization high under real, uneven traffic, instead of only under artificially uniform request lengths.

KV cache management at scale

A single request's KV cache (see inference for the per-request mechanism) is one thing; a serving system juggling hundreds of concurrent requests, each with its own growing cache, is another. Naive per-request cache allocation wastes memory through fragmentation, similar to how naive memory allocation in any system leaves gaps between allocated blocks. PagedAttention (the technique behind vLLM) borrows the idea of virtual memory paging from operating systems: it splits each request's KV cache into fixed-size blocks that don't need to be contiguous in memory, letting the system pack requests together far more densely than a naive contiguous allocation would allow.

Speculative decoding

Autoregressive generation is inherently sequential (see inference): one token at a time, each depending on the last. Speculative decoding speeds this up by using a small, fast "draft" model to guess several tokens ahead, then having the full, accurate model verify all of those guesses in a single parallel forward pass instead of one sequential step per token. Correct guesses are accepted for free; the first incorrect one gets corrected by the full model, and generation continues from there. When the draft model's guesses are frequently right, this can meaningfully speed up generation without changing the output distribution at all, since the full model still verifies (and can override) every guess.

Where to look further