Skip to content
Ishaan Reddy

Cookbook · Inference

Inference (generation)

4 mininferencesamplingkv-cachetransformers

The idea, in one analogy

Writing a sentence one word at a time, committing to each word before seeing what comes next, and never being able to go back and revise an earlier word once it's chosen. That's autoregressive generation: a trained transformer (see transformers) produces one token, appends it to the sequence, feeds the whole thing back in, and produces the next token, over and over until it decides to stop.

Why generation is fundamentally sequential

During training, a model sees a full sequence at once and predicts every position in parallel (see pretraining). Generation can't do that, because the tokens after the current position don't exist yet; each new token depends on every token generated before it. This is why generation is called "decode-bound": it's limited by how many sequential steps it takes, not by how much parallel compute is available, unlike training or a single forward pass over an existing sequence, both of which can use all of a GPU's parallelism at once.

Turning a probability distribution into a chosen token

At each step, the model outputs a probability distribution over the entire vocabulary for what comes next (the same distribution used to compute loss during training). Deciding which single token to pick from that distribution is a separate choice, called the sampling strategy:

  • Greedy decoding: always pick the single highest-probability token. Deterministic and simple, but tends to produce repetitive, generic text, since it never takes a slightly-lower-probability path even when that path leads somewhere better.
  • Temperature sampling: divide the logits by a temperature T before the softmax (the same operation covered in distillation's temperature scaling) before sampling. T < 1 sharpens the distribution toward the top choices; T > 1 flattens it, making less-likely tokens more competitive; T = 1 leaves it unchanged.
  • Top-k sampling: restrict sampling to only the k highest-probability tokens, zeroing out everything else, then sample from that restricted set. Prevents the rare, very-low-probability tail from ever getting picked, without being as rigid as greedy decoding.
  • Top-p (nucleus) sampling: instead of a fixed count, keep the smallest set of top tokens whose cumulative probability exceeds p (e.g. 0.9), then sample from that set. Adapts to the shape of the distribution: a confident prediction keeps a small set of tokens, an uncertain one keeps a larger set.
  • Beam search: instead of committing to one token per step, keep the n most likely full sequences so far ("beams") at each step, expand all of them, and keep only the top n again. Common for tasks with one clearly correct output (translation, summarization); less common for open-ended generation, where it tends to produce bland, safe text.

The KV cache

Recomputing attention over the entire sequence from scratch at every single generation step would mean the cost of generating token 100 requires redoing all the work already done for tokens 1 through 99. The KV cache avoids this: since a causally-masked position's keys and values (see attention) never change once computed, they're stored the first time and simply reused, reappended to as new tokens arrive, rather than recomputed every step. This is what makes long generations tractable at all, but the cache itself grows with sequence length, and at long context lengths its memory footprint can rival or exceed the model's own weights, a real constraint at the serving layer (see serving for how production systems manage this at scale).

Where to look further

  • Holtzman et al., "The Curious Case of Neural Text Degeneration": the paper introducing nucleus (top-p) sampling, with a clear explanation of why greedy/beam search produce degenerate repetitive text.
  • serving: how KV cache management and batching work at the scale of many concurrent requests.
  • distillation: temperature scaling used for a different purpose (softening a teacher's distribution rather than sampling from it).