Cookbook · Transformers
Attention
5 minattentiontransformersself-attention
The idea, in one analogy
Before guessing the next word in a sentence, you're allowed to glance back at every word you've already read and decide how much attention each one deserves. Take "The trophy didn't fit in the suitcase because it was too big": to figure out what "it" refers to, you weigh "trophy" and "suitcase" very differently depending on the rest of the sentence. That weighing act, deciding how much each earlier word matters to understanding the current one, is what "attention" mechanically computes.
The mechanism
Every position in the sequence produces three vectors from itself: a query ("what am I looking for"), a key ("what do I contain"), and a value ("what do I offer"). Attention compares each position's query against every other position's key, turns the resulting similarity scores into weights, and uses those weights to mix together the values, producing a new representation for that position that blends in whatever it needed from elsewhere in the sequence.
Causal vs. bidirectional attention
Causal (or "masked") attention restricts each position to only attend to itself and earlier positions, never the future. This is what makes autoregressive generation possible (see inference): the model is trained to predict each token using only what came before it, which matches exactly what's available at generation time, one token at a time. Causal masking sets the attention score to −∞ for any (query position, key position) pair where the key position is in the future, before the softmax, so those positions get exactly zero weight.
Bidirectional attention allows every position to attend to every other position, past and future. Models trained this way (BERT-style encoders) can build richer representations of a fixed piece of text since nothing is hidden from them, but they can't generate text left-to-right the way a causal model can, since at generation time the "future" tokens a bidirectional model wants to attend to don't exist yet. Decoder-only language models (see transformers) use causal attention specifically because generation is the goal.
A worked example: masking and attention on a real sentence
Take the 4-token sentence "The cat sat on." Causal masking means each token can only look at itself and the tokens before it:
| query \ key | The | cat | sat | on |
|---|---|---|---|---|
| The | ✓ | ✗ | ✗ | ✗ |
| cat | ✓ | ✓ | ✗ | ✗ |
| sat | ✓ | ✓ | ✓ | ✗ |
| on | ✓ | ✓ | ✓ | ✓ |
Before the softmax, the blocked (✗) positions get a score of −∞ so they vanish entirely once the softmax is applied. Suppose the raw QKᵀ scores for the allowed positions come out like this (blocked cells shown as −∞):
| query \ key | The | cat | sat | on |
|---|---|---|---|---|
| The | 0.8 | −∞ | −∞ | −∞ |
| cat | 0.3 | 0.9 | −∞ | −∞ |
| sat | 0.2 | 1.4 | 0.7 | −∞ |
| on | 0.1 | 0.6 | 1.2 | 0.4 |
After the softmax, each row turns into a probability distribution that sums to 100%:
| query \ key | The | cat | sat | on |
|---|---|---|---|---|
| The | 100% | 0% | 0% | 0% |
| cat | 35% | 65% | 0% | 0% |
| sat | 17% | 56% | 28% | 0% |
| on | 14% | 24% | 43% | 19% |
Reading the "sat" row: its output becomes 17% of "The"'s value vector, 56% of "cat"'s, and 28% of its own, mixed together. That mix, not just "sat" in isolation, is what gets passed on to the next layer. (Run notebooks/attention_mechanics.ipynb to see these exact numbers computed from the raw scores above.)
Why it beats what came before
Before transformers, sequence models (RNNs, LSTMs) processed tokens one at a time, in order, which meant training couldn't parallelize across the sequence length and long-range dependencies had to survive being passed through many sequential steps (they often didn't). Attention lets every position see every other position directly, in one operation, and lets the whole sequence be processed in parallel during training.
Where to look further
- Vaswani et al., "Attention Is All You Need": the original paper.
- transformers: how attention fits into the rest of a GPT-style architecture.
- inference: why causal attention specifically is what makes autoregressive generation work.