Cookbook · LLMs
Pretraining
5 minpretrainingtraining-loopcross-entropy-losscheckpointing
The idea, in one analogy
Pretraining is teaching someone a language by having them read an enormous pile of text and, at every single point, guess the next word before seeing it, then correcting them each time they're wrong. No grammar lessons, no labeled examples of "this sentence is correct," just relentless next-word prediction, at a scale where getting good at that game forces the model to implicitly learn grammar, facts, and reasoning patterns as a side effect.
What happens each training step
A batch of token sequences goes through the model (see transformers), producing a predicted probability distribution over the vocabulary at every position. That's the forward pass.
Next comes the loss: compare the predicted distribution at each position against the actual next token, using cross-entropy loss, a measure of how surprised the model was by the real answer. Lower is better; a loss of 0 would mean perfect, certain prediction.
The backward pass computes the gradient of that loss with respect to every parameter in the model: which direction each weight should move to make this specific mistake less likely next time.
Then the optimizer step nudges every weight slightly in that direction. AdamW is the standard choice for transformers; it adapts the step size per-parameter and applies weight decay (a mild pull toward zero that helps generalization).
Repeat this millions of times across a huge corpus, and "guess the next token" turns into a model that can hold a conversation.
A worked example: what loss means in practice
Say the training text is "...the cat sat on the mat" and the model has just seen everything up to "the" and needs to predict "mat." Early in training, the model's predicted probabilities for plausible next words might look like this:
| Candidate | Probability |
|---|---|
| mat (correct) | 4% |
| floor | 35% |
| dog | 28% |
| roof | 18% |
Cross-entropy loss for this step is −log(0.04) ≈ 3.22: a high loss, because the model assigned very little probability to the right answer. Later in training, after the model has seen this pattern many times:
| Candidate | Probability |
|---|---|
| mat (correct) | 72% |
| floor | 14% |
| dog | 4% |
| roof | 10% |
Loss is now −log(0.72) ≈ 0.33, much lower, because the model is now confident and correct. That's the entire training signal: nudge the weights so the probability mass shifts toward the token that came next, over and over, across billions of examples.
A rough sense of scale for what different loss values mean in practice (using a GPT-2-style ~50K vocabulary): a loss around 10.8 is what pure random guessing gets (ln(50257) ≈ 10.8), a loss around 5 is a small, lightly-trained model producing locally grammatical but incoherent text, a loss around 3.5 is roughly GPT-2 small (117M params) properly trained, and current frontier models run somewhere around 1.5-2.5. Perplexity (e^loss) converts this into "effective number of plausible next words the model is choosing between": a loss of 5.1 is a perplexity of about 164, a loss of 2.3 is a perplexity of about 10.
Checkpointing is not optional
Training runs take hours to weeks. Hardware or software will eventually fail mid-run. Saving the model weights, optimizer state, and current step count periodically means a crash costs you minutes, not the whole run. The optimizer state matters here, not just the weights: Adam-family optimizers carry per-parameter momentum terms, and resuming without them causes a visible bump in the loss curve right after resume as those statistics rebuild from scratch.
One trap: if your code that builds the optimizer's parameter groups ever iterates over an unordered collection (a Python set, for instance) instead of an ordered one (a list), the parameter order can come out different across runs, silently breaking checkpoint-resume, since the saved optimizer state no longer lines up with the right parameters. Use ordered containers throughout, and test resume-from-checkpoint at least once instead of assuming it works.
Where to look further
- Karpathy's nanoGPT
train.py: the training loop this whole pattern is descended from. - Hugging Face's
Trainer: a maintained, batteries-included implementation of the same loop, if you'd rather not hand-roll it. - This project's own from-scratch training loop and a real checkpoint-resume bug found along the way:
slm-from-scratchPhase 3 results.