Cookbook · GPUProgramming
Training frameworks and efficiency techniques
3 mintraining-frameworksdistributed-traininggradient-checkpointingmixed-precision
Picking a framework
Once you're not trying to learn the training loop itself (see pretraining), hand-rolling it stops being the point. Common options, roughly in order of how much they abstract away:
- Raw PyTorch: full control, most learning value, most code to maintain yourself.
- Hugging Face
Trainer: a maintained training loop covering most standard cases (mixed precision, checkpointing, logging, distributed training) with a config-driven interface. - Unsloth: custom Triton kernels (fused RoPE, MLP, attention) and smart sequence packing that can make fine-tuning (LoRA/QLoRA especially) noticeably faster and lower-memory on consumer GPUs. Built on top of the Hugging Face ecosystem, not a replacement for it.
- axolotl: a YAML-config-driven wrapper around the Hugging Face stack, popular for fine-tuning runs without writing training code at all.
Efficiency techniques that apply regardless of framework
Gradient checkpointing trades compute for memory: instead of storing every intermediate activation from the forward pass for use in the backward pass, it stores only a subset and recomputes the rest when needed. This can cut activation memory substantially at the cost of a slower backward pass (often 20-30% slower), and is usually the first thing to enable when a training run doesn't fit in memory.
Mixed precision (see pretraining's knobs section) does the forward/backward math in bf16 or fp16 instead of fp32, roughly halving memory for activations and increasing throughput on hardware with dedicated low-precision compute units, while keeping a higher-precision copy of weights for the actual optimizer update.
Distributed training splits training across multiple GPUs (or machines) when a single device isn't enough:
- Data parallelism: every device holds a full copy of the model and processes a different slice of the batch, then gradients are synchronized across devices. Simplest to reason about; doesn't help if the model itself doesn't fit on one device.
- Tensor parallelism: individual weight matrices are split across devices, so a single layer's computation is itself distributed. Needed when a model is too large for one device's memory even at batch size 1.
- Pipeline parallelism: different layers of the model live on different devices, and micro-batches flow through them like an assembly line. Reduces per-device memory at the cost of some idle time ("bubble") while later stages wait for earlier ones.
Large training runs typically combine several of these at once (data + tensor + pipeline parallelism together), which is what frameworks like DeepSpeed and Megatron-LM are built to orchestrate.
Where to look further
- DeepSpeed: Microsoft's distributed training library, including ZeRO (a memory-optimization technique for data parallelism that shards optimizer state, gradients, and even parameters across devices).
- Megatron-LM: NVIDIA's framework for large-scale tensor and pipeline parallelism.
- Pretraining: the training loop these techniques all optimize around.