Understanding Batch Size in Deep Learning
1. What Is Batch Size?
Batch size is the number of samples used each time the model parameters are updated during training.
An intuitive analogy with studying:
- Large batch = doing many practice problems at once (fast but coarse)
- Small batch = doing one problem at a time (slow but precise)
- Use a small per-device batch size, performing multiple forward and backward passes
- Accumulate gradients from several small batches
- Update the model parameters once at the end
- 💰 Pay 32 units at once
- 💾 Requires a lot of cash (GPU memory)
- 💸 Pay in 4 installments of 8 units each
- 💾 Needs only a small amount of cash (GPU memory)
- ✅ Same total amount paid in the end
- Start with
batch_size = 1or2 - Increase the effective batch size via gradient accumulation
- Pair larger batch sizes with larger learning rates
- Use gradient accumulation to simulate large batches
- Balance memory usage against training effectiveness
- Adjust flexibly according to your hardware
Key Formula
Effective batch size = per-device batch size × gradient accumulation steps
2. Large vs. Small Batch Sizes
Advantages of Large Batches
| Advantage | Explanation | |-----------|-------------| | 🎯 More stable training | More accurate gradient estimates, smoother convergence | | 🏆 Easier convergence to global optimum | Lower risk of getting stuck in local optima | | ⚡ Higher compute efficiency | Fully utilizes GPU parallel computing |
Advantages of Small Batches
| Advantage | Explanation | |-----------|-------------| | 💾 Saves GPU memory | Suitable for memory-constrained setups | | 🔍 Captures data details | Gradient noise helps escape local optima | | 🌟 Stronger generalization | Reduces overfitting risk |
3. Gradient Accumulation
Gradient accumulation works like an "installment payment" scheme:
The "Installment Payment" Analogy
Direct large batch:
Gradient accumulation:
4. Parameter Setting Tips
Small Models / Small Datasets
Memory Optimization
*Source: Easy AI Tutorial Series*