Data parallelism
Data parallelism runs replicas of one model on different slices of a batch, then combines their gradients before a synchronized update.
Data parallelism replicates a model across processes. Every replica computes local gradients from a different set of input samples.
Those gradients are averaged inside the data-parallel communicator before each optimizer step. In synchronous training, workers train on different input slices in sync and aggregate gradients at each step.
DistributedDataParallel is recommended when the model fits on one GPU. When a model does not fit, PyTorch’s overview points to model-parallel or sharded data-parallel techniques instead.
A model that fits on one device can train on more examples per step by dividing data work across devices.
Follow one synchronized data-parallel step.
- 1 · splitPartition the global batch so each replica receives different examples.
- 2 · computeEvery replica runs forward and backward on its local batch.
- 3 · reduceAn all-reduce combines local gradients across workers.
- 4 · updateEach replica applies the combined gradient to its copy of the model.
Data parallelism divides examples; it does not divide the ordinary DDP model weights.
| Who | What they ask | What it works with |
|---|---|---|
| Training engineer | “How can a model that fits on one GPU use eight GPUs?” | Examples per replica and all-reduce time |
| Research team | “Can we shorten the time for a fixed number of training examples?” | Step throughput and scaling efficiency |
| Platform operator | “Why are devices waiting at the end of each step?” | Input skew and gradient synchronization |
- It processes different input slices simultaneously on multiple devices.
- Synchronous gradient reduction keeps model replicas aligned after each update.
- The pattern can scale from multiple devices on one machine to multiple workers.
- Ordinary replicated data parallelism does not make a model fit if one replica exceeds device memory.
- Gradient communication adds overhead and can limit scaling.
- Synchronous replicas merge their local gradients at the end of every step.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsPyTorch Distributed Overview, PyTorch · read 28 Sept 2026
- docsWhat is Distributed Data Parallel, PyTorch · read 28 Sept 2026
- docsDistributed training with TensorFlow, TensorFlow · read 28 Sept 2026
- docsMulti-GPU and distributed training, TensorFlow · read 28 Sept 2026
- docsGetting Started with Fully Sharded Data Parallel, PyTorch · read 28 Sept 2026