Concepts

Data parallelism

4 min readintermediateUpdated 28 Sept 2026
1 · In one line

Data parallelism runs replicas of one model on different slices of a batch, then combines their gradients before a synchronized update.

1 · What it is

Data parallelism replicates a model across processes. Every replica computes local gradients from a different set of input samples.

Those gradients are averaged inside the data-parallel communicator before each optimizer step. In synchronous training, workers train on different input slices in sync and aggregate gradients at each step.

DistributedDataParallel is recommended when the model fits on one GPU. When a model does not fit, PyTorch’s overview points to model-parallel or sharded data-parallel techniques instead.

2 · Why it exists

A model that fits on one device can train on more examples per step by dividing data work across devices.

Replicated modelEach worker owns a model replica with the same starting parameters.
Split batchWorkers process non-overlapping local slices of the global batch.
SynchronizationLocal gradients must be reduced before replicas apply the same update.
3 · How it works

Follow one synchronized data-parallel step.

Different data, replicated model, synchronized gradients.
  1. 1 · splitPartition the global batch so each replica receives different examples.
  2. 2 · computeEvery replica runs forward and backward on its local batch.
  3. 3 · reduceAn all-reduce combines local gradients across workers.
  4. 4 · updateEach replica applies the combined gradient to its copy of the model.

Data parallelism divides examples; it does not divide the ordinary DDP model weights.

4 · Where it's used
WhoWhat they askWhat it works with
Training engineer“How can a model that fits on one GPU use eight GPUs?”Examples per replica and all-reduce time
Research team“Can we shorten the time for a fixed number of training examples?”Step throughput and scaling efficiency
Platform operator“Why are devices waiting at the end of each step?”Input skew and gradient synchronization
5 · What it solves, and what it doesn't
solves
  • It processes different input slices simultaneously on multiple devices.
  • Synchronous gradient reduction keeps model replicas aligned after each update.
  • The pattern can scale from multiple devices on one machine to multiple workers.
doesn't solve
  • Ordinary replicated data parallelism does not make a model fit if one replica exceeds device memory.
  • Gradient communication adds overhead and can limit scaling.
  • Synchronous replicas merge their local gradients at the end of every step.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsPyTorch Distributed Overview, PyTorch · read 28 Sept 2026
  2. docsWhat is Distributed Data Parallel, PyTorch · read 28 Sept 2026
  3. docsDistributed training with TensorFlow, TensorFlow · read 28 Sept 2026
  4. docsMulti-GPU and distributed training, TensorFlow · read 28 Sept 2026
  5. docsGetting Started with Fully Sharded Data Parallel, PyTorch · read 28 Sept 2026