U-Net
U-Net combines upsampled decoder features with matching high-resolution features from a contracting path.
U-Net is a convolutional network for dense image segmentation. Its contracting path repeatedly applies convolutions and pooling. Pooling reduces the feature-map resolution. The contracting path captures context. The expanding path increases output resolution.
The expanding path concatenates a correspondingly cropped feature map from the contracting path. At each resolution, a skip carries high-resolution features from the contracting side to the corresponding expanding stage. The decoder combines them with its upsampled features to assemble a more precise output.
The original system used paired images and masks. It used strong data augmentation. It used unpadded convolutions. Its official release included source code. The original paper describes overlap-tile segmentation. A 3D U-Net paper extended the architecture to volumetric segmentation learned from sparse annotations. Later Freiburg software supports 2D image data and 3D volumetric data. Attention U-Net added gates that learn to focus on target structures.
Image segmentation needs context and precise localization at the same time.
Trace an image down the contracting path and back up to a mask.
- 1 · contractMax pooling with stride two downsamples features along the contracting path.
- 2 · encodeThe contracting path captures context.
- 3 · copyHigh-resolution features cross to the matching expanding stage.
- 4 · expandUpsampling increases output resolution.
- 5 · classifyA pixel-wise classifier produces the segmentation map.
The U shape comes from a contracting path paired with a roughly symmetric expanding path.
| Who | What they ask | What it works with |
|---|---|---|
| Microscopy team | “Which pixels belong to each cell?” | A labelled cell mask |
| Medical-imaging team | “Where is the target structure in this scan?” | A dense anatomical mask |
| Volumetric-imaging team | “Can the same layout segment a 3D volume?” | A 3D U-Net with volumetric operations |
- The contracting path captures context while the expanding path supports precise localization.
- High-resolution features from the contracting path are combined with upsampled features.
- The original network can be trained end to end from very few images.
- U-Net still needs paired images and segmentation maps for supervised training.
- Unpadded convolutions make the original output smaller than its input by a fixed border.
- The original method uses an overlap-tile strategy for large images.
- Freiburg's 3D operations were added as an extension to its earlier 2D U-Net code.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperU-Net: Convolutional Networks for Biomedical Image Segmentation, Ronneberger, Fischer and Brox · read 28 Sept 2026
- officialU-Net project page, University of Freiburg · read 28 Sept 2026
- paper3D U-Net: Learning Dense Volumetric Segmentation from Sparse Annotation, Çiçek et al. · read 28 Sept 2026
- repoU-Net open-source software, University of Freiburg · read 28 Sept 2026
- paperAttention U-Net: Learning Where to Look for the Pancreas, Oktay et al. · read 28 Sept 2026