Concepts

Multilayer perceptron

3 min readbeginnerUpdated 28 Sept 2026
1 · In one line

A multilayer perceptron can fit a nonlinear model to training data.

1 · What it is

Nonlinear activations create complex mappings between model inputs and outputs. A dense layer computes an activation from an input-kernel dot product and a bias.

Hidden units can come to represent task features during training. A multiclass output can use softmax to return one probability per class.

Gradients are calculated using backpropagation. Large feedforward networks can overfit small training sets. Scikit-learn’s MLPClassifier supports only cross-entropy loss.

2 · Why it exists

Nonlinear activations create complex mappings between inputs and outputs.

Nonlinear patternsAn MLP can fit a nonlinear model to training data.
Learned featuresHidden units can come to represent task features during training.
Many outputsA multiclass output can use softmax to return one probability per class.
3 · How it works

Follow three features through a one-hidden-layer MLP.

Nonlinear activations create complex mappings between model inputs and outputs.
  1. 1 · receiveRead the input data.
  2. 2 · projectCompute an input-kernel dot product and add a bias.
  3. 3 · activateApply a nonlinear activation to the hidden values.
  4. 4 · outputApply softmax as the multiclass output function.
  5. 5 · learnBackpropagate the output error and adjust the connection weights.

Nonlinear activations create complex mappings between model inputs and outputs.

4 · Where it's used
WhoWhat they askWhat it works with
Risk team“Which class matches this tabular record?”Numeric and encoded categorical features
Forecasting team“What numeric value follows from these measurements?”A fixed-length feature vector
Research team“Does a nonlinear baseline improve on a linear classifier?”Labelled training examples
5 · What it solves, and what it doesn't
solves
  • An MLP can fit a nonlinear model to training data.
  • Gradients are calculated using backpropagation.
  • MLPClassifier applies softmax as its multiclass output function.
doesn't solve
  • Large feedforward networks can overfit small training sets.
  • Scikit-learn's MLPClassifier supports only cross-entropy loss.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsNeural network models (supervised), scikit-learn · read 28 Sept 2026
  2. docsMLPClassifier, scikit-learn · read 28 Sept 2026
  3. docsDense layer, Keras · read 28 Sept 2026
  4. docsBuild the Neural Network, PyTorch · read 28 Sept 2026
  5. paperLearning representations by back-propagating errors, Rumelhart, Hinton and Williams, Nature · read 28 Sept 2026
  6. paperImproving neural networks by preventing co-adaptation of feature detectors, Hinton et al. · read 28 Sept 2026