Concepts

Hyperparameters

5 min readbeginnerUpdated 28 Sept 2026
1 · In one line

Hyperparameters are the settings people choose to control training, such as the learning rate, rather than values the model learns itself.

1 · What it is

A hyperparameter is a setting that controls how a model trains. Parameters are the weights and biases that training adjusts. So a network’s weights are parameters, while its learning rate is a hyperparameter. Common examples are the learning rate, the batch size and the number of epochs. The learning rate sets how big each weight update is. The batch size is how many examples go through the model between one update and the next. An epoch is one full pass over the training set. The shape of the model can be a hyperparameter too, such as the number of hidden layers in a network. So can the strength of regularisation, such as alpha in scikit-learn’s Lasso model. Stronger regularisation fights overfitting, but too much of it can blunt the model’s predictions.

Choosing these values is called tuning, and it happens across many training runs. A search needs a set of candidate values, a way to pick candidates, data held out for scoring, and a score. Grid search tries every combination of the listed values. Random search samples a fixed number of candidates instead. A 2012 study showed random search matching or beating grid search while spending only a small slice of the compute. The reason is that on a typical dataset most settings barely move the result, and the handful that do change from one dataset to the next. Bayesian optimisation treats the score as an unknown function of the settings and models it statistically. It then uses the results so far to choose what to try next. In a 2012 paper, it matched or beat expert tuning on models including convolutional neural networks. KerasTuner ships with Bayesian optimisation, Hyperband and random search built in. Optuna’s default sampler is the Tree-structured Parzen Estimator. Its samplers use the record of earlier trials to narrow the search.

In practice, the validation set gets checked round after round, and the test set waits until the very end. Tuning against the test set instead would let the model slowly fit that set’s quirks. So the settings are chosen on validation scores, and the test set is used at the end to double-check the winner. Even the validation set wears out, since every extra decision based on it makes its score a weaker guide to new data. Google Research’s deep learning tuning playbook prefers quasi-random search while exploring. Once exploration is over, it recommends Bayesian optimisation to find the final configuration.

2 · Why it exists

Training can learn a model's weights, but someone still has to choose how it trains.

Training needs settingsBefore a run starts, someone must fix values such as the learning rate, the batch size and the number of epochs.
Bad settings waste runsA learning rate set too low makes training crawl. One set too high can stop the model from ever settling.
No universal best valueNo single learning rate suits every case; the best one shifts with the model and the data. The right batch size also depends on the data and the computing power available.
3 · How it works

Follow one search from candidate settings to a single final test.

Illustrative numbers. Each candidate trains its own model, the validation set picks the winner, and the test set is used once at the end.
  1. 1 · proposeChoose a range for each setting, then list every combination as a grid or sample a fixed number at random.
  2. 2 · trainTrain one model per candidate on the training set, so each learns its own weights.
  3. 3 · compareScore every trained model on the validation set and keep the setting with the best score.
  4. 4 · testCheck the chosen model once on the test set, which played no part in the choice.

Pick settings on the validation set; keep the test set for one final check.

4 · Where it's used
WhoWhat they askWhat it works with
Student on a first project“Why does my loss jump around instead of falling?”The learning rate, lowered and retried
Data science team“Which regularisation strength gives the best churn model?”A grid of alpha values scored by cross-validation
Deep learning researchers“Is the new optimiser better, or just better tuned?”The learning rate, re-tuned for each optimiser before comparing
Machine learning platform team“How do we tune twenty models without watching each run?”Automated trials in a tool such as Optuna or KerasTuner
5 · What it solves, and what it doesn't
solves
  • Separates what a model learns from what people decide about how it learns.
  • Turns guesswork about settings into a repeatable search with a clear score.
  • Random search can reach settings as good as a full grid's for much less computing time.
  • Keeps the test set out of the search, so the final score still says something about new data.
doesn't solve
  • No setting is best everywhere, so every new model and dataset needs fresh tuning.
  • Tuning deep networks still involves a lot of toil and guesswork.
  • Heavy reuse of the validation set wears it out, so its scores grow less trustworthy.
  • Grid search wastes effort when only a few of its settings really matter.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsMachine Learning Glossary, Google for Developers · read 27 Sept 2026
  2. docsLinear regression: Hyperparameters (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
  3. docsDatasets: Dividing the original dataset (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
  4. docsTuning the hyper-parameters of an estimator, scikit-learn · read 27 Sept 2026
  5. paperRandom Search for Hyper-Parameter Optimization, Bergstra and Bengio, Journal of Machine Learning Research 2012 · read 27 Sept 2026
  6. paperPractical Bayesian Optimization of Machine Learning Algorithms, Snoek, Larochelle and Adams, NeurIPS 2012 · read 27 Sept 2026
  7. repoDeep Learning Tuning Playbook, Google Research · read 27 Sept 2026
  8. docsEfficient Optimization Algorithms, Optuna · read 27 Sept 2026
  9. docsKerasTuner, Keras · read 27 Sept 2026