What Is Hyperparameter Optimization?
🏷️sec_what_is_hpo
As we have seen in the previous chapters, deep neural networks come with a
large number of parameters or weights that are learned during training. On
top of these, every neural network has additional hyperparameters that need
to be configured by the user. For example, to ensure that stochastic gradient
descent converges to a local optimum of the training loss
(see :numref:chap_optimization), we have to adjust the learning rate and batch
size. To avoid overfitting on training datasets,
we might have to set regularization parameters, such as weight decay
(see :numref:sec_weight_decay) or dropout (see :numref:sec_dropout). We can
define the capacity and inductive bias of the model by setting the number of
layers and number of units or filters per layer (i.e., the effective number
of weights).
Unfortunately, we cannot simply adjust these hyperparameters by minimizing the
training loss, because this would lead to overfitting on the training data. For
example, setting regularization parameters, such as dropout or weight decay
to zero leads to a small training loss, but might hurt the generalization
performance.
🏷️ml_workflow
Without a different form of automation, hyperparameters have to be set manually
in a trial-and-error fashion, in what amounts to a time-consuming and difficult
part of machine learning workflows. For example, consider training
a ResNet (see :numref:sec_resnet) on CIFAR-10, which requires more than 2 hours
on an Amazon Elastic Cloud Compute (EC2) g4dn.xlarge instance. Even just
trying ten hyperparameter configurations in sequence, this would already take us
roughly one day. To make matters worse, hyperparameters are usually not directly
transferable across architectures and datasets
:cite:feurer-arxiv22,wistuba-ml18,bardenet-icml13a, and need to be re-optimized
for every new task. Also, for most hyperparameters, there are no rule-of-thumbs,
and expert knowledge is required to find sensible values.
Hyperparameter optimization (HPO) algorithms are designed to tackle this
problem in a principled and automated fashion :cite:feurer-automlbook18a, by
framing it as a global optimization problem. The default objective is the error
on a hold-out validation dataset, but could in principle be any other business
metric. It can be combined with or constrained by secondary objectives, such as
training time, inference time, or model complexity.
Recently, hyperparameter optimization has been extended to neural architecture
search (NAS) :cite:elsken-arxiv18a,wistuba-arxiv19, where the goal is to find
entirely new neural network architectures. Compared to classical HPO, NAS is even
more expensive in terms of computation and requires additional efforts to remain
feasible in practice. Both, HPO and NAS can be considered as sub-fields of
AutoML :cite:hutter-book19a, which aims to automate the entire ML pipeline.
In this section we will introduce HPO and show how we can automatically find
the best hyperparameters of the logistic regression example introduced in
:numref:sec_softmax_concise.
The Optimization Problem
🏷️sec_definition_hpo
We will start with a simple toy problem: searching for the learning rate of the
multi-class logistic regression model SoftmaxRegression from
:numref:sec_softmax_concise to minimize the validation error on the Fashion
MNIST dataset. While other hyperparameters like batch size or number of epochs
are also worth tuning, we focus on learning rate alone for simplicity.