Adadelta
🏷️sec_adadelta
Adadelta is yet another variant of AdaGrad (:numref:sec_adagrad). The main difference lies in the fact that it decreases the amount by which the learning rate is adaptive to coordinates. Moreover, traditionally it referred to as not having a learning rate since it uses the amount of change itself as calibration for future change. The algorithm was proposed in :citet:Zeiler.2012. It is fairly straightforward, given the discussion of previous algorithms so far.
The Algorithm
In a nutshell, Adadelta uses two state variables, \mathbf{s}_t to store a leaky average of the second moment of the gradient and \Delta\mathbf{x}_t to store a leaky average of the second moment of the change of parameters in the model itself. Note that we use the original notation and naming of the authors for compatibility with other publications and implementations (there is no other real reason why one should use different Greek variables to indicate a parameter serving the same purpose in momentum, Adagrad, RMSProp, and Adadelta).
Here are the technical details of Adadelta. Given the parameter du jour is \rho, we obtain the following leaky updates similarly to :numref:sec_rmsprop:
$$\begin{aligned}
\mathbf{s}t & = \rho \mathbf{s}{t-1} + (1 - \rho) \mathbf{g}_t^2.
\end{aligned}$$
The difference to :numref:sec_rmsprop is that we perform updates with the rescaled gradient \mathbf{g}_t', i.e.,
$$\begin{aligned}
\mathbf{x}t & = \mathbf{x}{t-1} - \mathbf{g}_t'. \
\end{aligned}$$
So what is the rescaled gradient \mathbf{g}_t'? We can calculate it as follows:
$$\begin{aligned}
\mathbf{g}t' & = \frac{\sqrt{\Delta\mathbf{x}{t-1} + \epsilon}}{\sqrt{{\mathbf{s}_t + \epsilon}}} \odot \mathbf{g}_t, \
\end{aligned}$$
where \Delta \mathbf{x}_{t-1} is the leaky average of the squared rescaled gradients \mathbf{g}_t'. We initialize \Delta \mathbf{x}_{0} to be 0 and update it at each step with \mathbf{g}_t', i.e.,
$$\begin{aligned}
\Delta \mathbf{x}t & = \rho \Delta\mathbf{x}{t-1} + (1 - \rho) {\mathbf{g}_t'}^2,
\end{aligned}$$
and \epsilon (a small value such as 10^{-5}) is added to maintain numerical stability.
Implementation
Adadelta needs to maintain two state variables for each variable, \mathbf{s}_t and \Delta\mathbf{x}_t. This yields the following implementation.