Partial Derivatives and Gradients
🏷️subsec_calculus-grad
Thus far, we have been differentiating
functions of just one variable.
In deep learning, we also need to work
with functions of many variables.
We briefly introduce notions of the derivative
that apply to such multivariate functions.
Let y = f(x_1, x_2, \ldots, x_n) be a function with n variables.
The partial derivative of y
with respect to its i^\textrm{th} parameter x_i is
\frac{\partial y}{\partial x_i} = \lim_{h \rightarrow 0} \frac{f(x_1, \ldots, x_{i-1}, x_i+h, x_{i+1}, \ldots, x_n) - f(x_1, \ldots, x_i, \ldots, x_n)}{h}.
To calculate \frac{\partial y}{\partial x_i},
we can treat x_1, \ldots, x_{i-1}, x_{i+1}, \ldots, x_n as constants
and calculate the derivative of y with respect to x_i.
The following notational conventions for partial derivatives
are all common and all mean the same thing:
\frac{\partial y}{\partial x_i} = \frac{\partial f}{\partial x_i} = \partial_{x_i} f = \partial_i f = f_{x_i} = f_i = D_i f = D_{x_i} f.
We can concatenate partial derivatives
of a multivariate function
with respect to all its variables
to obtain a vector that is called
the gradient of the function.
Suppose that the input of function
f: \mathbb{R}^n \rightarrow \mathbb{R}
is an $n$-dimensional vector
\mathbf{x} = [x_1, x_2, \ldots, x_n]^\top
and the output is a scalar.
The gradient of the function f
with respect to \mathbf{x}
is a vector of n partial derivatives:
$$\nabla_{\mathbf{x}} f(\mathbf{x}) = \left[\partial_{x_1} f(\mathbf{x}), \partial_{x_2} f(\mathbf{x}), \ldots
\partial_{x_n} f(\mathbf{x})\right]^\top.$$
When there is no ambiguity,
\nabla_{\mathbf{x}} f(\mathbf{x})
is typically replaced
by \nabla f(\mathbf{x}).
The following rules come in handy
for differentiating multivariate functions:
- For all
\mathbf{A} \in \mathbb{R}^{m \times n} we have \nabla_{\mathbf{x}} \mathbf{A} \mathbf{x} = \mathbf{A}^\top and \nabla_{\mathbf{x}} \mathbf{x}^\top \mathbf{A} = \mathbf{A}.
- For square matrices
\mathbf{A} \in \mathbb{R}^{n \times n} we have that \nabla_{\mathbf{x}} \mathbf{x}^\top \mathbf{A} \mathbf{x} = (\mathbf{A} + \mathbf{A}^\top)\mathbf{x} and in particular
\nabla_{\mathbf{x}} \|\mathbf{x} \|^2 = \nabla_{\mathbf{x}} \mathbf{x}^\top \mathbf{x} = 2\mathbf{x}.
Similarly, for any matrix \mathbf{X},
we have \nabla_{\mathbf{X}} \|\mathbf{X} \|_\textrm{F}^2 = 2\mathbf{X}.
Chain Rule
In deep learning, the gradients of concern
are often difficult to calculate
because we are working with
deeply nested functions
(of functions (of functions...)).
Fortunately, the chain rule takes care of this.
Returning to functions of a single variable,
suppose that y = f(g(x))
and that the underlying functions
y=f(u) and u=g(x)
are both differentiable.
The chain rule states that
\frac{dy}{dx} = \frac{dy}{du} \frac{du}{dx}.
Turning back to multivariate functions,
suppose that y = f(\mathbf{u}) has variables
u_1, u_2, \ldots, u_m,
where each u_i = g_i(\mathbf{x})
has variables x_1, x_2, \ldots, x_n,
i.e., \mathbf{u} = g(\mathbf{x}).
Then the chain rule states that
\frac{\partial y}{\partial x_{i}} = \frac{\partial y}{\partial u_{1}} \frac{\partial u_{1}}{\partial x_{i}} + \frac{\partial y}{\partial u_{2}} \frac{\partial u_{2}}{\partial x_{i}} + \ldots + \frac{\partial y}{\partial u_{m}} \frac{\partial u_{m}}{\partial x_{i}} \ \textrm{ and so } \ \nabla_{\mathbf{x}} y = \mathbf{A} \nabla_{\mathbf{u}} y,
where \mathbf{A} \in \mathbb{R}^{n \times m} is a matrix
that contains the derivative of vector \mathbf{u}
with respect to vector \mathbf{x}.
Thus, evaluating the gradient requires
computing a vector--matrix product.
This is one of the key reasons why linear algebra
is such an integral building block
in building deep learning systems.
Discussion
While we have just scratched the surface of a deep topic,
a number of concepts already come into focus:
first, the composition rules for differentiation
can be applied routinely, enabling
us to compute gradients automatically.
This task requires no creativity and thus
we can focus our cognitive powers elsewhere.
Second, computing the derivatives of vector-valued functions
requires us to multiply matrices as we trace
the dependency graph of variables from output to input.
In particular, this graph is traversed in a forward direction
when we evaluate a function
and in a backwards direction
when we compute gradients.
Later chapters will formally introduce backpropagation,
a computational procedure for applying the chain rule.
From the viewpoint of optimization, gradients allow us
to determine how to move the parameters of a model
in order to lower the loss,
and each step of the optimization algorithms used
throughout this book will require calculating the gradient.
Exercises
- So far we took the rules for derivatives for granted.
Using the definition and limits prove the properties
for (i)
f(x) = c, (ii) f(x) = x^n, (iii) f(x) = e^x and (iv) f(x) = \log x.
- In the same vein, prove the product, sum, and quotient rule from first principles.
- Prove that the constant multiple rule follows as a special case of the product rule.
- Calculate the derivative of
f(x) = x^x.
- What does it mean that
f'(x) = 0 for some x?
Give an example of a function f
and a location x for which this might hold.
- Plot the function
y = f(x) = x^3 - \frac{1}{x}
and plot its tangent line at x = 1.
- Find the gradient of the function
f(\mathbf{x}) = 3x_1^2 + 5e^{x_2}.
- What is the gradient of the function
f(\mathbf{x}) = \|\mathbf{x}\|_2? What happens for \mathbf{x} = \mathbf{0}?
- Can you write out the chain rule for the case
where
u = f(x, y, z) and x = x(a, b), y = y(a, b), and z = z(a, b)?
- Given a function
f(x) that is invertible,
compute the derivative of its inverse f^{-1}(x).
Here we have that f^{-1}(f(x)) = x and conversely f(f^{-1}(y)) = y.
Hint: use these properties in your derivation.