Note that numpy normalizes the eigenvectors to be of length one,
whereas we took ours to be of arbitrary length.
Additionally, the choice of sign is arbitrary.
However, the vectors computed are parallel
to the ones we found by hand with the same eigenvalues.
Decomposing Matrices
Let's continue the previous example one step further. Let
\mathbf{W} = \begin{bmatrix}
1 & 1 \\
-1 & 2
\end{bmatrix},
be the matrix where the columns are the eigenvectors of the matrix \mathbf{A}. Let
\boldsymbol{\Sigma} = \begin{bmatrix}
1 & 0 \\
0 & 4
\end{bmatrix},
be the matrix with the associated eigenvalues on the diagonal.
Then the definition of eigenvalues and eigenvectors tells us that
\mathbf{A}\mathbf{W} =\mathbf{W} \boldsymbol{\Sigma} .
The matrix W is invertible, so we may multiply both sides by W^{-1} on the right,
we see that we may write
\mathbf{A} = \mathbf{W} \boldsymbol{\Sigma} \mathbf{W}^{-1}.
:eqlabel:eq_eig_decomp
In the next section we will see some nice consequences of this,
but for now we need only know that such a decomposition
will exist as long as we can find a full collection
of linearly independent eigenvectors (so that W is invertible).
Operations on Eigendecompositions
One nice thing about eigendecompositions :eqref:eq_eig_decomp is that
we can write many operations we usually encounter cleanly
in terms of the eigendecomposition. As a first example, consider:
\mathbf{A}^n = \overbrace{\mathbf{A}\cdots \mathbf{A}}^{\textrm{$n$ times}} = \overbrace{(\mathbf{W}\boldsymbol{\Sigma} \mathbf{W}^{-1})\cdots(\mathbf{W}\boldsymbol{\Sigma} \mathbf{W}^{-1})}^{\textrm{$n$ times}} = \mathbf{W}\overbrace{\boldsymbol{\Sigma}\cdots\boldsymbol{\Sigma}}^{\textrm{$n$ times}}\mathbf{W}^{-1} = \mathbf{W}\boldsymbol{\Sigma}^n \mathbf{W}^{-1}.
This tells us that for any positive power of a matrix,
the eigendecomposition is obtained by just raising the eigenvalues to the same power.
The same can be shown for negative powers,
so if we want to invert a matrix we need only consider
\mathbf{A}^{-1} = \mathbf{W}\boldsymbol{\Sigma}^{-1} \mathbf{W}^{-1},
or in other words, just invert each eigenvalue.
This will work as long as each eigenvalue is non-zero,
so we see that invertible is the same as having no zero eigenvalues.
Indeed, additional work can show that if \lambda_1, \ldots, \lambda_n
are the eigenvalues of a matrix, then the determinant of that matrix is
\det(\mathbf{A}) = \lambda_1 \cdots \lambda_n,
or the product of all the eigenvalues.
This makes sense intuitively because whatever stretching \mathbf{W} does,
W^{-1} undoes it, so in the end the only stretching that happens is
by multiplication by the diagonal matrix \boldsymbol{\Sigma},
which stretches volumes by the product of the diagonal elements.
Finally, recall that the rank was the maximum number
of linearly independent columns of your matrix.
By examining the eigendecomposition closely,
we can see that the rank is the same
as the number of non-zero eigenvalues of \mathbf{A}.
The examples could continue, but hopefully the point is clear:
eigendecomposition can simplify many linear-algebraic computations
and is a fundamental operation underlying many numerical algorithms
and much of the analysis that we do in linear algebra.
Eigendecompositions of Symmetric Matrices
It is not always possible to find enough linearly independent eigenvectors
for the above process to work. For instance the matrix
\mathbf{A} = \begin{bmatrix}
1 & 1 \\
0 & 1
\end{bmatrix},
has only a single eigenvector, namely (1, 0)^\top.
To handle such matrices, we require more advanced techniques
than we can cover (such as the Jordan Normal Form, or Singular Value Decomposition).
We will often need to restrict our attention to those matrices
where we can guarantee the existence of a full set of eigenvectors.
The most commonly encountered family are the symmetric matrices,
which are those matrices where \mathbf{A} = \mathbf{A}^\top.
In this case, we may take W to be an orthogonal matrix—a matrix whose columns are all length one vectors that are at right angles to one another, where
$\mathbf{W}^\top = \mathbf{W}^{-1}$—and all the eigenvalues will be real.
Thus, in this special case, we can write :eqref:eq_eig_decomp as
\mathbf{A} = \mathbf{W}\boldsymbol{\Sigma}\mathbf{W}^\top .
Gershgorin Circle Theorem
Eigenvalues are often difficult to reason with intuitively.
If presented an arbitrary matrix, there is little that can be said
about what the eigenvalues are without computing them.
There is, however, one theorem that can make it easy to approximate well
if the largest values are on the diagonal.
Let \mathbf{A} = (a_{ij}) be any square matrix (n\times n).
We will define r_i = \sum_{j \neq i} |a_{ij}|.
Let \mathcal{D}_i represent the disc in the complex plane
with center a_{ii} radius r_i.
Then, every eigenvalue of \mathbf{A} is contained in one of the \mathcal{D}_i.
This can be a bit to unpack, so let's look at an example.
Consider the matrix:
\mathbf{A} = \begin{bmatrix}
1.0 & 0.1 & 0.1 & 0.1 \\
0.1 & 3.0 & 0.2 & 0.3 \\
0.1 & 0.2 & 5.0 & 0.5 \\
0.1 & 0.3 & 0.5 & 9.0
\end{bmatrix}.
We have r_1 = 0.3, r_2 = 0.6, r_3 = 0.8 and r_4 = 0.9.
The matrix is symmetric, so all eigenvalues are real.
This means that all of our eigenvalues will be in one of the ranges of
[a_{11}-r_1, a_{11}+r_1] = [0.7, 1.3],
[a_{22}-r_2, a_{22}+r_2] = [2.4, 3.6],
[a_{33}-r_3, a_{33}+r_3] = [4.2, 5.8],
[a_{44}-r_4, a_{44}+r_4] = [8.1, 9.9].
Performing the numerical computation shows
that the eigenvalues are approximately 0.99, 2.97, 4.95, 9.08,
all comfortably inside the ranges provided.