Lazy Initialization
🏷️sec_lazy_init
So far, it might seem that we got away
with being sloppy in setting up our networks.
Specifically, we did the following unintuitive things,
which might not seem like they should work:
- We defined the network architectures
without specifying the input dimensionality.
- We added layers without specifying
the output dimension of the previous layer.
- We even "initialized" these parameters
before providing enough information to determine
how many parameters our models should contain.
You might be surprised that our code runs at all.
After all, there is no way the deep learning framework
could tell what the input dimensionality of a network would be.
The trick here is that the framework defers initialization,
waiting until the first time we pass data through the model,
to infer the sizes of each layer on the fly.
Later on, when working with convolutional neural networks,
this technique will become even more convenient
since the input dimensionality
(e.g., the resolution of an image)
will affect the dimensionality
of each subsequent layer.
Hence the ability to set parameters
without the need to know,
at the time of writing the code,
the value of the dimension
can greatly simplify the task of specifying
and subsequently modifying our models.
Next, we go deeper into the mechanics of initialization.