Multiscale Detection
Since we have generated multiscale anchor boxes,
we will use them to detect objects of various sizes
at different scales.
In the following
we introduce a CNN-based multiscale object detection
method that we will implement
in :numref:sec_ssd.
At some scale,
say that we have c feature maps of shape h \times w.
Using the method in :numref:subsec_multiscale-anchor-boxes,
we generate hw sets of anchor boxes,
where each set has a anchor boxes with the same center.
For example,
at the first scale in the experiments in :numref:subsec_multiscale-anchor-boxes,
given ten (number of channels) 4 \times 4 feature maps,
we generated 16 sets of anchor boxes,
where each set contains 3 anchor boxes with the same center.
Next, each anchor box is labeled with
the class and offset based on ground-truth bounding boxes. At the current scale, the object detection model needs to predict the classes and offsets of hw sets of anchor boxes on the input image, where different sets have different centers.
Assume that the c feature maps here
are the intermediate outputs obtained
by the CNN forward propagation based on the input image. Since there are hw different spatial positions on each feature map,
the same spatial position can be
thought of as having c units.
According to the
definition of receptive field in :numref:sec_conv_layer,
these c units at the same spatial position
of the feature maps
have the same receptive field on the input image:
they represent the input image information
in the same receptive field.
Therefore, we can transform the c units
of the feature maps at the same spatial position
into the
classes and offsets of the a anchor boxes
generated using this spatial position.
In essence,
we use the information of the input image in a certain receptive field
to predict the classes and offsets of the anchor boxes
that are
close to that receptive field
on the input image.
When the feature maps at different layers
have varying-size receptive fields on the input image, they can be used to detect objects of different sizes.
For example, we can design a neural network where
units of feature maps that are closer to the output layer
have wider receptive fields,
so they can detect larger objects from the input image.
In a nutshell, we can leverage
layerwise representations of images at multiple levels
by deep neural networks
for multiscale object detection.
We will show how this works through a concrete example
in :numref:sec_ssd.
Summary
- At multiple scales, we can generate anchor boxes with different sizes to detect objects with different sizes.
- By defining the shape of feature maps, we can determine centers of uniformly sampled anchor boxes on any image.
- We use the information of the input image in a certain receptive field to predict the classes and offsets of the anchor boxes that are close to that receptive field on the input image.
- Through deep learning, we can leverage its layerwise representations of images at multiple levels for multiscale object detection.
Exercises
- According to our discussions in :numref:
sec_alexnet, deep neural networks learn hierarchical features with increasing levels of abstraction for images. In multiscale object detection, do feature maps at different scales correspond to different levels of abstraction? Why or why not?
- At the first scale (
fmap_w=4, fmap_h=4) in the experiments in :numref:subsec_multiscale-anchor-boxes, generate uniformly distributed anchor boxes that may overlap.
- Given a feature map variable with shape
1 \times c \times h \times w, where c, h, and w are the number of channels, height, and width of the feature maps, respectively. How can you transform this variable into the classes and offsets of anchor boxes? What is the shape of the output?