We show that spatial transformations of CNN feature maps cannot align the feature maps of a transformed image to match those of it’s original for general affine transformations. This implies that methods that spatially transform CNN feature maps, such as spatial transformer networks, dilated or deformable convolutions or spatial pyramid pooling cannot enable true invariance. Our proof is based on elementary analysis for both the single- and multi-layer network cases.
QC 20201224