I think that convolutional networks are really mostly a way to address space shift invariance too; and I even wonder if the same job could be done by some special purpose code that just runs the underlying neural net with a bunch of different rotations of the same image. That's probably how they do it now.. I feel like that's probably close to the optimal approach.