A Deep Net (to be specific: a deep belief network which is a series of stacked RBMs, not Stacked Denoising AutoEncoders for clarification that there's a difference) usually can benefit from a moving window approach (slicing up an image in to chunks) to simulate a convolutional net. This can help a deep net generalize better.
That being said: even deep learning requires some sort of feature engineering at times (even if its pretty good with either hessian free training or pretraining).
The main thing with images is ensuring scaling them.
The trick with deep belief networks in particular is to make sure the RBMs have the right visible and hidden units (Hinton recommends Gaussian Visible, Rectified Linear Hidden).
That being said: even deep learning requires some sort of feature engineering at times (even if its pretty good with either hessian free training or pretraining).
The main thing with images is ensuring scaling them.
The trick with deep belief networks in particular is to make sure the RBMs have the right visible and hidden units (Hinton recommends Gaussian Visible, Rectified Linear Hidden).
Happy to answer other questions as well!