Diffusion Gemma sucks to train if you are not already mostly in distribution. The reason diffusion Gemma is fast is a learned denoising that balances quality and speed. So the further your data is from what already happens, the slower it gets. In almost every case I’ve tried, you might as well just train Nemotron.