I think a bigger issue is that retraining isn't deterministic (ie bitwise identical) across different hardware, IIRC due to differing float behaviour. I guess folks interested in ML reproducible builds will have to ignore small differences in model float values, but I wonder what affect those differences on model outputs.