Apparently op cares a lot about speed, which is fine, but ML researchers care about correctness first, speed second. And it makes sense, because they are not as resource constrained as OP.
Most PyTorch tensor operations are cython not python. So imo rewriting in rust is not going to have an enormous speed up. If that really was the concern we should see a throughput comparison vs PyTorch or something.
Are you trying to make an argument here that speed is more important than correctness? I'm finding it difficult to interpret - the purpose of - this comment otherwise, and if you are, I'd consider such an argument pretty wild.