Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

OP, you should not have written this in Rust. It should be in PyTorch, which is by far the most popular. We can’t tell if this architecture is good or whether there is a problem in your implementation.

You can test the whole thing for free on a GPU with Google Colab. Test both the transformer and your new architecture on a larger dataset. Something that maxes out the GPU for an hour each run.

Also, the readme mentions keeping the same optimizer schedule which sounds nice at first but they are completely different architectures. The loss is high on the transformer, did you try raising the learning rate on it?

In general I’m interested in parameter efficient architectures. I don’t think transformers are optimal, and indeed many improvements have been made to vanilla transformers. But if you have an idea for something better you need to show it.

 help



I dunno, I could probably be convinced to try a new tool purely on the basis of not having to deal with installing pytorch

Yeah likewise. Pytorch needs to die.

That said I don't know what's wrong with using a Rust AI framework like Candle.


I’m curious, what’s the criticism for PyTorch?

One criticism is that you have to install the same package, torch, but from different Python indexes in order to install the cpu version or gpu version, on Linux. On windows, `pip install torch` gets you the cpu version. On linux, that gets you a ton of Nvidia extras that take a lot of space.

GPU support should really be a optional extra eg `torch[gpu]` or `torch[nvidia]`.


I would argue that we need a `torch[stubs]` package (or similar) that doesn't install any form of libtorch. The CPU version of libtorch is 160MB compressed and 400MB extracted.

I don't think it's caused by PyTorch on its own, but every AI-related Python project I try out locally manages to depend on a version of PyTorch that isn't in my disk cache yet. Having to download a gigabyte of dependencies for every project gets tiresome.

The Rust compile cycle will probably generate a gigabyte of files locally as well, but at least they can be `rm`'d out of `target/` once it's done.

It should be said that for this project that's entirely irrelevant of course, but seeing PyTorch has made me skip over projects on the HN homepage before and probably will again in the future.


The entire Python ecosystem is horrible and shouldn’t have been as falsely boosted by institutions as it was in the 2010s. Yes, it got less bad with 3.8 or whatever version added type annotations. But making so much of ML depend on Python has made it distasteful to a lot of devs who would otherwise have contributed more to it.

We really need to move all AI/ML research off PyTorch to Candle or… just anything that isn’t Python or another old-gen, broken language like it.


Apparently op cares a lot about speed, which is fine, but ML researchers care about correctness first, speed second. And it makes sense, because they are not as resource constrained as OP.

Most PyTorch tensor operations are cython not python. So imo rewriting in rust is not going to have an enormous speed up. If that really was the concern we should see a throughput comparison vs PyTorch or something.

Are you trying to make an argument here that speed is more important than correctness? I'm finding it difficult to interpret - the purpose of - this comment otherwise, and if you are, I'd consider such an argument pretty wild.

They aren't using Rust for speed.

yeah that is why they compute in fp32 lol



Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: