Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

The problem is less of scheduling and more of communication. Typically all nodes will run (facilitated by a program like mpirun or mpiexec) the same executable which will be launched simultaneously in each OS. This is done inside a job scheduler and the resulting host list results in a defined communication pattern.

By the way, even systems completely across the room may be "close" to each other, depending on the topology of the system. You can imagine this being important for solving a physical system which is periodic in certain dimensions, so there should be little interconnect distance between physically-distant nodes.



At that scale, I find odd people didn't come up with an optimized OS that sees and manages the resources and also plan the communications between them.

It looks more like a private data center with 4000 dedicated machines in the same network to run distributed algorithms than a "single super computer". Are we just "wow-ing" at what is basically a data center here?


For HPC applications that actually need low-latency coordination between nodes, the application code itself manages communication. The communication can't be better optimized by the OS.

If you have an embarrassingly parallel problem, it will run well on this machine but it will also be a waste of the machine's expensive design. Embarrassingly parallel problems run just as well on generic data center hardware. This machine is built for problems that only parallelize effectively with low-latency coordination between nodes. Such problems come up a lot in scientific/engineering simulations but are comparatively rare in general purpose computing environments. General purpose nodes in a cloud computing environment cannot run some of the harder problems this machine runs, at any price. For any non-trivial parallel computing job there comes a crossover point where adding more nodes makes the total time-to-solution longer rather than shorter. This point comes a lot sooner if you don't have dedicated high-bandwidth, low-latency interconnects between nodes.


Precisely, if we are talking about low-latency, why would we let the communications go through the application (in user-land), to the kernel, then to the network stack, then be received by a kernel from another node, and then finally received again by the application in user-land. I would imagine as a first guess that bypassing several linux kernels and directly accessing remote hardwares would be mandatory for best low-latency.

If they are some info on internet about the software stack/architecture of the entire system, I would document myself on that. I didn't explore all the links I posted above yet.

I'm nowhere an expert, and HPC is really specific use case, but there is surely interesting bits to learn from it


One place to start might be http://infiniband.sourceforge.net/




Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: