Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Since the current NVL72 are still at the ~15%/yr failure rate it's not clear your new data center is going to have half it's compute in 3 years. If you're still running H100s they draw >10x kWh/Mtoken as new designs. All of these systems become dated, but not all of them require entirely new infrastructure.

If a ROM rack running a near frontier agent model at >10ktoken/sec costs <$1M (rather than $4-8M for NVL72s) and draws only 10-20kW (rather than 100-200kW), and doesn't require a completely new cooling and power system every time you update? There'll be lots of demand for GLM5.3 in a year.

What these don't do is TRAINING, they only do INFERENCE, but they could do it pretty well.



> Since the current NVL72 are still at the ~15%/yr failure rate it's not clear your new data center is going to have half it's compute in 3 years

This would be a stupidly bad failure rate, basically the worst business decision you could make, especially if you're somehow on the hook for eating those losses (which seems to be the implication?). Is there a linkable source on this?

The only thing I could find is SemiAnalysis claims that 15% of Blackwells end up RMAed[1]. That appears to be a total failure rate, though, and if you're RMAing them, you're getting replacements. So that appears to be a pretty different state of things.

[1] https://www.dwarkesh.com/p/dylan-patel#:~:text=GPUs%20are%20...


I'll note CoreWeave didn't commission the very first Blackwells until ~Jan 2025 so it hasn't been very long for RMAs. Furthermore, for training even when the NVL backplane can use the remaining GPUs a ~20% drop in performance means it isn't in the training cluster.

I've personally heard this from several sources in the data centers (installers, training, network). It's not uncommon for 10% of racks to fail on delivery. I hear that's improved somewhat from GB200 to GB300, but the number of FW updates from the time they ship, until they're commissioned is >>10. If an HBM or GPU or backplane supply/cooling fails, it is basically not swappable or repairable. You have a "dead" rack, and deliveries are on allocation so you don't get a replacement for months (eg some "RMAs" for early delivered parts in late 2025 are still dead racks 9 months later). "Tray" swaps are technically possible, but still quite rare, perhaps because debugging takes as much time as commissioning a new rack.

I don't want to out anyone, but these are similar comments:

https://www.linkedin.com/posts/neelmaster1_aiinfrastructure-...

https://www.hostzealot.com/blog/news/nvidia-gb200-nvl72-is-n...

https://introl.com/blog/gb200-nvl72-deployment-72-gpu-liquid...


They fail that often? Damn.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: