Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

> On my 128GB M4 Max Mac Studio, generating a 15s 480p video with MiniMax H3 in ComfyUI takes an hour and a half.

That's crazy, a RTX Pro 6000 does that in in 2-3 minutes (give or take, depending on your exact settings). LLMs don't make the difference between standalone GPU vs unified memory + CPU so obvious as diffusion models seems to do.



It’s always been the case, it’s more the anomaly that LLMs work at comparable speeds on M series because almost all other ML runs way faster on Nvidia cards.


LLM prompt processing and diffusion models are compute bound, while LLM token generation is memory bandwidth bound.


An RTX6000 is a completely different class of hardware.


Really? No wonder I keep trying to type on it like a laptop but it doesn't work and doesn't even have a display!


Right? Do you also type on a Mac Studio without a keyboard plugged in? Like tap the ethernet port 3 times in a row then this sequence of sticking your fingers into TB5 ports? I mean, it’s clear you stick your RTX into a computer.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: