Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Maybe switching effort routes you to a different rack of gpu’s which don’t have the cache


The KV cache is probably offloaded to RAM after a generation is complete. Then it is pulled back into any rack in the data center that has the model you’re using.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: