Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

No, because the cache is local to the GPUs and there would be no cost/compute reason to have long-lived caches outside of the 1 hour cache already done by LLM APIs.


What if the real business isn't caching generic prompts/but caching outputs for a specific vertical where one can know two requests are truly equivalent n not just embedding similar?




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: