Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Haven't read the full report yet, just a quick question. Are your numbers for cold start without pre fill or is it after warmed cache?


We do graph capture etc at startup (same as vLLM) but this model variant doesn’t require prefix caching - the prefix is just 10 tokens.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: