Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

You cannot infer this because they only show the tokens per second per user. One way to get a higher number is to have fewer users per chip.

I'm pretty sure Cerebras has a confidentiality agreement with OpenAI, and this press release was carefully constructed to avoid leaking details about the model weights. For example, the graph of tokens per second vs. tokens per second per user doesn't have any numbers that would allow you to translate between the two. (And in any case the relationship depends on the model.)



They show that CS-4 can't really do batching (or rather it can't properly benefit from it), total throughput barely changes (25%?): https://cdn.sanity.io/images/e4qjo92p/production/6a132331880...

Which I think makes it feasible to approximate activation from CS-4 tokens per second per user.




Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: