Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I think a number of us are using the method and these folks decided to do a paper and release their findings.


Lots of speedups already on Llama3 like Groq (hardware arch) and Modal [1] getting 200 t/s

From the below Modal link they use: >continuous batching, so multiple generations can take place at the same time on a single container

>PagedAttention, which applies memory paging to the attention mechanism’s key-value cache, increasing throughput

[1]https://modal.com/docs/examples/text_generation_inference




Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: