From the below Modal link they use: >continuous batching, so multiple generations can take place at the same time on a single container
>PagedAttention, which applies memory paging to the attention mechanism’s key-value cache, increasing throughput
[1]https://modal.com/docs/examples/text_generation_inference