Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

> This could make video games take up so much less space and have much more robust speech, especially from NPCs.

Maybe, maybe not. You'll see some of the model sizes I posted in comments above. These are quite large, and adding models for multiple speakers gets quite large. These have to live in memory and probably can't be paged in selectively.

Once we achieve high fidelity multi-speaker embedding models (where multiple speakers are encoded in a singular model), then we'll have something compelling. I imagine the models will become less dense over time as well.

Furthermore, if the models are deterministic, then the designers will know what each line will sound like exactly before it's produced.



Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: