Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Thank you! Under good conditions it achieves approx a 67% cache hit rate with 16 expert slots


That's great, now I wonder how cache hit rate scales for larger models. Do you have any plans trying Qwen 3.6 or larger?


Check for colibri, dwarf star and flash-moe. they do similar things with bigger models

https://github.com/JustVugg/colibri https://github.com/antirez/ds4 https://github.com/danveloper/flash-moe




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: