Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I'm of the mind that distillation is happening, but from what I've read, it's not some magical activity. It's basically adjusting the weights from fuzzy to more precise; It's not specifically a training method, and likely, it's not an ongoing need. Once they bootstrapped the models, they can try to do the same as the "SOTA" and generate their own synthetic data and try to curate that to completion.

Aside from it being a completely hollow cry for attention from US labs, its also just what these models are designed around: taking in data and forming a way for them to operate on some level of intelligence and context. It's like a dictionary maker getting made that someone used their complicated words, and another dictionary maker heard the word and wrote it down then checked what the first dictionary maker wrote about that word. It is not dictionary maker B reads dictionary A directly, as the labs want to imply.

 help



Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: