Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Can you explain the difference?


Distillation originally meant matching the distribution of the student model to the teacher model using something like a KL divergence.

When you instead fine-tune the student on the samples from the teacher, which is what people mean by distillation today, you are in effect doing a monte-carlo version of the same thing. While in theory this is higher variance, given modern setups where the student and teacher are both large and are RLd heavily (leading to a sharp teacher distribution), and given that you typically use lots and lots of samples, it ends up OK.


Fine-tuning is done against a dataset. Distillation is done against a model.


Distillation is used to build part of a data set for fine-tuning (loosely interpreted). Advanced model traces are useless if you don't have a base model that is good enough to be improved by them.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: