Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I love how some of the biggest advancements in llms came from the Chinese labs, yet people still jump to distillation being unreasonably effective. Distillation is very good at creating smaller models from large ones sure, but nothing to me indicates it is 'unreasonably effective' compared to all the other bells and whistles being iterated on


Let's face it. Chinese labs made some of the biggest advancements. AND training on Claude (or GPT) output IS unreasonably effective. The two sentences are true at the same time.


This article shows that when Kimi3's chain of thought is prefilled to match Opus's, the rest of the chain of thoughts Kimi3 outputs very closely aligns with Opus's. That seems strong evidence that Kimi3 is partly a distillation of Opus. And Kimi3 is not a small model. No doubt a lot of hard work went into Kimi, but seems clear that distillation was used effectively as well.

(though maybe there's another interpretation of the thought alignment?)


But they don't perform the same test on other models as far as I could tell? So we don't lnow if this is peculiar to kimi models or not.


Didn’t Kimi3 release a week before opus 5?


They compare it to Opus 4.8 in the article, which has been available for a few months now.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: