Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

The takeaway for that Pfau et al. paper is slightly more nuanced than that: It can only solve a subclass of problems without CoT, and that subclass can be equivalently solved with a larger model _without_ '...'

But arguably, a larger model will not need the chain of thought a smaller model does, which means simply by scaling we're already reducing CoT.

If the people who were relying on CoT are panicking now, they should've been panicking when perceptrons became multi-layer perceptrons.

 help



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: