Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I suspect it’s a side effect of heavy RL that rewards solved problems but not writing clarity.
 help



With human RL, sounding like you've solved a problem is even better. So many times Codex writes some enthusiastic paragraph, then I learn later that it never reran the tests, or had to add some insane hard-coded hack that renders the feature useless for the general case, etc.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: