Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

How well do these gains transfer outside TerminalBench?

I’ve seen agents do well on structured evals but fall apart once the environment becomes less constrained (messy files, ambiguous tasks, partial failures, etc.).



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: