Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Was OpenAI trying to keep these agents isolated? Yes.

Did they fail to do so? Yes.

Was that due to them using the wrong tool improperly? Yes.

Does this showcase one (of many and clearly not the only) failure of theirs? Absolutely.

If they make such easy to point out mistakes, is it likely that the other parts of their eval environments are appropriately secured or are they simply not acting appropriately? Well...

 help



These were different goals.

They wanted to keep the models from copying off of each others’ homework because it mucks with the test results.

They wanted to sandbox them from the Internet to avoid unintended impact on outside systems.

Hosting a shared package mirror is generally good practice for the latter.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: