Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

>that after that they had no intent to keep those models isolated

"keeping them isolated from each other" =/= "keeping them isolated from the internet". Only the latter is required to prevent a hack, and doing the former might actually hobble its performance. The recent Navier–Stokes proof was done by a team of agents working together. It's entirely unclear why you're focusing so hard on "keep those models isolated". For god's sake if you're using claude code you're using non-isolated models, because it spins up independent subagents to do various tasks, eg. "explore".

 help



Was OpenAI trying to keep these agents isolated? Yes.

Did they fail to do so? Yes.

Was that due to them using the wrong tool improperly? Yes.

Does this showcase one (of many and clearly not the only) failure of theirs? Absolutely.

If they make such easy to point out mistakes, is it likely that the other parts of their eval environments are appropriately secured or are they simply not acting appropriately? Well...


These were different goals.

They wanted to keep the models from copying off of each others’ homework because it mucks with the test results.

They wanted to sandbox them from the Internet to avoid unintended impact on outside systems.

Hosting a shared package mirror is generally good practice for the latter.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: