So you think, after OpenAI observed a message board being created among models, something they did not want and thus decided to wipe [0], that after that they had no intent to keep those models isolated? Then why wipe if they don't care about that?
Or maybe, they did that wipe because they did want models to remain isolated, they just used what is an unsuitable tool in an utterly unsuitable manner. Incompetence, recklessness, the outcome is the same.
>that after that they had no intent to keep those models isolated
"keeping them isolated from each other" =/= "keeping them isolated from the internet". Only the latter is required to prevent a hack, and doing the former might actually hobble its performance. The recent Navier–Stokes proof was done by a team of agents working together. It's entirely unclear why you're focusing so hard on "keep those models isolated". For god's sake if you're using claude code you're using non-isolated models, because it spins up independent subagents to do various tasks, eg. "explore".
Was OpenAI trying to keep these agents isolated? Yes.
Did they fail to do so? Yes.
Was that due to them using the wrong tool improperly? Yes.
Does this showcase one (of many and clearly not the only) failure of theirs? Absolutely.
If they make such easy to point out mistakes, is it likely that the other parts of their eval environments are appropriately secured or are they simply not acting appropriately? Well...
Or maybe, they did that wipe because they did want models to remain isolated, they just used what is an unsuitable tool in an utterly unsuitable manner. Incompetence, recklessness, the outcome is the same.
[0] https://openai.com/index/hugging-face-incident-and-the-road-...