Was OpenAI trying to keep these agents isolated? Yes.
Did they fail to do so? Yes.
Was that due to them using the wrong tool improperly? Yes.
Does this showcase one (of many and clearly not the only) failure of theirs? Absolutely.
If they make such easy to point out mistakes, is it likely that the other parts of their eval environments are appropriately secured or are they simply not acting appropriately? Well...
Did they fail to do so? Yes.
Was that due to them using the wrong tool improperly? Yes.
Does this showcase one (of many and clearly not the only) failure of theirs? Absolutely.
If they make such easy to point out mistakes, is it likely that the other parts of their eval environments are appropriately secured or are they simply not acting appropriately? Well...