The above is evidence that Astra can do substantially more difficult reasoning without verbalizing a CoT at all, and when it does produce one, it clearly has much greater control over what the CoT looks like - including pretending to reason about a different question.
How does 'It changes zero' follow here when we have direct evidence that it in fact seems to matter? It literally couldn't be more clear that GPT-6 can hide details outside of its CoT at a much greater degree than any model out there, including other models in its price and size class. Enough to hide all potential plotting ? I don't know, but it certainly seems like this architecture is enabling things you don't think it's enabling.
Granted it probably doesn't really matter, because no-one will be monitoring these agents anyway.
I'd love to see an ablation study on looping being the reason CoT control increased. Until then I'll remain agnostic about a causal relationship between the two.