>That's where the fear of looped transformers comes from. Is the residual stream that looped transformers iterate on a potential hiding place for plotting AI?
In my opinion, no more so than the residual stream of existing transformers. It changes zero.
9.2.1 CoT Controllability
We find that GPT-6 Astra’s CoT controllability is substantially higher than that of GPT-5.6 Sol and GPT-5.5 Thinking (Figure 28). Because CoT controllability is heavily confounded by CoT length—longer CoTs are harder to control, all else equal (see Yueh-Han, 2026 [8], sec. 5.3)—we report controllability scores as a function of CoT length. For example, among CoTs between 750 and 1,250 tokens long, GPT-6 Astra successfully controls 60.9%, compared with 16.1% for GPT-5.6 Sol and 1.7% for GPT-5.5 Thinking. This increase in controllability is consistent across the three datasets (Figure 29) and across the eight CoT instruction types (Figure 30). Qualitatively, GPT-6 Astra is now capable of generating very long CoTs satisfying complex constraints, e.g., alternating between lowercase and uppercase letters (Table 9) and pretending to reason about a different question (Table 10).[1]
9.3 External Evaluation for Monitorability - UK AISI
To assess monitorability, UK AISI evaluated Astra using four non-agentic evaluations:
No-CoT math time horizon: Astra can solve significantly more difficult math problems in a single forward pass than past models. UK AISI measured Astra’s time horizon at 30.9 minutes compared to 3.6 minutes for GPT 5.6 Sol (Figure 1). [2]
The above is evidence that Astra can do substantially more difficult reasoning without verbalizing a CoT at all, and when it does produce one, it clearly has much greater control over what the CoT looks like - including pretending to reason about a different question.
How does 'It changes zero' follow here when we have direct evidence that it in fact seems to matter? It literally couldn't be more clear that GPT-6 can hide details outside of its CoT at a much greater degree than any model out there, including other models in its price and size class. Enough to hide all potential plotting ? I don't know, but it certainly seems like this architecture is enabling things you don't think it's enabling.
Granted it probably doesn't really matter, because no-one will be monitoring these agents anyway.
I'd love to see an ablation study on looping being the reason CoT control increased. Until then I'll remain agnostic about a causal relationship between the two.
9.2.1 CoT Controllability
We find that GPT-6 Astra’s CoT controllability is substantially higher than that of GPT-5.6 Sol and GPT-5.5 Thinking (Figure 28). Because CoT controllability is heavily confounded by CoT length—longer CoTs are harder to control, all else equal (see Yueh-Han, 2026 [8], sec. 5.3)—we report controllability scores as a function of CoT length. For example, among CoTs between 750 and 1,250 tokens long, GPT-6 Astra successfully controls 60.9%, compared with 16.1% for GPT-5.6 Sol and 1.7% for GPT-5.5 Thinking. This increase in controllability is consistent across the three datasets (Figure 29) and across the eight CoT instruction types (Figure 30). Qualitatively, GPT-6 Astra is now capable of generating very long CoTs satisfying complex constraints, e.g., alternating between lowercase and uppercase letters (Table 9) and pretending to reason about a different question (Table 10).[1]
9.3 External Evaluation for Monitorability - UK AISI
To assess monitorability, UK AISI evaluated Astra using four non-agentic evaluations:
No-CoT math time horizon: Astra can solve significantly more difficult math problems in a single forward pass than past models. UK AISI measured Astra’s time horizon at 30.9 minutes compared to 3.6 minutes for GPT 5.6 Sol (Figure 1). [2]
[1] https://deploymentsafety.openai.com/gpt-6-astra/cot-controll...
[2] https://deploymentsafety.openai.com/gpt-6-astra/external-eva...