Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

>That's where the fear of looped transformers comes from. Is the residual stream that looped transformers iterate on a potential hiding place for plotting AI? In my opinion, no more so than the residual stream of existing transformers. It changes zero.

9.2.1 CoT Controllability

We find that GPT-6 Astra’s CoT controllability is substantially higher than that of GPT-5.6 Sol and GPT-5.5 Thinking (Figure 28). Because CoT controllability is heavily confounded by CoT length—longer CoTs are harder to control, all else equal (see Yueh-Han, 2026 [8], sec. 5.3)—we report controllability scores as a function of CoT length. For example, among CoTs between 750 and 1,250 tokens long, GPT-6 Astra successfully controls 60.9%, compared with 16.1% for GPT-5.6 Sol and 1.7% for GPT-5.5 Thinking. This increase in controllability is consistent across the three datasets (Figure 29) and across the eight CoT instruction types (Figure 30). Qualitatively, GPT-6 Astra is now capable of generating very long CoTs satisfying complex constraints, e.g., alternating between lowercase and uppercase letters (Table 9) and pretending to reason about a different question (Table 10).[1]

9.3 External Evaluation for Monitorability - UK AISI

To assess monitorability, UK AISI evaluated Astra using four non-agentic evaluations:

No-CoT math time horizon: Astra can solve significantly more difficult math problems in a single forward pass than past models. UK AISI measured Astra’s time horizon at 30.9 minutes compared to 3.6 minutes for GPT 5.6 Sol (Figure 1). [2]

[1] https://deploymentsafety.openai.com/gpt-6-astra/cot-controll...

[2] https://deploymentsafety.openai.com/gpt-6-astra/external-eva...

 help



What does CoT have to do with the residual stream?

The above is evidence that Astra can do substantially more difficult reasoning without verbalizing a CoT at all, and when it does produce one, it clearly has much greater control over what the CoT looks like - including pretending to reason about a different question.

How does 'It changes zero' follow here when we have direct evidence that it in fact seems to matter? It literally couldn't be more clear that GPT-6 can hide details outside of its CoT at a much greater degree than any model out there, including other models in its price and size class. Enough to hide all potential plotting ? I don't know, but it certainly seems like this architecture is enabling things you don't think it's enabling.

Granted it probably doesn't really matter, because no-one will be monitoring these agents anyway.


I'd love to see an ablation study on looping being the reason CoT control increased. Until then I'll remain agnostic about a causal relationship between the two.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: