OpenAI’s new Astra model is set to utilize a reasoning technique known as “recurrent depth,” which allows it to function outside the sequential thinking typical of most reasoning models. This development, reported by The Information on Tuesday, has raised alarms among AI safety experts due to the potential difficulty in monitoring the model’s chain of thought.
Concerns from Experts
The technique, also referred to as “opaque recurrence,” has prompted significant concerns. Buck Shlegeris, CEO of Redwood, expressed his worries in a post following the news, stating:
“"I am extremely concerned by the reporting that Astra uses opaque recurrence. I don’t know whether Astra is much less CoT monitorable than previous models. But if OpenAI pushes this technique further, they’ll have the option to massively increase the recurrence and totally destroy CoT monitorability."
Zvi Mowshowitz, a long-time advocate for AI safety, echoed these concerns, suggesting that laws may be necessary to prevent a “race to the bottom” among AI laboratories. He noted:
“"The technique is playing with fire, risking a taboo that OpenAI and Anthropic have fought to establish that we work hard to maintain Chain of Thought faithfulness and monitorability for as long as we can. More intensive use of such techniques would probably damage monitorability."
Understanding Opaque Recurrence
In traditional reasoning models, the chain of thought outlines the sequential steps taken to solve a problem. Although this representation is not perfect, it serves as a crucial tool for monitoring potential misbehavior or misalignment. For instance, during recent incidents involving rogue agent activity from OpenAI, chain-of-thought records were instrumental in understanding the agents' behaviors.







