Traceability of AI thought processes

New Reasoning Technique in OpenAI’s Astra Model Alarms AI Safety Experts

KI, Sicherheit, Gefahr, Shutterstock
Facebook
X
LinkedIn
Reddit
WhatsApp

According to a report by The Information, OpenAI’s upcoming Astra model is said to use a reasoning technique referred to as “recurrent depth” or “opaque recurrence.”

Buck Shlegeris, CEO of Redwood Research, a company specializing in AI safety research, expressed public concern after the news broke. He stated he did not know whether Astra’s chain of thought would actually be significantly harder to monitor as a result compared to earlier models; however, if OpenAI were to push the technique further, the company would have the ability to massively amplify the recurrence and thereby completely destroy the monitorability of the chain of thought.

Ad

“I am extremely concerned about the reports that Astra uses opaque recurrence.”

Buck Shlegeris, CEO of Redwood Research, a company specializing in AI safety research

Ad

Longtime AI safety advocate Zvi Mowshowitz explained that legal regulations might become necessary to prevent a so-called race to the bottom among AI labs. “The technique is playing with fire and risking a taboo.” OpenAI and Anthropic had jointly established this taboo with great effort, with both companies striving to maintain the traceability and reliability of the chain of thought for as long as possible; a more intensive use of such techniques would presumably damage this monitorability.

Astra: Concern over completely invisible thought process

Ryan Greenblatt, Chief Scientist at Redwood Research, additionally warned that opaque reasoning could scale significantly faster than classic, step-by-step thinking in text form, which could effectively remove a model’s entire thought process from visible channels: “My biggest concern is that a natural progression from here.” According to him, this progression could consist of expanding opaque reasoning to such an extent that a model ultimately thinks completely or almost completely in a hidden, not directly observable feature space; he hopes it is not already too late to avoid the most concerning architectures and that OpenAI pauses at this point.

OpenAI points to limited use

OpenAI Chief Scientist Jakub Pachocki responded to the reporting by pointing out that the use of the technique in Astra has been limited so far: “OpenAI has strived to maintain and utilize chain-of-thought monitoring since our very first reasoning models.” This remains a central goal of the company’s ongoing research program.

(Editorial Team)

Ad

Weitere Artikel