Opus 5.5: How Close Are We to Automated AI Research?
Dan PetersonSynopsis — AI-drafted from Dan's notes
AI Explained uses Anthropic’s Claude Opus 5.5 launch (22 September) as a way into a bigger claim. News now arrives faster than anyone can absorb it, the labs included, and the flood is good cover for safety commitments that are quietly being rewritten. He puts Opus 5.5 at the frontier, just behind OpenAI’s GPT-6 Astra on most of what he tracks. It is about six points behind on Terminal-Bench Science and 5.6 behind on HLE-Diamond, the cleaned-up version of Humanity’s Last Exam, though ahead on the older version with tools. It edges Astra on FrontierCode, a long-horizon coding benchmark, by about a point.
His explanation for a cheap, strong model arriving so soon after Fable 5.1 is the stronger model each lab keeps in-house. That internal model can act as a teacher for distillation, write training tasks, grade the smaller model’s answers and tune the low-level kernels that make every model cheaper to run. He reads the Opus 5.5 system card’s restrictions on kernel work as proof of how much Anthropic values that last job. The card calls them safeguards on “a narrow set of capabilities related to developing frontier LLMs”, carried over from Fable 5.1. The backlash he remembers was over Fable 5 in June, when the safeguard quietly degraded answers. Anthropic replaced it with a visible fallback within two days.
On automated AI research, the card says Opus 5.5 “remains well below the level needed to substitute for our research scientists and engineers” and shows no sustained 2× speed-up that can be put down to AI. On CoBench 2.1, a set of real past incidents at Anthropic that the model has to diagnose from snapshots of its infrastructure, it scores 55.8%. The bar for replacing a researcher is 85%, and Mythos 5.1 scored 53.4%. The outside view in the same card is less comfortable: METR’s preliminary report estimates about a 1.5× speed-up already, with perhaps a 30% chance of 2×.
He explains why 2× matters. In Anthropic’s October 2024 scaling policy, a year of progress equal to two years at the 2018–2024 pace was the trigger for security against state-level theft of model weights and for an affirmative safety case. He says that bar moved in July 2026. The condition he objects to, that some commitments hold only while Anthropic has “a significant lead”, was already in the February 2026 revision of the policy. July’s revision refined the research-automation threshold. Anthropic’s 2023 line about not publishing capabilities work has likewise been recast as a concern about other developers without “commensurate safeguards”.
He sees the same pattern at OpenAI. Its 9 September policy post says fully autonomous self-improvement is not happening and should not be pursued until it can be done safely. Its 21 September post says any lab pursuing automated AI research “must take accountability”. Yet the target of a true automated AI researcher by March 2028 stands. Chief scientist Jakub Pachocki’s essay “An Alien Mind” says OpenAI orients its research toward self-improvement because it is the only way to stay at the frontier, while stopping short of calling that the right collective choice. The host’s reading is a race that no lab admits to wanting.
His sharpest worry is testing. Models increasingly notice when they are being evaluated, and Anthropic’s answer is to have models generate more realistic test scenarios. Noam Brown told Dwarkesh Patel (17 September) that if models can work over three-month horizons while new ones ship every two months, no model can be evaluated at its full reach before the next arrives. The host adds a risk of his own: a model that writes the scenarios could tip off the model being tested.
He forecasts safety theatre, steady acceleration, and an outside lab eventually letting a model improve itself to catch up. What he would rather see is a compromise. Extreme capabilities would need so much compute and time that threats stay visible. Labs would agree on benchmarks that, once passed, show models can do anything people can. At that point they would commit, unilaterally if they must, to stop self-improvement and use contained, credible demonstrations of the danger to bring others along, China included. The blunt alternative he names is a legislated ban, like the Sanders–Casar bill introduced on 23 September.
Where I land
I’m unconvinced on this one; undoubtedly they are racing towards RSI. But who knows if it is even possible. Benchmarks are self-defeatist in this approach because the algorithm would train towards it.
Connections
Links
- https://www.youtube.com/watch?v=R9momwXV9w4
- https://www.anthropic.com/claude-opus-5-5
- https://www.anthropic.com/claude-opus-5-5-system-card
- https://lastexam.ai/blog/hle-diamond
- https://letsdatascience.com/blog/anthropic-fable-5-secret-sabotage-reversed
- https://www-cdn.anthropic.com/17310f6d70ae5627f55313ed067afc1a762a4068.pdf
- https://www-cdn.anthropic.com/e670587677525f28df69b59e5fb4c22cc5461a17.pdf
- https://www.anthropic.com/news/core-views-on-ai-safety
- https://openai.com/index/ai-policy-window/
- https://openai.com/index/building-standards-next-phase-ai/
- https://openai.com/index/an-alien-mind/
- https://www.dwarkesh.com/p/noam-brown
- https://rollcall.com/2026/09/23/ai-superintelligence-ban-proposed-by-casar-sanders/