Like Sabine Hossenfelder, I Was Offered Money to Tell You AI Will Kill Us

usefulHouse of Elvideo2026-09-22read ai llmsphilosophycognitive science

Synopsis — AI-drafted from Dan's notes

A computer scientist works through twelve papers to find where the published evidence stops on the route from recursive self-improvement to human extinction. Her finding is that the route is not one claim but four, at different stages of proof: self-improvement with a scoreboard, self-improvement without one, a runaway loop, and extinction.

Bounded self-improvement is demonstrated: systems rewrite their own agent code and score better, and one of them improved the infrastructure that trains AI. However, every one of those results had an external scoreboard telling it whether it had succeeded. Take the scoreboard away and you are asking for scientific judgement instead of optimisation, which is the step nobody has shown. The runaway loop is modelled rather than observed (the models disagree about whether the feedback is strong enough yet). Extinction is a chain of its own, which she splits into three further claims, each needing its own evidence: capability, propensity and opportunity. Even then it needs a pathogen or a cyber-attack that reaches every last person, including the ones living off-grid.

Her objection is therefore narrow. She is not arguing against safety research or regulation; she argues the opposite. The objection is to collapsing “worth investigating”, “plausible”, “forecast” and “demonstrated” into one category, and then making policy as though the last word applied to all four.

Where I land

House of El cuts through the hype and doom/gloom and objectively presents what the bleeding-edge research shows. I also liked one of the papers that shows how rationalists flatten 3 very substantial claims into one (capability, intent, opportunity).

I think her position is stronger than Hinton’s (for whom I hold tremendous respect). Hinton’s position isn’t based on any data, except for the “expert intuition” which he cites. A model of the conditions under which something could happen is still not evidence that it is happening. How would you ever know a priori? There are so many unknowns here: we don’t even know if RSI as it is stated, without any human intervention, is even possible.

Of the three claims being flattened together, I think intent does the most damage. The media, and through their reporting the general public, are leaning on the expert opinion of the safety researchers. We expect labs to spin things towards whatever direction is in their corporate interests; that is understood, and it is the first argument most people have when reading about their claims. However, the safety researchers themselves (especially ones not aligned with any particular frontier lab, e.g. Hinton) are the ones that make everyone get fired up and want to burn everything down, despite the benefits LLMs have (and I believe will continue to have) as a technological tool. Or worse: people become apathetic to the repetitious noise and do not see trouble if and when it ever comes.

Connections