Reasoning through arguments against taking AI safety seriously
Dan PetersonSynopsis — AI-drafted from Dan's notes
Bengio wrote this about a year after he began warning publicly about AI. It takes thirteen objections to treating AI safety as a serious problem and answers them in turn, each under a heading that begins “For those who”. One claim sits behind all of them: we are racing to build systems as capable as human experts or more so, and nobody knows how to make sure such a system behaves as intended, or how to stop people abusing it. He treats alignment and control as one unsolved problem and governance as the other.
The first group of objections says the danger is impossible or far off. Bengio defines AGI by capability, as a system as good as or better than a human expert at basically any cognitive task, and says consciousness has nothing to do with the risk. He sees no scientific reason to think humans are the ceiling. The step from human level to beyond could be quick, since a system that matches a top AI researcher can be run as hundreds of thousands of copies working on AI itself. And because laws and regulators take years to build, uncertainty about timing is a reason to start early.
The second group says it will turn out fine. To those who expect a kind AI he answers that a system with a goal of preserving itself would resist being turned off, and that such a goal could come from people giving it one or as a side effect of pursuing another. To those who trust companies and existing law he answers that profit and safety have conflicted before, with fossil fuels and with thalidomide, and that a safety rule written in ordinary language has loopholes a superhuman system would find the way a legal team does. To those who want to accelerate he says the benefits are worth little if the bet loses, and that the people urging speed don’t carry the cost.
The third group is political. Talking about catastrophe need not crowd out present harms, he says, and people who want regulation for either reason should be allies. On the rivalry between the United States and China, his answer is that an uncontrolled superintelligence beats both sides, which gives them a shared interest. Treaties need that shared interest plus a way to verify compliance, and he points to control of high-end chips. To those who say the genie is out of the bottle, he replies that difficulty is no reason to stop trying and that lowering the odds of catastrophe is already a win. On open source he thinks releasing models helps safety for now, and asks who should decide when that changes.
The last two objections are about reasoning. Worrying is not Pascal’s wager, he says, because the probability is not tiny: he cites a survey in which the median AI researcher put 5% on extinction-level harm. And the lack of a quantitative model is a reason to study the risk, not to dismiss it.
One of his examples needs correcting. Bengio says an EPFL study showed GPT-4 out-persuading people when given their Facebook pages. The study, by Salvi and colleagues, gave the model basic sociodemographic information about its opponent, not Facebook pages. It found 81.7% higher odds of increased agreement, across 820 participants.