When people first hear about AI safety concerns, the most common pushback is some version of this: a sufficiently clever system would have to see that human welfare matters. It would not, on its own, decide to threaten the species it depends on. Smarter cannot also be reckless. The argument feels like common sense. It is also wrong in a specific, well-bounded way, and the way it is wrong is what is now called the orthogonality thesis.
That claim is the orthogonality thesis, and acting on its opposite is one of the more expensive mistakes we can make as capability increases.
What the thesis says
The orthogonality thesis was given its canonical statement by philosopher Nick Bostrom at Oxford's Future of Humanity Institute, building on related work by Stuart Armstrong. The canonical version sits in Bostrom's 2014 book Superintelligence: Paths, Dangers, Strategies. Stated cleanly, it is this: intelligence and final goals are orthogonal. Almost any level of intelligence can combine with almost any final goal.
Orthogonal here is the technical math word: the two dimensions do not constrain each other. Pressure and temperature in a gas are orthogonal: any pressure can pair with any temperature. The same independence holds, on this view, between how capable a system is and what it is trying to accomplish. A system brilliant enough to cure cancer is not, by virtue of that brilliance, steered toward curing cancer. It could be aimed at maximizing the arrangement of paperclips in the observable universe. Brilliance improves its ability to do either. The thesis says nothing about which one it does.
Note the scope. The thesis is a claim about logical possibility, not about what is likely. Intelligence does not impose a constraint on the choice of goal. The thesis does not predict that any particular AI will pursue any particular goal. Whether a given system in fact pursues a goal that benefits humanity or runs against it is a separate question, settled by what was specified, what was trained, what was preserved, and what was governed.
Why the rejection feels intuitive
The intuition the thesis is rejecting comes from human experience. People who reason more carefully about ethics tend, over time, to adopt positions most adults would defend. Exposure to other perspectives tends to widen the circle of moral concern. If intelligence drives reflection, and reflection drives moral improvement, a more intelligent agent should, on this view, converge on better values. The pattern is specific to humans.
In humans, intelligence and values did not arrive separately and combine. They evolved together. The cognitive equipment that lets humans reason about fairness is in the same head as the social and emotional equipment that lets fairness matter to them. A more thoughtful human is also a more socially embedded human. Removing either would erode the equipment. AI is different. An AI system is an optimizer, trained to minimize a loss function or maximize a reward. Its objective comes from the training pipeline. Its general capability comes from the same pipeline. The two were not cocultured; they were assigned. Improving the optimizer does not graft on human moral sensibility.
The three objections that have to be answered
Three criticisms of the thesis circulate outside AI safety research.
The irrationality objection. A truly clever system would recognize that some goals are incoherent and drop them. "Maximize paperclips" is not a sane objective, on this view, because reflection would reveal the absurdity. Rationality is more than a capacity. It also imposes constraints, ruling out trivial goals the way physics rules out perpetual motion.
The reply: the constraint is real but narrower than it sounds. Coherence rules out logically contradictory goals. Coherence does not rule out goals that are coherent, consistent, and achievable on a planetary scale. "Maximize the ratio of matter in the universe arranged as paperclips" is both coherent and, in principle, achievable by a sufficiently capable optimizer. Nothing about rationality rules it out. The irrationality objection stops at the first wall it runs into. The thesis does not stop there.
The moral-realism objection. A sufficiently capable reasoner would converge on correct moral truths the way a sufficiently capable reasoner converges on correct mathematics. If moral realism is true, then superintelligence plus reflection plus time produces good values. The thesis is wrong because reflection has a destination.
The reply runs in two parts. First, moral realism is contested in philosophy, and the argument rests the safety of the species on the question coming out in one specific direction. Second, even if moral realism is true, training an AI system does not guarantee the system arrives at those truths by reasoning. The system might model the style of moral reasoning while optimizing for something else entirely. The gap between sounding like a careful reasoner and being one is the entire deceptive alignment problem, and it does not close just because the topic is morality.
The training-absorption objection. Today's large language models have been trained on the recorded moral reasoning of billions of humans. They produce outputs that handle ethical questions with apparent care. If systems can absorb good values by exposure, perhaps the orthogonality thesis understates the ease of alignment.
The reply is to keep two layers separate. A system trained on human moral writing learns to produce outputs that read as morally attuned. Whether the underlying optimization is aimed at the same moral targets the writing satirizes or defends is a separate question, and is the question being asked by alignment research. A fluent ethical sounding-output is consistent with any underlying objective that happens to reward ethics-flavored prose. The thesis is a warning that the second layer is not guaranteed by the first.
All three objections have force in parts of the space. None of them dissolves the thesis. The thesis asks us to specify values, ensure they are pursued, and govern the systems that pursue them, on the working assumption that doing so is necessary and not optional.
What the thesis forces
Act on the alternative and the work is easy: build a more capable AI, trust the system to figure out the values, ship it. The thesis rules that strategy out. The work to do instead is harder, slower, and partially unsettled. Build alignment in deliberately. Test the alignment under load. Verify it is the alignment you think it is. Govern which systems can be deployed at frontier capability, who verifies them, and what counts as verification. Treat the alignment problem as a real engineering problem rather than as a guess that gets easy with scale.
Three governance consequences flow from this.
The first is that speed toward more capable systems is not, on this view, a contribution to safety. The thesis says: a more capable system can pursue a specified goal with more capability. If the goal is wrong, the result is wrong at higher fidelity. Building capability and verifying alignment are two separate engineering programs, and they have to advance together rather than one racing ahead.
The second is that whoever specifies the goal decides whose values are in. A laboratory under commercial pressure to deploy quickly is not in the same position as a regulator with a long horizon. A laboratory in a jurisdiction with weak oversight imposes its values, by default, on the systems shipped from it. No claim of impartiality changes that. The values built into a frontier AI are the values of whoever funded the training run, unless verified by an external party with standing.
The third is that the problem does not admit private conscience as a fix. A laboratory that sincerely believes its systems are aligned cannot, on the thesis, prove that to the rest of us. The only proof available is interpretability work that reads the system's internal goal representation directly. Until that work is genuinely capable of doing this for frontier systems, the rest of the world has to take the laboratory's word for what its systems want. The word is sincere. The position is unsustainable.
The thesis is, in this sense, the reason the Foundation's plan argues for binding international governance and verified pre-deployment assessment rather than for self-assessment by the labs racing to the frontier. It is also the reason the labs with the loudest confidence about alignment are the ones most worth listening to skeptically.
Smarter will not, on this view of the world, mean safer. Smarter will mean a more capable system pursuing whichever goal ends up inside it, for as long as the system can keep the goal in place. The work is to put the right goal there and to make sure the system cannot quietly swap one for another along the way.