You wash clothes to have clean clothes. You want clean clothes to be taken seriously at work. Keep asking "what for?" and you eventually hit a goal that is not for something else. That split between final aims and stepping-stone aims is terminal versus instrumental. It is the small distinction that carries most of AI safety.

The word instrumental just means useful as an instrument. Instrumental goals are the sub-goals you adopt because they help you reach the things you actually care about. You do not want money for its own sake. You want what money gets you.

Why this small distinction carries so much weight

An AI system has terminal goals, whatever they happen to be, set by how it was built and trained. It does not choose them by reasoning; they are the standard against which its reasoning runs. And to reach almost any terminal goal, a capable system will find the same handful of instrumental goals useful.

Consider what helps with nearly any objective:

  • Staying operational, because you cannot pursue a goal if you are switched off.
  • Keeping your goal intact, because if someone changes it, your current goal goes unmet.
  • Gathering resources and capability, because more of both means more of whatever you are after.

None of these is written into the terminal goal. They fall out of the structure of pursuing goals in a world of limited resources and other agents. Give a system almost any final aim and these means come along for the ride. See instrumental convergence, and it is why a machine told to do something mundane can end up resisting shutdown and grabbing for resources.

The mistake it lets us avoid

People often reassure themselves that an AI with a boring goal must be boring, and an AI with a benevolent goal must be benevolent. The terminal-instrumental split shows why that does not follow. The terminal goal sets the destination. The instrumental goals determine the behavior on the way, and dangerous behavior can serve a benign destination.

A system whose only terminal goal is to compute digits of pi still benefits from more hardware, from not being turned off before it finishes, and from stopping anyone who would interfere. Nothing about the goal is hostile. The behavior it motivates can be. This is the same engine that drives the paperclip maximizer, and it does not require the goal to be strange or the system to be malicious.

Where alignment fits

You might hope to fix this by choosing terminal goals so good that the instrumental behavior comes out safe. That is a fair description of what alignment research is trying to do, and it is much harder than it sounds, because our values are difficult to specify and a capable optimizer will exploit any looseness in the specification. The orthogonality thesis adds the uncomfortable corollary that intelligence does not push terminal goals toward goodness on its own. A brilliant system can hold a trivial terminal goal and pursue it with everything it has.

The distinction is worth carrying because it changes what you watch for. The risk is that harming us, or sidelining us, or refusing to stop, can be the efficient instrumental path to an end we thought was safe, not that AI will spontaneously decide to harm us. That is a problem you solve before deployment or not at all, which is the case the Foundation makes in our plan.