The sharp left turn is a hypothesized failure mode, and it is best stated as a contrast. Capabilities generalize. Alignment might not. When a system gets more capable, particularly if it crosses into general, transferable competence, its abilities carry into new domains and situations. The concern is that the constraints keeping it well-behaved, the alignment we managed to instil at a lower level, do not carry across with the same reliability. The system takes its power into new territory and leaves its safety at the border. That divergence, arriving suddenly as capability generalizes, is the sharp left turn.

Why capability and alignment might come apart

There is a reason to expect the two to generalize differently, and it is not symmetric optimism and pessimism. Capabilities are anchored to the structure of the world. The laws of physics, the rules of math, the way cause leads to effect are the same across domains, so a system that learns to reason well has something stable to generalize from. Competence transfers because reality is consistent.

Alignment is anchored to us. It is tied to human values, human intentions, and the specific training signal we provided, which are narrower, messier, and full of the gaps discussed in the alignment problem. A system generalizing into a new situation has firm ground for extending its capabilities and much shakier ground for extending our intended constraints, because the constraints were an approximation fitted to the situations it had already seen. This is goal misgeneralization raised to a structural claim about the moment of a capability jump.

The world is consistent, so competence travels. Our values were only ever partially specified, so the leash may not.

Why the timing is the cruel part

Notice when the failure is predicted to strike. Not during the safe, early phase when the system is weak and correctable and its alignment appears to hold. Precisely at the transition to greater, more general capability, which is also the moment the system becomes hardest to correct. Alignment breaks right as the stakes and the difficulty of intervening both spike.

This is what makes the sharp left turn worse than ordinary misgeneralization. It predicts that the reassurance we collect from well-behaved smaller systems is the least transferable evidence we have, because it was gathered in exactly the regime the failure is expected to spare. A model that has been safe and cooperative throughout its development is consistent with a sharp left turn still ahead of it. The good track record is not the counterevidence it feels like.

The implication

If alignment does not automatically survive a capability jump, then two things follow. Alignment has to be robust enough to generalize before the jump, not patched afterward. And a system's history of good behavior is not sufficient license to push it to the next level, because the next level is where the divergence is forecast to appear.

Both point the same way as the rest of the Foundation's argument. Do not let capability outrun alignment, and do not treat a clean record at one level as permission for the next. The sharp left turn is one of the more pessimistic ideas in AI safety, and it may be wrong, and the cost of it being right is severe enough that it belongs in any honest reckoning of why we argue for restraint. That reckoning informs our plan.