A formal proof demonstrates that it is possible to construct an advanced AI system that is reliably aligned and controllable at superhuman performance levels. The authors establish this by explicitly constructing a policy, environment, and control mechanism that satisfy all required mathematical conditions for alignment, control effectiveness, and reliability with parameters fixed in advance.

  • The system achieves superhuman performance on a defined benchmark, exceeding the 99th percentile of expert human performance on at least 99% of tasks while falling below the median on no more than 1%.
  • Alignment is verified through objective fidelity, action optimality, and preference stability, with the internal objective identically matching the principal utility.
  • Control effectiveness includes robustness against drift, adversarial perturbation, and model misspecification, with minimal invasiveness ensured by an empty intervention set.
  • Reliability is proven for a horizon of T=1, showing that the probability of failure to align is zero within the specified environment class.

The construction disproves the hypothesis that advanced AI systems cannot be reliably aligned, while clarifying that this existence proof does not contradict impossibility results regarding the verification or certification of arbitrary systems.