Tag: Testing for Goal Alignment

Verification of the alignment between an AI system’s stated goals and its actual behavior. Covers testing techniques to detect deviations, unintended emergent behaviors, goal misalignment, and situations where the model optimizes surrogate metrics instead of real objectives, with a focus on reward hacking and specification gaming risks.