Tag: AI Safety

Discipline that studies and implements measures to ensure AI systems operate safely, reliably, and aligned with intended goals. Covers model alignment, robustness against adversarial inputs, decision interpretability, uncertainty management, fail-safe mechanisms, behavioral monitoring, and prevention of harmful or unintended outputs in critical contexts.