The protocol defines threat models, test prompts, tools, scoring rules, and escalation thresholds. Results guide mitigations and deployment decisions but cover only the tested conditions.
A safety evaluation tests an AI system for specified harmful capabilities, behaviors, or policy violations.
The protocol defines threat models, test prompts, tools, scoring rules, and escalation thresholds. Results guide mitigations and deployment decisions but cover only the tested conditions.