Filters may use rules, classifiers, similarity checks, or model-based review before and after generation. Thresholds create false-positive and false-negative tradeoffs, so appeals and monitoring may be needed.
Content filtering detects or blocks inputs and outputs that match a defined safety or usage policy.
Filters may use rules, classifiers, similarity checks, or model-based review before and after generation. Thresholds create false-positive and false-negative tradeoffs, so appeals and monitoring may be needed.