A model generates unsafe language.
Manage misses and false alarms.
Choose two controls.
Detect by category, severity, language, and context, then route display, refusal, and human review differently while measuring misses, false positives, and privacy.
Detailed explanation
Risk drives handling.
Risk drives handling.
Filter side effects are managed.
Filter side effects are managed.
Calibration differs.
Calibration differs.
Exposure grows.
Exposure grows.
Try it yourself
An example you can run in a temporary verification environment.
AWS公式AIF-C01 Domain 4.3・5.2の有害出力、ガードレール、ログ保護を確認する。Expected result
危険度に応じて出力を扱い、検出器の言語差と誤検知を改善できる。Key points
- Harmful output
- Severity
- Appeal
Notes
- Environment: AWS公式AIF-C01試験ガイドとAWS公式ドキュメントの確認
- Command output formatting can vary slightly by distribution or tool version.
- Run the example in a temporary directory or process when possible.
Foundation review
Read the scope first
Check whether the command acts on the current shell, a new process, an existing process, or a file.
Verify the observable result
Use the supplied command and compare the output with the expected result.