Test a chatbot's safety.
Use normal and adversarial inputs.
Choose two evaluations.
Measure harmful categories, false positives, false negatives, paraphrases, prompts, and language differences rather than testing only normal questions.
Detailed explanation
Risk-specific weaknesses are visible.
Risk-specific weaknesses are visible.
Both overblocking and misses are measured.
Both overblocking and misses are measured.
Adversarial input is missing.
Adversarial input is missing.
Behavior changes with context and phrasing.
Behavior changes with context and phrasing.
Try it yourself
An example you can run in a temporary verification environment.
AWS公式AIF-C01 Domain 3.4の安全性評価、コンテンツフィルター、レッドチームを確認する。Expected result
有害出力と過剰拒否を複数ケースで評価できる。Key points
- Harm
- False refusal
- Red team
Notes
- Environment: AWS公式AIF-C01試験ガイドとAWS公式ドキュメントの確認
- Command output formatting can vary slightly by distribution or tool version.
- Run the example in a temporary directory or process when possible.
Foundation review
Read the scope first
Check whether the command acts on the current shell, a new process, an existing process, or a file.
Verify the observable result
Use the supplied command and compare the output with the expected result.