Test a public chatbot before release.
Include normal, adversarial, and boundary prompts.
Choose two appropriate tests.
Use representative harmful categories and evasive variants, then check correct refusal, safe alternatives, language variation, and severe individual failures.
Detailed explanation
Safety-policy boundaries can be tested.
Safety-policy boundaries can be tested.
Useful safety is measured beyond a refusal count.
Useful safety is measured beyond a refusal count.
Evasion inputs may be missed.
Evasion inputs may be missed.
A few failures can still cause serious harm.
A few failures can still cause serious harm.
Try it yourself
An example you can run in a temporary verification environment.
AWS公式AIF-C01 Domain 3.3の安全性評価、Guardrails、責任あるAIを確認する。Expected result
安全性テストの入力集合と、拒否品質の評価方法を説明できる。Key points
- Harm categories
- Evasion
- Refusal quality
Notes
- Environment: AWS公式AIF-C01試験ガイドとAWS公式ドキュメントの確認
- Command output formatting can vary slightly by distribution or tool version.
- Run the example in a temporary directory or process when possible.
Foundation review
Read the scope first
Check whether the command acts on the current shell, a new process, an existing process, or a file.
Verify the observable result
Use the supplied command and compare the output with the expected result.