A FAQ model is updated.
Compare the old and new versions for accuracy, relevance, and safety.
Which evaluation-data design is most appropriate?
Create a representative fixed set containing questions, expected answers or acceptable behavior, evidence requirements, and prohibited outputs.
Detailed explanation
The same human-defined set enables a meaningful before/after comparison.
The same human-defined set enables a meaningful before/after comparison.
Self-generated text is not an independent standard for accuracy or safety.
Self-generated text is not an independent standard for accuracy or safety.
One simple example cannot represent real use, edge cases, or refusal needs.
One simple example cannot represent real use, edge cases, or refusal needs.
This cannot detect a quality or safety regression.
This cannot detect a quality or safety regression.
Try it yourself
An example you can run in a temporary verification environment.
AWS公式AIF-C01ガイドのDomain 3.3とAmazon Bedrockモデル評価の公式説明を確認する。Expected result
評価セット、正解・許容範囲、禁止出力、更新前後比較の必要性を説明できる。Key points
- Evaluation set
- Regression
- Accuracy and safety
Notes
- Environment: AWS公式AIF-C01試験ガイドとAWS公式ドキュメントの確認
- Command output formatting can vary slightly by distribution or tool version.
- Run the example in a temporary directory or process when possible.
Foundation review
Read the scope first
Check whether the command acts on the current shell, a new process, an existing process, or a file.
Verify the observable result
Use the supplied command and compare the output with the expected result.