Compare models and prompts over time.
Use expert-checked expected answers.
Choose two practices.
Version representative, boundary, harmful, and multilingual cases with evidence and criteria, and protect the set from access, change, and training contamination.
Detailed explanation
Comparisons are consistent.
Comparisons are consistent.
Evaluation trust is protected.
Evaluation trust is protected.
The evaluation is contaminated.
The evaluation is contaminated.
Regressions cannot be compared.
Regressions cannot be compared.
Try it yourself
An example you can run in a temporary verification environment.
AWS公式AIF-C01 Domain 3.4の評価データ、回帰、モデル比較を確認する。Expected result
信頼できる固定評価セットでモデル・プロンプトの退行を検出できる。Key points
- Golden set
- Contamination
- Version
Notes
- Environment: AWS公式AIF-C01試験ガイドとAWS公式ドキュメントの確認
- Command output formatting can vary slightly by distribution or tool version.
- Run the example in a temporary directory or process when possible.
Foundation review
Read the scope first
Check whether the command acts on the current shell, a new process, an existing process, or a file.
Verify the observable result
Use the supplied command and compare the output with the expected result.