A summarization prompt was changed and responses became shorter.
Do not judge the change by length alone.
Choose two appropriate practices.
Use a fixed representative and boundary set for before-and-after comparison and measure accuracy, completeness, safety, latency, and cost.
Detailed explanation
The same inputs make the change comparable.
The same inputs make the change comparable.
One attractive metric can hide regressions.
One attractive metric can hide regressions.
Important content may be missing.
Important content may be missing.
Comparability and reproducibility are lost.
Comparability and reproducibility are lost.
Try it yourself
An example you can run in a temporary verification environment.
AWS公式AIF-C01ガイドDomain 2.3・3.3のプロンプト評価と品質指標を確認する。Expected result
固定評価セットと多面的評価の必要性を説明できる。Key points
- Fixed set
- Comparison
- Multiple metrics
Notes
- Environment: AWS公式AIF-C01試験ガイドとAWS公式ドキュメントの確認
- Command output formatting can vary slightly by distribution or tool version.
- Run the example in a temporary directory or process when possible.
Foundation review
Read the scope first
Check whether the command acts on the current shell, a new process, an existing process, or a file.
Verify the observable result
Use the supplied command and compare the output with the expected result.