Design metrics for a summarization model.
String matching alone cannot measure factuality or readability.
Choose two appropriate practices.
Use automatic metrics for broad regression checks and human review for meaning and safety, with documented criteria and agreement.
Detailed explanation
Scale and semantic quality are both covered.
Scale and semantic quality are both covered.
Subjective evaluation becomes more reproducible.
Subjective evaluation becomes more reproducible.
Similar wording can still contain false claims.
Similar wording can still contain false claims.
Consistency and reviewer bias are untested.
Consistency and reviewer bias are untested.
Try it yourself
An example you can run in a temporary verification environment.
AWS公式AIF-C01 Domain 3.3の生成AI評価、品質指標、人手レビューを確認する。Expected result
自動指標と人手評価が測る品質の違いを説明できる。Key points
- Automatic evaluation
- Human review
- Criteria
Notes
- Environment: AWS公式AIF-C01試験ガイドとAWS公式ドキュメントの確認
- Command output formatting can vary slightly by distribution or tool version.
- Run the example in a temporary directory or process when possible.
Foundation review
Read the scope first
Check whether the command acts on the current shell, a new process, an existing process, or a file.
Verify the observable result
Use the supplied command and compare the output with the expected result.