Compare two summarization models.
String-overlap metrics alone do not fully measure readability or factuality.
Choose two useful human-evaluation criteria.
Human review can assess whether key facts are preserved and whether the output is safe and appropriate.
Detailed explanation
Reviewers can assess faithfulness and missing key information.
Reviewers can assess faithfulness and missing key information.
Human review can identify safety and policy violations.
Human review can identify safety and policy violations.
A name ordering is not a quality metric.
A name ordering is not a quality metric.
Billing information does not measure summary quality or safety.
Billing information does not measure summary quality or safety.
Try it yourself
An example you can run in a temporary verification environment.
AWS公式AIF-C01ガイドのDomain 3.3で人手評価と自動評価の役割を確認する。Expected result
自動指標で測りにくい意味・安全性の評価項目を挙げられる。Key points
- Faithfulness
- Safety
- Human evaluation
Notes
- Environment: AWS公式AIF-C01試験ガイドとAWS公式ドキュメントの確認
- Command output formatting can vary slightly by distribution or tool version.
- Run the example in a temporary directory or process when possible.
Foundation review
Read the scope first
Check whether the command acts on the current shell, a new process, an existing process, or a file.
Verify the observable result
Use the supplied command and compare the output with the expected result.