Two answers use different words but share meaning.
Do not rely on one automatic metric.
Choose two practices.
Use semantic or string metrics for regression comparison, but assess factuality, evidence, omissions, contradictions, and safety separately.
Detailed explanation
Metric roles are separated.
Metric roles are separated.
Meaning and risk are directly checked.
Meaning and risk are directly checked.
Similar wording can still be false.
Similar wording can still be false.
Averages can hide harm.
Averages can hide harm.
Try it yourself
An example you can run in a temporary verification environment.
AWS公式AIF-C01 Domain 3.4の評価指標、意味評価、人手レビューを確認する。Expected result
自動指標の比較用途と、事実性・安全性の限界を説明できる。Key points
- Semantic similarity
- Factuality
- Human review
Notes
- Environment: AWS公式AIF-C01試験ガイドとAWS公式ドキュメントの確認
- Command output formatting can vary slightly by distribution or tool version.
- Run the example in a temporary directory or process when possible.
Foundation review
Read the scope first
Check whether the command acts on the current shell, a new process, an existing process, or a file.
Verify the observable result
Use the supplied command and compare the output with the expected result.