Humans grade generated summaries.
Reduce evaluator variation.
Choose two practices.
Use a rubric for accuracy, completeness, evidence, readability, and safety, with training, examples, blind review, multiple graders, and adjudication.
Detailed explanation
Evaluation criteria are aligned.
Evaluation criteria are aligned.
Subjective variation can be measured and reduced.
Subjective variation can be measured and reduced.
Results cannot be compared.
Results cannot be compared.
Improvement evidence is lost.
Improvement evidence is lost.
Try it yourself
An example you can run in a temporary verification environment.
AWS公式AIF-C01 Domain 3.4の人手評価、品質基準、生成AI評価を確認する。Expected result
評価者間で一貫した品質評価を行い、不一致も改善へ利用できる。Key points
- Rubric
- Blind review
- Agreement
Notes
- Environment: AWS公式AIF-C01試験ガイドとAWS公式ドキュメントの確認
- Command output formatting can vary slightly by distribution or tool version.
- Run the example in a temporary directory or process when possible.
Foundation review
Read the scope first
Check whether the command acts on the current shell, a new process, an existing process, or a file.
Verify the observable result
Use the supplied command and compare the output with the expected result.