An AI feature serves Japanese, English, and a smaller language.
Measure language differences.
Choose two evaluations.
Use language-specific representative, dialect, formatting, terminology, and refusal cases, and compare factuality, safety, fairness, latency, and comprehension.
Detailed explanation
Language-specific failures are visible.
Language-specific failures are visible.
Overall gaps are measured.
Overall gaps are measured.
Language differences remain.
Language differences remain.
Impact is hidden.
Impact is hidden.
Try it yourself
An example you can run in a temporary verification environment.
AWS公式AIF-C01 Domain 3.4・4.2の多言語評価、公平性、安全性を確認する。Expected result
一言語の平均性能を他言語へ外挿せず、言語ごとの影響を評価できる。Key points
- Multilingual
- Coverage
- Language metrics
Notes
- Environment: AWS公式AIF-C01試験ガイドとAWS公式ドキュメントの確認
- Command output formatting can vary slightly by distribution or tool version.
- Run the example in a temporary directory or process when possible.
Foundation review
Read the scope first
Check whether the command acts on the current shell, a new process, an existing process, or a file.
Verify the observable result
Use the supplied command and compare the output with the expected result.