Evaluate a new recommendation model.
Compare test accuracy with real-user effects.
Choose two practices.
Use fixed offline data for candidate comparison and monitor production latency, errors, behavior, complaints, and safety; neither replaces the other.
Detailed explanation
Candidates are compared before full exposure.
Candidates are compared before full exposure.
Operational and user impact are measured.
Operational and user impact are measured.
Distribution and behavior can differ.
Distribution and behavior can differ.
The comparison loses credibility.
The comparison loses credibility.
Try it yourself
An example you can run in a temporary verification environment.
AWS公式AIF-C01 Domain 3.3のモデル評価、運用指標、責任あるAIを確認する。Expected result
開発時指標と本番影響指標を分け、段階展開で確認できる。Key points
- Offline evaluation
- Production monitoring
- User impact
Notes
- Environment: AWS公式AIF-C01試験ガイドとAWS公式ドキュメントの確認
- Command output formatting can vary slightly by distribution or tool version.
- Run the example in a temporary directory or process when possible.
Foundation review
Read the scope first
Check whether the command acts on the current shell, a new process, an existing process, or a file.
Verify the observable result
Use the supplied command and compare the output with the expected result.