Production AI Evaluation and Observability
Ý chính
AI production không ổn định chỉ bằng prompt tốt. Cần eval đại diện, trace được trajectory, đo retrieval/latency/cost, và buộc answer bám evidence.
Bài học từ các case
- Anthropic multi-agent: chấm outcome bằng rubric, LLM-as-judge calibrate với human, tracing decision pattern để debug agent không deterministic.
- Yelp Assistant: tách retrieval khỏi generation, đo latency từng stage, source selection để tránh search sai store, few-shot/prompt được version như code.
- Bits AI SRE: benchmark trên incident thật đã label, theo causal chain thay vì nhồi mọi telemetry vào context.
Mental model
production traffic / labeled cases
-> retrieval và tool trajectory
-> grounded output + citations
-> rubric eval + human calibration
-> trace/failure analysis
-> prompt/model/retrieval update