Production AI Evaluation and Observability

Ý chính

AI production không ổn định chỉ bằng prompt tốt. Cần eval đại diện, trace được trajectory, đo retrieval/latency/cost, và buộc answer bám evidence.

Bài học từ các case

  • Anthropic multi-agent: chấm outcome bằng rubric, LLM-as-judge calibrate với human, tracing decision pattern để debug agent không deterministic.
  • Yelp Assistant: tách retrieval khỏi generation, đo latency từng stage, source selection để tránh search sai store, few-shot/prompt được version như code.
  • Bits AI SRE: benchmark trên incident thật đã label, theo causal chain thay vì nhồi mọi telemetry vào context.

Mental model

production traffic / labeled cases
-> retrieval và tool trajectory
-> grounded output + citations
-> rubric eval + human calibration
-> trace/failure analysis
-> prompt/model/retrieval update

Liên kết