Production LLM System Design

Luồng hệ thống

Production LLM system không chỉ là model endpoint. Một request thật thường đi qua auth, context assembly, tokenization, inference, streaming, tool execution, safety checks, memory retrieval, logging và monitoring.

user/client
-> gateway / auth
-> context engineering
-> model serving / inference
-> streaming or action loop
-> safety and output handling
-> observability and eval feedback

Các trục thiết kế

TrụcCâu hỏi cần hỏiConcept liên quan
LatencyUser cần thấy token/action nhanh đến mức nào?LLM Inference Engineering, AI Model Serving
ContextModel cần thấy gì và bỏ gì?Context Engineering, LLM Memory
HardwareBottleneck là compute hay memory bandwidth?AI Hardware Accelerator, KV Cache
Tool/actionModel có quyền làm gì ngoài việc trả lời?Tool Use, Excessive Agency
SecurityUntrusted content đi vào đâu, outbound channel ở đâu?LLM Security, Prompt Injection
QualityLàm sao biết version mới tốt hơn?LLM Evaluation, Model Benchmarking
ObservabilityKhi answer sai thì sai ở retrieval, tool, prompt hay generation?LLM Observability, Agent Tracing
RoutingRequest nên chạy bằng model nào?Model Router, AI Model Serving

Bài học từ case study

  • ChatGPT: trải nghiệm đơn giản che giấu một pipeline dài, trong đó context, tool, safety, memory và streaming cùng tham gia.
  • Cursor: code completion là workload latency cực thấp; codebase indexing giúp chat agent lấy context rộng hơn mà không gửi toàn bộ repo mỗi lần.
  • OpenAI voice AI: với realtime audio, network/protocol architecture có thể quyết định cảm giác hội thoại nhiều như tốc độ model.
  • Codex/ChatGPT Work: agent efficiency là tối ưu end-to-end qua harness, API và inference, trong đó mỗi lớp tránh lặp lại work đã làm.
  • Anthropic multi-agent: agent production cần trace decision pattern, checkpoint/retry và eval outcome bằng rubric vì cùng prompt có thể đi nhiều path khác nhau.
  • Yelp Assistant: production assistant nên tách retrieval/source selection khỏi final generation, dùng SSE để giảm perceived latency, và đánh giá tone/grounding bằng example/rubric.
  • Bits AI SRE: AI ops agent hiệu quả khi follow causal evidence theo hypothesis, không nhồi mọi telemetry vào context rồi summarize.

Ghi nhớ

Thiết kế LLM production là bài toán hệ thống. Model quality là một phần, nhưng trải nghiệm cuối phụ thuộc vào context đúng, inference nhanh, quyền tool hẹp, guardrail nhiều lớp và eval có thể bắt regression.

Liên kết