AI Search and Recommendation Systems

Mental model

AI search là chuỗi quyết định về vị trí đặt intelligence:

  • trước retrieval: Query Understanding, rewrite, taxonomy mapping;
  • trong retrieval: Semantic Search, Two-Tower Retrieval, ANN;
  • quanh retrieval: guardrail, hard filter, domain context;
  • sau retrieval: ranking, reranking, business logic;
  • ngoài hot path: graph enrichment, cache, offline feature generation.

Ba độ sâu tích hợp LLM

Độ sâuVí dụKhi hợp
LLM ở offline/peripheryDoorDash enrich graph, parse query có output ràng buộcĐã có taxonomy/knowledge graph mạnh
LLM ở query understandingInstacart dùng RAG/cache cho head query và fine-tuned model cho tail queryMuốn hợp nhất nhiều model query cũ
LLM là embedding backboneUber Eats dùng fine-tuned Qwen trong two-tower retrievalCần shared semantic space đa domain/ngôn ngữ

Bài học từ case study

  • Amazon COSMO: LLM tốt để sinh hypothesis commonsense, nhưng production cần filter, annotation, classifier và graph serving.
  • Instacart Postgres/pgvector: đưa compute gần data có thể giảm latency hơn việc thêm một service chuyên dụng.
  • Airtable Milvus: data shape quyết định kiến trúc, nhất là tenant isolation và hot/cold pattern.
  • DoorDash/Instacart/Uber Eats: model choice ít quan trọng hơn câu hỏi LLM nên nằm ở đâu trong stack.
  • LinkedIn/Meta/YouTube: cùng chuyển retrieval từ engagement proxy sang meaning, nhưng chọn consolidation, specialization funnel hoặc generative retrieval tùy data shape và rollback/cost trade-off.

Ghi nhớ

Hybrid là mặc định. Keyword search, vector search, knowledge graph, cache, filter và reranker vẫn cùng tồn tại. LLM hữu ích nhất khi nó được đặt đúng vị trí: hiểu intent, tạo representation, sinh tri thức có kiểm chứng, hoặc thu hẹp output vào taxonomy an toàn.

Liên kết