2022 - Retrieval-Augmented Multimodal Language Modeling - arXiv 2211.12561v2
Nguồn
- PDF gốc: 2022 - Retrieval-Augmented Multimodal Language Modeling - arXiv 2211.12561v2.pdf
- Vai trò trong CS224N: paper nối retrieval augmentation với multimodal language modeling.
Câu hỏi trung tâm
Retrieval có thể giúp multimodal model tận dụng memory ngoài tham số khi xử lý/generate text-image không?
Kiến thức cốt lõi
- Multimodal models cần xử lý thông tin không chỉ trong text mà cả image/visual context.
- Retrieval augmentation thêm nguồn tri thức hoặc examples liên quan ngoài parametric memory.
- Cách này mở rộng intuition của RAG sang multimodal setting.
- Challenge gồm representation chung, retrieval relevance và cách fuse retrieved items vào generator.
- Source này nằm trong trục RAG và multimodal CS224N.
Cơ chế / công thức / kiến trúc
multimodal input/query
-> retrieve text/image/multimodal neighbors
-> condition generator trên retrieved context
-> sinh output multimodal hoặc text grounded hơnKhi áp dụng
- Dùng khi parametric model thiếu tri thức visual cụ thể.
- Retriever phải hiểu cross-modal similarity, không chỉ text overlap.
- Cần đánh giá cả retrieved evidence và generated answer.
Kết quả / bằng chứng đáng giữ
- Title nêu retrieval-augmented multimodal language modeling.
- CS224N đặt paper này cạnh RAG và multimodal generation.
- Nó mở rộng vấn đề provenance/update knowledge sang không gian multimodal.
Cách hiểu bằng lời của tôi
Nếu RAG là “mở sách” cho text LM, multimodal RAG là “mở cả thư viện ảnh-văn bản” cho model nhìn và nói.
Câu hỏi review
- Multimodal retrieval khác text retrieval ở đâu?
- Retrieved visual context có thể giúp generation như thế nào?
- Evaluation cần đo thêm thành phần nào?