Multimodal Annotation Fusion
Định nghĩa
Multimodal Annotation Fusion là bước hợp nhất output từ nhiều model/modalities thành một representation chung có thể index và query.
Cách hiểu bằng lời của tôi
Trong video search, mỗi model nhìn footage theo cách khác nhau: character model trả label, scene model trả embedding, dialogue model trả transcript có timestamp. Nếu không fuse chúng vào cùng mốc thời gian, query kiểu “nhân vật X ở địa điểm Y nói câu Z” sẽ rất khó chạy nhanh.
Cơ chế từ nguồn Netflix
raw model outputs
-> map interval liên tục thành bucket thời gian
-> intersect annotation trong cùng bucket
-> ghi record fused theo asset + second
-> index parent/child document để query cross-annotationQuyết định thiết kế
- Bucket nhỏ tăng precision nhưng tăng số record.
- Fusion offline giữ ingestion nhẹ và query nhanh, đổi lại freshness kém hơn.
- Update record theo bucket giúp thêm model mới mà không tạo duplicate timeline.
- Parent/child index giúp exact label, transcript và vector embedding cùng nằm trong một context truy vấn.