Mixture of Experts
Định nghĩa
Mixture of Experts (MoE) là kiến trúc gồm nhiều expert functions và một gating/router quyết định expert nào đóng góp cho từng input. Mục tiêu là tăng capacity hoặc specialization mà không nhất thiết chạy toàn bộ experts cho mọi sample.
Công thức cơ bản
Với experts và score functions :
Dense MoE trộn mọi expert. Sparse MoE chỉ giữ top- scores:
Cách hiểu bằng lời của tôi
MoE tách hai câu hỏi:
- “Ai biết xử lý input này?” — router/gate.
- “Xử lý như thế nào?” — expert.
Capacity có thể lớn vì có nhiều experts, nhưng compute mỗi token/sample vẫn giới hạn nếu chỉ route tới top-.
Dense và sparse
| Dạng | Experts chạy mỗi input | Điểm mạnh | Rủi ro |
|---|---|---|---|
| Dense mixture | Gần như tất cả | Smooth combination | Compute tăng theo số experts |
| Sparse MoE | Top- | Tăng capacity mà compute thấp hơn | Router imbalance, communication overhead |
Attention nhìn như MoE
WAVE/WAVE++ chỉ ra một attention-head output tại position :
có cùng hình thức với MoE:
- value-transformed token đóng vai expert output ;
- query-key score đóng vai gate ;
- mỗi query position dùng gate riêng nhưng chia sẻ expert set.
Vì vậy một attention head có thể được nhìn như nhiều gated mixtures, không phải Transformer đã cài một sparse-MoE feed-forward layer theo nghĩa triển khai phổ biến. WAVE++ PDF, tr. 6 WAVE++ PDF, tr. 7
Prefix tuning như thêm experts
Khi thêm prefix key/value:
- prefix value tạo expert output mới;
- prefix key tạo score mới với mỗi query;
- attention trộn prefix experts với pretrained token experts.
Lens này giải thích vì sao Prompt Pool giống sparse expert pool: query chọn top- prompts/prefix experts theo keys.
Nhưng prefix experts trong formulation này là simple offset/constant functions. Chúng không có capacity tùy ý như MLP experts, nên “tương đương MoE” là tương đương về dạng weighted mixture, không phải mọi thuộc tính kiến trúc.
Router là điểm nghẽn
MoE chỉ mạnh khi routing tốt. Failure modes:
- load imbalance: một vài experts nhận hầu hết tokens;
- expert collapse/dead experts;
- router noise/instability;
- capacity overflow và dropped tokens;
- communication cost giữa devices;
- semantic specialization không xuất hiện dù loss tốt;
- task/input shift làm router chọn sai experts.
Trong continual learning, router còn có thể quên hoặc prompt keys bị drift, nên routing accuracy cần được đo riêng.
MoE trong continual learning
Một expert/pool riêng theo task giảm interference:
task cũ -> freeze experts cũ
task mới -> thêm experts mớiNhưng capacity tăng theo task và inference không biết sẵn expert nào đúng. WAVE++ xử lý bằng cascade voting rồi route bên trong task pool.
Khi áp dụng
MoE hữu ích khi:
- data có nhiều modes/domains/tasks;
- muốn tăng model capacity nhưng giữ active compute có giới hạn;
- có đủ traffic/data để experts chuyên môn hóa;
- hệ thống chịu được routing và distributed communication complexity.
Không nên dùng chỉ vì “nhiều experts nghe mạnh”: với dataset nhỏ, router có thể học kém và overhead lớn hơn lợi ích.
Câu hỏi review
- Gate và expert có vai trò gì?
- Sparse MoE giảm compute bằng cách nào?
- Attention giống MoE ở dạng toán nào?
- Prefix token trở thành expert theo cách hiểu nào?
- Vì sao task-specific experts không tự giải quyết continual learning?
Gợi ý trả lời
- Gate chọn/trộn; expert biến đổi input.
- Chỉ active top- experts thay vì tất cả.
- Output là weighted sum các value/expert outputs với normalized query-key/gating scores.
- Prefix key tạo gate score, prefix value tạo output được trộn vào attention.
- Phải suy đúng task/expert ở test, shared components vẫn có thể quên, và memory tăng theo task.