2017 - Attention Is All You Need - arXiv 1706.03762v7

Nguồn

Vấn đề paper giải quyết

Sequence transduction trước đây dựa vào recurrent/convolutional networks và attention. Paper đặt câu hỏi: có thể bỏ recurrence và convolution hoàn toàn, chỉ dùng attention để xây encoder-decoder không?

Đóng góp chính

  • Đề xuất Transformer dựa hoàn toàn trên attention mechanisms.
  • Dùng multi-head self-attention để trộn thông tin giữa token.
  • Kết hợp positional encoding, feed-forward layers, residual connections và normalization.
  • Cho thấy kiến trúc attention-only có thể đạt kết quả mạnh cho machine translation và train hiệu quả hơn nhờ parallelization.

Cơ chế Transformer

input embeddings + positional encoding
-> encoder stack: self-attention + feed-forward
-> decoder stack: masked self-attention + cross-attention + feed-forward
-> output distribution

Attention lõi:

Vì sao quan trọng với CS224N

Đây là paper trung tâm của khoá. Từ Lecture 05 trở đi, hầu hết nội dung như BERT, GPT, pretraining, RAG, agents và multimodal models đều là biến thể hoặc hệ sinh thái xung quanh Transformer.

Hạn chế / câu hỏi

  • Self-attention có cost theo sequence length.
  • Cần positional signal vì attention không tự có thứ tự.
  • Kiến trúc mở đường cho scaling nhưng cũng tạo bài toán memory/compute lớn.

Câu hỏi review

  1. Transformer bỏ recurrence bằng cách nào?
  2. Vì sao cần scale bằng ?
  3. Masked self-attention khác encoder self-attention ở đâu?
  4. Multi-head attention đem lại lợi ích gì?