2026-08-19 - Attention Is All You Need - Gemini Notebook Workflow

Ranh giới

Đây là working note scaffold từ paper note/PDF. Các phần closed-book recall và oral exam vẫn để trống cho bạn tự trả lời; không coi note này là bằng chứng đã đọc xong paper.

Setup

  • Paper: Attention Is All You Need
  • PDF: Attention Is All You Need.pdf
  • Paper note chính: Attention Is All You Need
  • Gemini Notebook / NotebookLM URL:
  • Mục tiêu buổi đọc: hiểu vì sao Transformer bỏ recurrence/convolution, cách scaled dot-product attention và multi-head attention hoạt động, và Table 2-3 chứng minh gì.
  • Phần cần đọc trước: Abstract, Introduction, Figure 1-2, Sections 3.1-3.5, Table 1-3, Conclusion.
  • PDF count đã kiểm tra: 15 trang.

Từ điển khái niệm nhanh

Khái niệmĐịnh nghĩa ngắn trong paper nàyVì sao quan trọngLink
Sequence transductionBài toán biến input sequence thành output sequence, ví dụ machine translation.Đây là task mà Transformer được đánh giá chính trong paper.Transformer
TransformerEncoder-decoder architecture chỉ dùng attention và feed-forward layers, không dùng recurrence/convolution.Đây là đóng góp kiến trúc trung tâm của paper.Transformer
Self-attentionMỗi token tạo representation bằng cách attend tới các token khác trong cùng sequence.Thay thế recurrence/convolution để kết nối dependency xa với path length ngắn.Self-Attention
Scaled dot-product attentionAttention tính bằng rồi nhân với .Scaling giúp logits ổn định hơn khi dimension key lớn.Self-Attention
Multi-head attentionChạy nhiều attention heads song song trên các projection khác nhau rồi concat lại.Cho model học nhiều kiểu quan hệ token-token ở các subspace khác nhau.Multi-Head Attention
Encoder-decoder attentionDecoder queries attend tới encoder outputs làm keys/values.Kết nối output đang sinh với input sentence trong translation.Cross-Attention
Positional encodingVector vị trí cộng vào token embedding để cung cấp thứ tự sequence.Attention-only model không tự có notion về vị trí nếu thiếu positional signal.Positional Embeddings
Position-wise FFNCùng một feed-forward network áp dụng độc lập cho từng position.Thêm nonlinear transformation sau attention mà vẫn giữ parallelism theo position.Transformer
Residual connection + LayerNormSkip connection quanh sublayer rồi chuẩn hóa representation.Giúp train stack encoder/decoder sâu ổn định hơn.Transformer, Layer Normalization
BLEUMetric đánh giá machine translation dựa trên n-gram overlap với reference.Là metric chính trong Table 2 để so sánh WMT results.BLEU

Phase 1 — Paper Map

Prompt gửi Gemini

Do not summarize the paper in detail yet. Create a structural map of this paper and identify problem, motivation, gap, contributions, pipeline, components, equations, datasets, baselines, metrics, experiments, ablations, and limitations. For every item, point to the relevant section, figure, table, or equation. The purpose is to tell me WHERE to read, not to replace my reading.

Paper map — scaffold từ nguồn

  • Problem: sequence transduction models mạnh trước đó dựa vào recurrence/convolution, gây sequential bottleneck và khó song song hóa trong training. PDF tr. 2
  • Motivation: attention đã giúp nối dependency xa; paper hỏi liệu attention có thể trở thành cơ chế chính thay RNN/CNN không. PDF tr. 1
  • Gap: chưa có architecture transduction mạnh chỉ dựa vào attention, không recurrence/convolution. PDF tr. 2
  • Main idea: Transformer encoder-decoder dùng stacked multi-head self-attention, encoder-decoder attention, FFN, residual connections, layer norm và positional encoding.
  • Main contributions: attention-only architecture, scaled dot-product attention, multi-head attention, positional encoding, WMT results với training cost thấp.
  • Important figure: Figure 1 architecture; Figure 2 scaled dot-product attention và multi-head attention. PDF tr. 3, PDF tr. 4
  • Important equations: Eq. 1 attention; Eq. 2 multi-head attention; Eq. 3 FFN; positional encoding formulas. PDF tr. 4, PDF tr. 5, PDF tr. 6
  • Main result table: Table 2 WMT 2014 BLEU/cost. PDF tr. 8
  • Ablation table: Table 3 architecture variations. PDF tr. 9
  • Limitations: self-attention quadratic in sequence length; autoregressive generation vẫn tuần tự; future work nhắc local/restricted attention cho image/audio/video. PDF tr. 6, PDF tr. 10

Chỗ cần đọc trước

  • Abstract + Introduction
  • Figure 1 architecture
  • Eq. 1-3
  • Table 1 complexity/path length
  • Table 2 main results
  • Table 3 ablation
  • Conclusion/future work

Phase 2 — Pass 1 Recall

Closed-book recall của tôi

Problem

Why does it matter?

Research gap

Main idea

Main contribution

Main result

Prompt kiểm tra recall

I have completed the first pass of the paper.
 
Here is my understanding:
 
[PASTE MY NOTES]
 
Compare my understanding against the paper. Return what is correct, inaccurate, missing, confusing, and which sections/citations I should revisit. Do not rewrite the entire paper for me.

Phase 3 — Problem / Motivation / Gap

MụcDiễn giải bằng lời của tôiEvidence / citation
General problemSequence transduction cần map input sequence sang output sequence, ví dụ machine translation.PDF tr. 2
Why it mattersRNN phải tính tuần tự theo position, làm train chậm và dependency xa khó học hơn.PDF tr. 2
What prior work solvesRNN/CNN encoder-decoder + attention đã đạt kết quả mạnh trong translation.PDF tr. 2
What prior work fails to solveRecurrence/convolution vẫn gây sequential operations hoặc path length lớn hơn self-attention.PDF tr. 6
Exact research gapThiếu một architecture transduction attention-only mạnh, song song hóa tốt, không recurrence/convolution.PDF tr. 1
Hypothesis / intuitionAttention đủ để model hóa quan hệ giữa token; positional encoding bù thông tin thứ tự.PDF tr. 6
Contribution addressing the gapTransformer dùng attention + FFN + positional encoding trong encoder-decoder stack.PDF tr. 3

Câu hỏi tự kiểm tra

  • Paper bỏ recurrence/convolution ở đâu, nhưng vẫn giữ autoregressive decoding ở đâu?
  • Vì sao self-attention có path length nhưng complexity ?
  • Vì sao positional encoding là bắt buộc trong attention-only architecture?

Phase 4 — Method / Architecture

Tôi tự vẽ trước

Source tokens
-> token embeddings + positional encodings
-> encoder stack N=6
   -> multi-head self-attention
   -> feed-forward network
-> encoder memory
Shifted target tokens
-> token embeddings + positional encodings
-> decoder stack N=6
   -> masked self-attention
   -> encoder-decoder attention
   -> feed-forward network
-> linear + softmax
-> next-token distribution

Component map

ComponentInputOperationOutputPurposeEvidence
Embedding + positional encodingtoken ids, positionssum token embedding and position encodingposition-aware token vectorsinject order without recurrencePDF tr. 6
Encoder self-attentionencoder statesQ/K/V from same sequencecontextualized source stateseach source token attends to all source tokensPDF tr. 5
Decoder masked self-attentionshifted target statescausal mask prevents future attentionprefix-aware target statesautoregressive trainingPDF tr. 5
Encoder-decoder attentiondecoder queries, encoder keys/valuesattention over source memorysource-conditioned target statesalign/condition target generationPDF tr. 5
Multi-head attentionQ, K, Vproject to multiple heads, concat, projectmixed attention representationmultiple relation subspacesPDF tr. 5
Position-wise FFNeach position vectortwo linear layers + ReLUtransformed vectorper-token nonlinear processingPDF tr. 5
Residual + layer normsublayer input/outputstabilized hidden statetrain deep stackPDF tr. 3

Điều tôi vẫn chưa hiểu

  • Table 1 complexity có giả định gì về , , ?
  • Vì sao learned positional embeddings gần ngang sinusoidal trong ablation nhưng paper vẫn chọn sinusoidal?

Phase 5 — Section Recall

Section 3.1 — Encoder and Decoder Stacks

  • Input: source/target embeddings cộng positional encodings.
  • Process: encoder dùng self-attention + FFN; decoder thêm masked self-attention và encoder-decoder attention.
  • Output: encoder memory và decoder hidden states để dự đoán next token.
  • Purpose: thay recurrent/convolutional blocks bằng attention-based blocks nhưng vẫn giữ encoder-decoder formulation.

Section 3.2 — Attention

  • Input: query, keys, values.
  • Process: scaled dot-product attention và multi-head projections.
  • Output: weighted sum of values, concat multi-head output.
  • Purpose: cho mỗi token truy cập trực tiếp token liên quan ở toàn sequence.
  • Still unclear: multi-head attention học các quan hệ khác nhau bằng cơ chế nào nếu không có supervision riêng cho từng head?

Phase 6 — Equations

Prompt operational equation walkthrough

Walk me through the key equations or formal blocks in this paper.
 
For each equation/block, explain:
1. Input: what variables or objects go into it.
2. Output: what it produces.
3. Where it is used in the training/inference pipeline.
4. What behavior it encourages.
5. What would likely break or become weaker if removed.
6. Which table, figure, ablation, or result supports its usefulness.
 
Do not summarize the whole paper. Focus only on operational understanding of the equations and formal mechanisms.
Eq.Dùng để làm gì?Biến chínhBehavior được khuyến khíchEvidence / ablationStatus
1Tính scaled dot-product attentionattend theo similarity nhưng scale để softmax ổn định, tránh softmax quá sắc khi lớnPDF tr. 4source-checked
2Multi-head attentionhọc nhiều attention subspaces song song; single-head kém multi-head khoảng 0.9 BLEU trong Table 3PDF tr. 5, PDF tr. 9source-checked
3Position-wise FFNnonlinear transform độc lập từng position sau attentionPDF tr. 5source-checked
PESin/cos positional encodinginject thứ tự token; learned positional embeddings gần như ngang sinusoidal trong Table 3 row EPDF tr. 6, PDF tr. 9source-checked
LR schedulewarmup + inverse sqrt decaytăng learning rate lúc đầu để ổn định, sau đó decay theo inverse square rootPDF tr. 7source-checked

Phase 7 — Loss Functions

Training objective
├── token-level cross-entropy / likelihood -> predict next target token
├── label smoothing epsilon_ls=0.1 -> regularize output distribution
└── dropout -> regularize residual/attention/embedding paths
Loss / regularizerEquationInputsTrains componentBehaviorWeightAblation
Cross-entropy / NLLnot expanded as a named equationpredicted next-token distribution + target tokenfull encoder-decodermaximize translation likelihoodstandard objectivenot ablated as core
Label smoothingdescribed in training sectiontarget distributionoutput probabilitiesreduce overconfidence, improve BLEU/perplexity trade-offnot isolated in Table 3
Dropouttraining setupresidual/attention/embedding pathsfull modelregularization baseTable 3 shows no dropout hurts

Phase 8 — Experiments

ExperimentResearch questionDatasetBaselinesMetricTable/FigureMain resultCaveat
Main translationTransformer có đạt SOTA/cost tốt hơn không?WMT 2014 EN-DE, EN-FRprior NMT modelsBLEU, training costTable 2Big: 28.4 EN-DE, 41.8 EN-FRreported, not reproduced
Architecture ablationhead count, dimensions, dropout, PE ảnh hưởng thế nào?EN-DE devTransformer variantsBLEU/perplexityTable 3multi-head, dropout, bigger model matterdev setting only
Complexity comparisonself-attention trade-off gì so với recurrent/convolution?theoretical layer comparisonrecurrent/convolution/separable convoperations, complexity, path lengthTable 1self-attention path length/sequential ops tốt nhưng assumes sequence length regime
Parsing generalizationTransformer có generalize ngoài MT không?WSJ parsingparsing baselinesF1Table 491.3 WSJ-only, 92.7 semi-supervisedsmall experiments, not deeply tuned
Attention visualizationheads học dependency nào?example sentencesvisualizationqualitativeAppendix figuressome heads track long-distance/syntactic relationsillustrative, not proof

Protocol fingerprint

  • Dataset and split: WMT 2014 English-German, English-French; parsing on WSJ.
  • Tokenization: EN-DE BPE shared vocab about 37K; EN-FR word-piece 32K.
  • Scenario / label space: supervised sequence transduction / machine translation.
  • Backbone: Transformer base/big, no recurrence/convolution.
  • Model base: , , , , .
  • Model big: , , .
  • Training: base 100K steps, big 300K steps.
  • Hardware: 8 NVIDIA P100 GPUs.
  • Optimizer: Adam .
  • Metric: BLEU for translation; F1 for parsing.
  • Result type: reported/observed from paper, not reproduced local.

Phase 9 — Claim → Evidence

ClaimWhere claim appearsExperimentEvidenceMy judgmentCaveat
Attention-only architecture can replace recurrence/convolution for MT.Abstract/IntroWMT Table 2Transformer big reaches 28.4 BLEU EN-DE and 41.8 EN-FR. PDF tr. 8Strong for WMT protocolnot proof for every sequence task
Transformer trains faster / with lower cost than prior models.Abstract/ResultsTable 2 cost comparisonreported lower training cost than listed baselines. PDF tr. 8Strong reported evidencehardware/framework differences matter
Multi-head attention matters.AblationTable 3single head underperforms multi-head variants. PDF tr. 9Supportedablation dev setting
Positional encoding is needed, but learned vs sinusoidal similar.Method/AblationTable 3 row Elearned positional embeddings give similar result to sinusoidal. PDF tr. 9Supportedextrapolation claim is hypothesis
Self-attention has better path length but quadratic cost.Table 1/methodtheoretical comparisonself-attention sequential ops/path length O(1), complexity . PDF tr. 6Strong theoretical framinglong sequence cost remains limitation

Phase 10 — Ablation Study

ComponentIntended purposeWith componentWithout / changed componentDifferenceConclusion justifiedNot justified
Multi-head attentionmultiple relation subspacesbaseline BLEU in Table 3single head lowerabout -0.9 BLEU in dev setupmulti-head helpsexact optimal head count universal
Attention key/value dimensioncapacity per headbaselinesmaller variants lowervariesdimension mattersbigger always better without cost
Dropoutregularizationbaselineno dropout worseclear degradationdropout importantdropout value universal
Model sizecapacitybase/big variantssmaller variants lowerbigger generally bettercapacity helpsscaling law proven
Positional encoding typeorder signalsinusoidallearned PE similarsmall/no major differenceboth viable in setupsinusoidal always superior

Ranh giới evidence

Các dòng trên là reported/observed từ paper và paper note chính, chưa phải kết quả reproduce local.

Phase 11 — Critical Reading

  • Strongest contribution: turning attention into the main sequence modeling primitive and showing strong WMT results with practical training speed.
  • Weakest part: evidence for interpretability of heads is qualitative; state-of-the-art claim is protocol-bound.
  • Main assumption: machine translation setup with moderate sequence lengths where attention is affordable.
  • Alternative explanation: gains come from architecture plus engineering/training recipe, not attention formula alone.
  • Missing experiment: broader long-sequence tasks, non-autoregressive generation, systematic head interpretability.
  • Generalization risk: image/audio/video require restricted/local attention, as authors note.
  • Reproducibility risk: modern exact reproduction depends on preprocessing/tokenization, batching, and framework details.

Phase 12 — Reproduction Check

ItemStatusDetailMissing detail / risk
Dataset and splitClearly specifiedWMT 2014 EN-DE/EN-FR, WSJ parsingexact preprocessing scripts needed
PreprocessingPartially specifiedBPE/word-piece vocab sizesexact vocab/training pipeline
Input representationClearly specifiedembeddings + positional encodings
Model / backboneClearly specifiedTransformer base/big
ArchitectureClearly specifiedencoder/decoder stacks, attention, FFN
Training procedurePartially specifiedsteps, optimizer, schedulebatching/token batching details
Sampling procedureMissing/Not applicablesupervised MTdecoding details for BLEU need scripts
Memory / replay strategyNot applicableno continual memory
Loss functionsPartially specifiedCE with label smoothingexact implementation details
OptimizerClearly specifiedAdam params
Learning rateClearly specifiedwarmup + inverse sqrt
Batch sizePartially specifiedapproximate tokens/batch in paperexact batching can vary
EpochsNot specified as epochstraining by steps
HyperparametersClearly specified
Random seedsMissingnot reported in scaffoldvariance risk
Evaluation protocolPartially specifiedBLEU newstest2014tokenization BLEU script matters
Inference procedurePartially specifiedautoregressive decodingbeam/search settings need full paper/code

Phase 13 — Completeness / Oral Exam

  • Giải thích bottleneck sequential của RNN.
  • Vẽ encoder layer và decoder layer.
  • Giải thích Eq. 1 scaling .
  • Phân biệt self-attention, masked self-attention, encoder-decoder attention.
  • Giải thích positional encoding.
  • Đọc Table 1 complexity/path length.
  • Đọc Table 2 BLEU/cost và nêu caveat.
  • Đọc Table 3 ablation.
  • Nêu limitation quadratic attention và autoregressive decoding.

Prompt oral exam

Act as my PhD advisor. Quiz me on Attention Is All You Need one question at a time. Start from problem/gap, then architecture, equations, experiments/ablation, limitations, and finally ask me to transfer the idea to another sequence modeling setting. Do not reveal the ideal answer before I attempt it.

Final Paper Note Handoff

Chỉ chuyển sang paper note chính những ý đã tự kiểm tra lại bằng PDF citation.

Ý cần chuyển sang paper note

  • Problem: sequential bottleneck của RNN/CNN.
  • Method overview: encoder-decoder attention-only stack.
  • Important equations: attention, multi-head, FFN, positional encoding.
  • Protocol fingerprint: WMT 2014, BPE/word-piece, base/big, Adam schedule.
  • Main results: Table 2 BLEU/cost.
  • Ablation: head count, model size, dropout, positional encoding.
  • Limitations: quadratic attention, autoregressive generation, qualitative head visualization.
  • Concepts: Transformer, Self-Attention, Multi-Head Attention, Positional Embeddings, Encoder-Decoder Architecture.

Liên kết