2026-09-23 - Mean Teacher - Gemini Notebook Workflow
Ranh giới
Đây là working note scaffold từ PDF/paper note để hỗ trợ đọc. Các phần factual dùng citation từ paper; phần recall cá nhân để trống và không đánh dấu đã đọc xong.
Setup
- Paper: Mean Teachers are Better Role Models
- PDF: Mean Teachers are Better Role Models.pdf
- Paper note chính: Mean Teachers are Better Role Models
- Gemini Notebook / NotebookLM URL:
- Mục tiêu buổi đọc: hiểu EMA teacher, consistency regularization, vai trò noise/augmentation, hyperparameter EMA decay và liên hệ với continual KD.
- Phần cần đọc trước: Abstract, Figure 1-2, Section 3, Figure 4, Conclusion, Appendix B-C.
- PDF count đã kiểm tra: 16 trang.
Từ điển khái niệm nhanh
| Khái niệm | Định nghĩa ngắn trong paper này | Vì sao quan trọng | Link |
|---|---|---|---|
| Mean Teacher | Teacher model có trọng số là exponential moving average của student weights. | Là đóng góp chính: target ổn định hơn teacher cùng trọng số với student. | Mean Teachers are Better Role Models |
| EMA weights | Cập nhật teacher bằng . | Tạo teacher mượt hơn, cập nhật theo step thay vì theo epoch. | Model Distillation |
| Consistency regularization | Ép prediction của student và teacher nhất quán khi có noise/augmentation. | Cho phép dùng unlabeled data như regularizer. | Knowledge Distillation |
| Temporal Ensembling | EMA trên prediction của từng training example. | Baseline trực tiếp; Mean Teacher thay prediction-average bằng weight-average. | Knowledge Distillation |
| Confirmation bias | Teacher tự sinh target sai và student học theo nếu target kém. | Motivation cho việc cải thiện quality của teacher target. | Model Distillation |
| MSE consistency cost | Paper chủ yếu dùng MSE giữa prediction student và teacher. | Là loss chính cho consistency; Appendix C so với KL-divergence. | Knowledge Distillation |
Phase 1 — Paper Map
Paper map — scaffold từ nguồn
- Problem: Semi-supervised learning cần tận dụng unlabeled data để giảm overfitting khi label đắt; consistency target trên unlabeled data dễ bị confirmation bias nếu teacher yếu. PDF tr. 1, PDF tr. 2
- Motivation: Temporal Ensembling cải thiện target bằng EMA prediction nhưng mỗi example chỉ cập nhật target một lần mỗi epoch, nên chậm và khó mở rộng với dataset lớn. PDF tr. 2
- Main idea: Dùng teacher có weights là EMA của student weights; student học supervised classification loss trên labeled data và consistency loss với teacher output trên labeled/unlabeled data. PDF tr. 3
- Important figure: Figure 2 mô tả student/teacher, noise riêng , classification cost, consistency cost và EMA update. PDF tr. 3
- Important equations: Consistency cost ; EMA update . PDF tr. 3
- Main result tables: Table 1-2 cho SVHN/CIFAR-10 ConvNet; Table 4 cho ResNet Mean Teacher trên CIFAR-10 và ImageNet. PDF tr. 4, PDF tr. 7
- Ablation: Figure 4 kiểm tra noise, teacher dropout, consistency weight, EMA decay, dual output, MSE vs KL. PDF tr. 6
- Limitation/caveat: Paper thuộc semi-supervised image classification, không phải continual learning; khi mượn cho CRE/TAPTA cần nói rõ đây là nguồn cơ chế EMA teacher, không phải baseline CRE.
Chỗ cần đọc trước
- Abstract + Introduction
- Figure 1-2
- Section 3 experiments
- Figure 4 ablation
- Appendix B-C cho training details và MSE/KL
Phase 2 — Pass 1 Recall
Closed-book recall của tôi
Problem
Why does it matter?
Research gap
Main idea
Main contribution
Main result
Prompt kiểm tra recall
I have completed the first pass of the paper.
Here is my understanding:
[PASTE MY NOTES]
Compare my understanding against the paper.
Return what I understood correctly, what is inaccurate, important points I missed, confusing concepts, and source sections to revisit.
Do not rewrite the entire paper for me.Phase 3 — Problem / Motivation / Gap
| Mục | Diễn giải bằng lời của tôi | Evidence / citation |
|---|---|---|
| General problem | Deep models overfit khi label ít; unlabeled data cần được khai thác bằng regularization. | PDF tr. 1 |
| Why it matters | Gán nhãn thủ công đắt, còn model lớn cần nhiều data để học abstraction tốt. | PDF tr. 1 |
| What prior work solves | model và Temporal Ensembling dùng consistency target để tận dụng unlabeled data. | PDF tr. 2 |
| What prior work fails to solve | Temporal Ensembling cập nhật target theo example chậm, mỗi epoch một lần, nên feedback loop kém nhanh khi dataset lớn. | PDF tr. 2 |
| Exact research gap | Cần teacher target ổn định hơn student tức thời nhưng cập nhật nhanh hơn EMA prediction theo epoch. | PDF tr. 3 |
| Hypothesis / intuition | Averaging weights tạo teacher có representation trung gian và prediction target tốt hơn, từ đó consistency learning nhanh và ổn định hơn. | PDF tr. 3 |
| Contribution addressing the gap | Mean Teacher: teacher weights là EMA của student weights, dùng consistency loss với teacher output. | PDF tr. 3 |
Phase 4 — Method / Architecture
Tôi tự vẽ trước
Input image
-> augmentation/noise cho student và teacher
-> Student model theta
-> classification loss nếu có label
-> Teacher model theta' = EMA(theta)
-> consistency loss giữa prediction student và teacher
-> update theta bằng SGD/Adam
-> update theta' bằng EMA, không backprop qua teacherComponent map
| Component | Input | Operation | Output | Purpose | Evidence |
|---|---|---|---|---|---|
| Student model | noisy/augmented input | forward + gradient update | prediction | học supervised + consistency objective | PDF tr. 3 |
| Teacher model | noisy/augmented input | forward only với EMA weights | prediction | sinh target ổn định hơn | PDF tr. 3 |
| EMA update | student weights | teacher weights mới | average qua training steps | PDF tr. 3 | |
| Consistency cost | student/teacher predictions | thường dùng MSE | regularization signal | học invariance từ unlabeled data | PDF tr. 3 |
| Noise/augmentation | input/layer noise | random translation, flip, Gaussian noise, dropout | hai views khác nhau | tạo perturbation để consistency có ý nghĩa | PDF tr. 4 |
Điều tôi vẫn chưa hiểu
- EMA teacher khác checkpoint averaging cuối training ở điểm nào về gradient/training dynamics?
- Với CRE/TAPTA, EMA nên áp dụng lên backbone, prompt, router, hay teacher logits/prototypes?
Phase 6 — Equations
Prompt operational equation walkthrough
Walk me through the key equations or formal blocks in this paper.
For each equation/block, explain input, output, where it is used, what behavior it encourages, what weakens if removed, and which table/figure/ablation supports it.
Focus only on operational understanding of the equations and formal mechanisms.Equation queue
| Eq. | Dùng để làm gì? | Biến chính | Behavior được khuyến khích | Evidence / ablation | Status |
|---|---|---|---|---|---|
| Consistency cost giữa prediction của student và teacher dưới noise. | Prediction ổn định quanh data manifold. | Figure 2, Section 2/3; PDF tr. 3 | scaffold | ||
| Cập nhật teacher weights bằng EMA. | Teacher chậm hơn, mượt hơn, target ít nhiễu hơn. | Figure 4(d) sensitivity; PDF tr. 6 | scaffold | ||
| MSE consistency | Loss thực nghiệm chính cho consistency. | teacher/student probabilities | Penalize prediction mismatch trên labeled/unlabeled data. | Appendix C so MSE tốt hơn KL trong setting này; PDF tr. 15 | scaffold |
Phase 7 — Loss Functions
Total Loss
├── Classification cost -> học nhãn thật trên labeled examples
├── Consistency cost -> khớp prediction student với EMA teacher
├── Dual-output/logit MSE trick -> tách phần classification và consistency trong một số setting
└── Weight decay -> regularization| Loss | Equation | Inputs | Trains component | Behavior | Weight | Ablation |
|---|---|---|---|---|---|---|
| Classification cost | task CE | labeled input + one-hot label | student | học nhãn thật | theo setup | baseline supervised |
| Consistency cost | MSE hoặc KL family | student/teacher prediction | student | ổn định prediction dưới perturbation | ramp-up; SVHN/CIFAR/ImageNet khác nhau | Figure 4(c/f), Appendix C |
| EMA teacher | không backprop loss trực tiếp | student weights | teacher update only | target mượt | Figure 4(d) |
Phase 8 — Experiments
Experiment map
| Experiment | Research question | Dataset | Baselines | Metric | Table/Figure | Main result | Caveat |
|---|---|---|---|---|---|---|---|
| ConvNet SSL | Mean Teacher có hơn model/Temporal Ensembling không? | SVHN, CIFAR-10 | Supervised-only, , Temporal Ensembling, VAT | error rate | Table 1-2 | SVHN 250 labels: 4.35%; CIFAR-10 4000 labels ConvNet: 12.31%. | VAT vẫn mạnh ở vài setting. |
| Extra unlabeled data | Có tận dụng thêm unlabeled data không? | SVHN extra | model | error rate | Table 3/Figure 3 | Mean Teacher tốt hơn khi thêm unlabeled data. | vẫn cải thiện lâu hơn khi extra data rất lớn. |
| Ablation | Noise, EMA decay, consistency weight nhạy ra sao? | SVHN 250 labels | variants | validation error | Figure 4 | Noise/augmentation cần thiết; EMA decay và consistency weight có range tốt. | One-factor-at-a-time ablation. |
| ResNet/ImageNet | Có scale sang architecture/dataset lớn không? | CIFAR-10, ImageNet 2012 | SOTA | error rate | Table 4 | CIFAR-10 4000 labels 6.28%; ImageNet 10% labels 9.11%. | ImageNet dùng validation vì test set không public. |
Protocol fingerprint
- Dataset and split: SVHN, CIFAR-10, ImageNet 2012; nhiều mức label fraction.
- Scenario / label space: semi-supervised classification, unlabeled examples dùng consistency.
- Backbone: 13-layer ConvNet; ResNet/ResNeXt ở experiment lớn.
- Frozen/trainable components: student trainable; teacher updated by EMA, không backprop.
- Seeds / number of runs: 10 runs cho nhiều SVHN/CIFAR setting; ImageNet 2 runs.
- Metric and averaging: classification error percentage.
- External data / teacher / generated data: unlabeled data, EMA teacher từ chính student.
- Evaluation after each task: không phải continual task sequence.
Phase 9 — Claim → Evidence
| Claim | Where claim appears | Experiment | Evidence | My judgment | Caveat |
|---|---|---|---|---|---|
| Weight-averaged teacher cải thiện target so với /Temporal Ensembling. | Abstract/Section 2 | SVHN/CIFAR | Table 1-2; Figure 3 | Strong trong SSL image benchmarks. | Không chứng minh trực tiếp cho continual KD. |
| EMA update theo step scale tốt hơn Temporal Ensembling. | Section 2 | dataset lớn/online argument + ImageNet | Figure 2, Table 4 | Plausible + supported by large-scale run. | Temporal Ensembling comparison chủ yếu ở small image benchmarks. |
| Noise/augmentation vẫn cần thiết. | Section 3.4 | ablation | Figure 4(a/b) | Strong trong SVHN 250 labels. | Có thể khác với NLP/CRE augmentation. |
| MSE consistency tốt hơn KL trong setting này. | Appendix C | cost function sweep | Figure 5 | Reported/observed. | Authors nói nguyên nhân chưa rõ. |
Phase 10 — Ablation Study
| Component | Intended purpose | With component | Without / changed component | Difference | Conclusion justified | Not justified |
|---|---|---|---|---|---|---|
| Noise/augmentation | tạo perturbation cho consistency | passable validation error | no noise/input noise only kém hơn | qualitative từ Fig. 4 | cần perturbation phù hợp | không kết luận mọi noise đều tốt |
| EMA decay | teacher smoothing | good around selected ranges | quá thấp/quá cao degrade | qualitative từ Fig. 4(d) | EMA là hyperparameter quan trọng | không có công thức chọn tối ưu chung |
| Consistency weight | cân bằng supervised/consistency | range tốt khoảng order-of-magnitude | ngoài range degrade | qualitative từ Fig. 4(c) | cần ramp/weight tuning | không tự động ổn định mọi setting |
| Cost shape | đo mismatch prediction | MSE tốt nhất trong sweep | KL/C tau tệ hơn | qualitative từ Fig. 5 | MSE hợp setting này | không phủ định KL trong bài toán khác |
Phase 11 — Critical Reading
Reviewer notes
- Strongest contribution: đổi đơn vị ensembling từ prediction per-example sang weight EMA, tạo teacher target cập nhật nhanh hơn.
- Weakest part: paper không phải continual learning; khi dùng cho CRE phải tránh overclaim rằng EMA teacher tự giải quyết forgetting.
- Main assumption: consistency target chất lượng hơn sẽ cải thiện semi-supervised learning.
- Alternative explanation: architecture/augmentation/training schedule đóng góp lớn cùng với EMA.
- Missing experiment: kết hợp Mean Teacher với VAT như authors gợi ý nhưng không làm.
- Generalization risk: từ image SSL sang NLP/CRE cần thiết kế augmentation và teacher target khác.
- Reproducibility risk: nhiều training details/hyperparameters nằm ở appendix; cần tái lập schedule chính xác.
Phase 12 — Reproduction Check
| Item | Status | Detail | Missing detail / risk |
|---|---|---|---|
| Dataset and split | Clearly specified | SVHN/CIFAR/ImageNet label fractions | label subset randomization cần seed cụ thể |
| Preprocessing | Clearly specified | Appendix B | cần follow đúng từng dataset |
| Model / backbone | Clearly specified | ConvNet, ResNet/ResNeXt | implementation details nhiều |
| Training procedure | Partially specified | ramp-up, EMA update, optimizer | exact code hữu ích hơn paper |
| Loss functions | Clearly specified | CE + consistency + weight decay | cost variants cần cẩn thận |
| Hyperparameters | Partially specified | tables/appendix | tuning range rộng |
| Evaluation protocol | Clearly specified | error rate, runs | ImageNet validation-only |
Phase 13 — Completeness / Oral Exam
Completion criteria
- Giải thích vì sao EMA weights khác Temporal Ensembling.
- Viết được công thức EMA teacher và consistency cost.
- Nói được khi nào Mean Teacher cần noise/augmentation.
- Đọc được Figure 4 ablation.
- Phân biệt Mean Teacher với Knowledge Distillation hậu training.
- Nói rõ liên hệ và giới hạn khi mượn Mean Teacher cho continual KD.
Prompt oral exam
Act as my PhD advisor. Quiz me one question at a time about Mean Teacher: problem, method, equations, experiments, ablations, limitations, and how to transfer the idea to continual relation extraction. Wait for my answer before giving feedback.Final Paper Note Handoff
- Tóm tắt một câu:
- Problem:
- Gap:
- Method overview:
- Important equations:
- Protocol fingerprint:
- Main results:
- Ablation:
- Limitations:
- Critical judgment:
- Concepts cần tạo/cập nhật: