2026-09-23 - Mean Teacher - Gemini Notebook Workflow

Ranh giới

Đây là working note scaffold từ PDF/paper note để hỗ trợ đọc. Các phần factual dùng citation từ paper; phần recall cá nhân để trống và không đánh dấu đã đọc xong.

Setup

Từ điển khái niệm nhanh

Khái niệmĐịnh nghĩa ngắn trong paper nàyVì sao quan trọngLink
Mean TeacherTeacher model có trọng số là exponential moving average của student weights.Là đóng góp chính: target ổn định hơn teacher cùng trọng số với student.Mean Teachers are Better Role Models
EMA weightsCập nhật teacher bằng .Tạo teacher mượt hơn, cập nhật theo step thay vì theo epoch.Model Distillation
Consistency regularizationÉp prediction của student và teacher nhất quán khi có noise/augmentation.Cho phép dùng unlabeled data như regularizer.Knowledge Distillation
Temporal EnsemblingEMA trên prediction của từng training example.Baseline trực tiếp; Mean Teacher thay prediction-average bằng weight-average.Knowledge Distillation
Confirmation biasTeacher tự sinh target sai và student học theo nếu target kém.Motivation cho việc cải thiện quality của teacher target.Model Distillation
MSE consistency costPaper chủ yếu dùng MSE giữa prediction student và teacher.Là loss chính cho consistency; Appendix C so với KL-divergence.Knowledge Distillation

Phase 1 — Paper Map

Paper map — scaffold từ nguồn

  • Problem: Semi-supervised learning cần tận dụng unlabeled data để giảm overfitting khi label đắt; consistency target trên unlabeled data dễ bị confirmation bias nếu teacher yếu. PDF tr. 1, PDF tr. 2
  • Motivation: Temporal Ensembling cải thiện target bằng EMA prediction nhưng mỗi example chỉ cập nhật target một lần mỗi epoch, nên chậm và khó mở rộng với dataset lớn. PDF tr. 2
  • Main idea: Dùng teacher có weights là EMA của student weights; student học supervised classification loss trên labeled data và consistency loss với teacher output trên labeled/unlabeled data. PDF tr. 3
  • Important figure: Figure 2 mô tả student/teacher, noise riêng , classification cost, consistency cost và EMA update. PDF tr. 3
  • Important equations: Consistency cost ; EMA update . PDF tr. 3
  • Main result tables: Table 1-2 cho SVHN/CIFAR-10 ConvNet; Table 4 cho ResNet Mean Teacher trên CIFAR-10 và ImageNet. PDF tr. 4, PDF tr. 7
  • Ablation: Figure 4 kiểm tra noise, teacher dropout, consistency weight, EMA decay, dual output, MSE vs KL. PDF tr. 6
  • Limitation/caveat: Paper thuộc semi-supervised image classification, không phải continual learning; khi mượn cho CRE/TAPTA cần nói rõ đây là nguồn cơ chế EMA teacher, không phải baseline CRE.

Chỗ cần đọc trước

  • Abstract + Introduction
  • Figure 1-2
  • Section 3 experiments
  • Figure 4 ablation
  • Appendix B-C cho training details và MSE/KL

Phase 2 — Pass 1 Recall

Closed-book recall của tôi

Problem

Why does it matter?

Research gap

Main idea

Main contribution

Main result

Prompt kiểm tra recall

I have completed the first pass of the paper.
 
Here is my understanding:
 
[PASTE MY NOTES]
 
Compare my understanding against the paper.
Return what I understood correctly, what is inaccurate, important points I missed, confusing concepts, and source sections to revisit.
Do not rewrite the entire paper for me.

Phase 3 — Problem / Motivation / Gap

MụcDiễn giải bằng lời của tôiEvidence / citation
General problemDeep models overfit khi label ít; unlabeled data cần được khai thác bằng regularization.PDF tr. 1
Why it mattersGán nhãn thủ công đắt, còn model lớn cần nhiều data để học abstraction tốt.PDF tr. 1
What prior work solves model và Temporal Ensembling dùng consistency target để tận dụng unlabeled data.PDF tr. 2
What prior work fails to solveTemporal Ensembling cập nhật target theo example chậm, mỗi epoch một lần, nên feedback loop kém nhanh khi dataset lớn.PDF tr. 2
Exact research gapCần teacher target ổn định hơn student tức thời nhưng cập nhật nhanh hơn EMA prediction theo epoch.PDF tr. 3
Hypothesis / intuitionAveraging weights tạo teacher có representation trung gian và prediction target tốt hơn, từ đó consistency learning nhanh và ổn định hơn.PDF tr. 3
Contribution addressing the gapMean Teacher: teacher weights là EMA của student weights, dùng consistency loss với teacher output.PDF tr. 3

Phase 4 — Method / Architecture

Tôi tự vẽ trước

Input image
-> augmentation/noise cho student và teacher
-> Student model theta
-> classification loss nếu có label
-> Teacher model theta' = EMA(theta)
-> consistency loss giữa prediction student và teacher
-> update theta bằng SGD/Adam
-> update theta' bằng EMA, không backprop qua teacher

Component map

ComponentInputOperationOutputPurposeEvidence
Student modelnoisy/augmented inputforward + gradient updateprediction học supervised + consistency objectivePDF tr. 3
Teacher modelnoisy/augmented inputforward only với EMA weightsprediction sinh target ổn định hơnPDF tr. 3
EMA updatestudent weights teacher weights mớiaverage qua training stepsPDF tr. 3
Consistency coststudent/teacher predictionsthường dùng MSEregularization signalhọc invariance từ unlabeled dataPDF tr. 3
Noise/augmentationinput/layer noiserandom translation, flip, Gaussian noise, dropouthai views khác nhautạo perturbation để consistency có ý nghĩaPDF tr. 4

Điều tôi vẫn chưa hiểu

  • EMA teacher khác checkpoint averaging cuối training ở điểm nào về gradient/training dynamics?
  • Với CRE/TAPTA, EMA nên áp dụng lên backbone, prompt, router, hay teacher logits/prototypes?

Phase 6 — Equations

Prompt operational equation walkthrough

Walk me through the key equations or formal blocks in this paper.
 
For each equation/block, explain input, output, where it is used, what behavior it encourages, what weakens if removed, and which table/figure/ablation supports it.
Focus only on operational understanding of the equations and formal mechanisms.

Equation queue

Eq.Dùng để làm gì?Biến chínhBehavior được khuyến khíchEvidence / ablationStatus
Consistency cost giữa prediction của student và teacher dưới noise.Prediction ổn định quanh data manifold.Figure 2, Section 2/3; PDF tr. 3scaffold
Cập nhật teacher weights bằng EMA.Teacher chậm hơn, mượt hơn, target ít nhiễu hơn.Figure 4(d) sensitivity; PDF tr. 6scaffold
MSE consistencyLoss thực nghiệm chính cho consistency.teacher/student probabilitiesPenalize prediction mismatch trên labeled/unlabeled data.Appendix C so MSE tốt hơn KL trong setting này; PDF tr. 15scaffold

Phase 7 — Loss Functions

Total Loss
├── Classification cost -> học nhãn thật trên labeled examples
├── Consistency cost -> khớp prediction student với EMA teacher
├── Dual-output/logit MSE trick -> tách phần classification và consistency trong một số setting
└── Weight decay -> regularization
LossEquationInputsTrains componentBehaviorWeightAblation
Classification costtask CElabeled input + one-hot labelstudenthọc nhãn thậttheo setupbaseline supervised
Consistency costMSE hoặc KL familystudent/teacher predictionstudentổn định prediction dưới perturbationramp-up; SVHN/CIFAR/ImageNet khác nhauFigure 4(c/f), Appendix C
EMA teacherkhông backprop loss trực tiếpstudent weightsteacher update onlytarget mượtFigure 4(d)

Phase 8 — Experiments

Experiment map

ExperimentResearch questionDatasetBaselinesMetricTable/FigureMain resultCaveat
ConvNet SSLMean Teacher có hơn model/Temporal Ensembling không?SVHN, CIFAR-10Supervised-only, , Temporal Ensembling, VATerror rateTable 1-2SVHN 250 labels: 4.35%; CIFAR-10 4000 labels ConvNet: 12.31%.VAT vẫn mạnh ở vài setting.
Extra unlabeled dataCó tận dụng thêm unlabeled data không?SVHN extra modelerror rateTable 3/Figure 3Mean Teacher tốt hơn khi thêm unlabeled data. vẫn cải thiện lâu hơn khi extra data rất lớn.
AblationNoise, EMA decay, consistency weight nhạy ra sao?SVHN 250 labelsvariantsvalidation errorFigure 4Noise/augmentation cần thiết; EMA decay và consistency weight có range tốt.One-factor-at-a-time ablation.
ResNet/ImageNetCó scale sang architecture/dataset lớn không?CIFAR-10, ImageNet 2012SOTAerror rateTable 4CIFAR-10 4000 labels 6.28%; ImageNet 10% labels 9.11%.ImageNet dùng validation vì test set không public.

Protocol fingerprint

  • Dataset and split: SVHN, CIFAR-10, ImageNet 2012; nhiều mức label fraction.
  • Scenario / label space: semi-supervised classification, unlabeled examples dùng consistency.
  • Backbone: 13-layer ConvNet; ResNet/ResNeXt ở experiment lớn.
  • Frozen/trainable components: student trainable; teacher updated by EMA, không backprop.
  • Seeds / number of runs: 10 runs cho nhiều SVHN/CIFAR setting; ImageNet 2 runs.
  • Metric and averaging: classification error percentage.
  • External data / teacher / generated data: unlabeled data, EMA teacher từ chính student.
  • Evaluation after each task: không phải continual task sequence.

Phase 9 — Claim → Evidence

ClaimWhere claim appearsExperimentEvidenceMy judgmentCaveat
Weight-averaged teacher cải thiện target so với /Temporal Ensembling.Abstract/Section 2SVHN/CIFARTable 1-2; Figure 3Strong trong SSL image benchmarks.Không chứng minh trực tiếp cho continual KD.
EMA update theo step scale tốt hơn Temporal Ensembling.Section 2dataset lớn/online argument + ImageNetFigure 2, Table 4Plausible + supported by large-scale run.Temporal Ensembling comparison chủ yếu ở small image benchmarks.
Noise/augmentation vẫn cần thiết.Section 3.4ablationFigure 4(a/b)Strong trong SVHN 250 labels.Có thể khác với NLP/CRE augmentation.
MSE consistency tốt hơn KL trong setting này.Appendix Ccost function sweepFigure 5Reported/observed.Authors nói nguyên nhân chưa rõ.

Phase 10 — Ablation Study

ComponentIntended purposeWith componentWithout / changed componentDifferenceConclusion justifiedNot justified
Noise/augmentationtạo perturbation cho consistencypassable validation errorno noise/input noise only kém hơnqualitative từ Fig. 4cần perturbation phù hợpkhông kết luận mọi noise đều tốt
EMA decay teacher smoothinggood around selected rangesquá thấp/quá cao degradequalitative từ Fig. 4(d)EMA là hyperparameter quan trọngkhông có công thức chọn tối ưu chung
Consistency weightcân bằng supervised/consistencyrange tốt khoảng order-of-magnitudengoài range degradequalitative từ Fig. 4(c)cần ramp/weight tuningkhông tự động ổn định mọi setting
Cost shapeđo mismatch predictionMSE tốt nhất trong sweepKL/C tau tệ hơnqualitative từ Fig. 5MSE hợp setting nàykhông phủ định KL trong bài toán khác

Phase 11 — Critical Reading

Reviewer notes

  • Strongest contribution: đổi đơn vị ensembling từ prediction per-example sang weight EMA, tạo teacher target cập nhật nhanh hơn.
  • Weakest part: paper không phải continual learning; khi dùng cho CRE phải tránh overclaim rằng EMA teacher tự giải quyết forgetting.
  • Main assumption: consistency target chất lượng hơn sẽ cải thiện semi-supervised learning.
  • Alternative explanation: architecture/augmentation/training schedule đóng góp lớn cùng với EMA.
  • Missing experiment: kết hợp Mean Teacher với VAT như authors gợi ý nhưng không làm.
  • Generalization risk: từ image SSL sang NLP/CRE cần thiết kế augmentation và teacher target khác.
  • Reproducibility risk: nhiều training details/hyperparameters nằm ở appendix; cần tái lập schedule chính xác.

Phase 12 — Reproduction Check

ItemStatusDetailMissing detail / risk
Dataset and splitClearly specifiedSVHN/CIFAR/ImageNet label fractionslabel subset randomization cần seed cụ thể
PreprocessingClearly specifiedAppendix Bcần follow đúng từng dataset
Model / backboneClearly specifiedConvNet, ResNet/ResNeXtimplementation details nhiều
Training procedurePartially specifiedramp-up, EMA update, optimizerexact code hữu ích hơn paper
Loss functionsClearly specifiedCE + consistency + weight decaycost variants cần cẩn thận
HyperparametersPartially specifiedtables/appendixtuning range rộng
Evaluation protocolClearly specifiederror rate, runsImageNet validation-only

Phase 13 — Completeness / Oral Exam

Completion criteria

  • Giải thích vì sao EMA weights khác Temporal Ensembling.
  • Viết được công thức EMA teacher và consistency cost.
  • Nói được khi nào Mean Teacher cần noise/augmentation.
  • Đọc được Figure 4 ablation.
  • Phân biệt Mean Teacher với Knowledge Distillation hậu training.
  • Nói rõ liên hệ và giới hạn khi mượn Mean Teacher cho continual KD.

Prompt oral exam

Act as my PhD advisor. Quiz me one question at a time about Mean Teacher: problem, method, equations, experiments, ablations, limitations, and how to transfer the idea to continual relation extraction. Wait for my answer before giving feedback.

Final Paper Note Handoff

  • Tóm tắt một câu:
  • Problem:
  • Gap:
  • Method overview:
  • Important equations:
  • Protocol fingerprint:
  • Main results:
  • Ablation:
  • Limitations:
  • Critical judgment:
  • Concepts cần tạo/cập nhật:

Liên kết