2026-09-23 - LwF - Gemini Notebook Workflow

Ranh giới

Đây là working note scaffold từ canonical paper note/PDF để hỗ trợ đọc. Không đánh dấu đã đọc xong; phần recall cá nhân để trống.

Setup

  • Paper: Learning without Forgetting
  • PDF: Learning without Forgetting.pdf
  • Paper note chính: Learning without Forgetting
  • Gemini Notebook / NotebookLM URL:
  • Mục tiêu buổi đọc: nắm LwF như baseline KD cho continual learning, đặc biệt recorded old responses trên dữ liệu task mới, warm-up, loss balance và failure mode khi task mới lệch phân phối.
  • Phần cần đọc trước: Figure 1-3, Section 3, Table 1, Figure 4, Table 2(b), Figure 7, Discussion.
  • PDF count đã kiểm tra: 13 trang.

Từ điển khái niệm nhanh

Khái niệmĐịnh nghĩa ngắn trong paper nàyVì sao quan trọngLink
LwFHọc task mới bằng dữ liệu mới, đồng thời dùng output cũ trên dữ liệu mới làm distillation targets cho old tasks.Baseline kinh điển cho rehearsal-free distillation-based continual learning.Learning without Forgetting
Recorded responses Output của original model trên dữ liệu task mới trước khi update.Thay thế old task data bằng soft targets.Knowledge Distillation
Shared parameters Các layer chung được update khi học task mới.Nơi xảy ra trade-off old/new tasks.Continual Learning
Old task parameters Head/parameters cho task cũ.LwF vẫn joint-optimize old head để match old responses.Catastrophic Forgetting
New task parameters Head mới random initialized cho task mới.Cần warm-up để tránh gradient lớn phá shared representation.Continual Learning
Distillation loss Modified cross-entropy với temperature giữa old responses và current old predictions.Cơ chế giữ old behavior mà không có old data.Knowledge Distillation
Distribution mismatchTask mới không kích hoạt feature/domain của old task.Failure mode chính, quan trọng khi chuyển sang CRE/TAPTA.Catastrophic Forgetting

Phase 1 — Paper Map

Paper map — scaffold từ nguồn

  • Problem: Thêm task/capability mới vào CNN đã train mà không có dữ liệu huấn luyện của task cũ; cần tránh catastrophic forgetting. PDF tr. 1, PDF tr. 2
  • Motivation: Old data có thể quá lớn, không được lưu vì privacy/proprietary, hoặc không thực tế để retrain joint model. PDF tr. 1
  • Main idea: Chạy original model trên dữ liệu task mới để ghi lại old-task responses ; train model mới bằng loss cho task mới + distillation loss giữ output cũ trên . PDF tr. 5
  • Important figure: Figure 2 so sánh fine-tuning, feature extraction, joint training và LwF; Figure 3 procedure. PDF tr. 3, PDF tr. 5
  • Important equations: recorded responses, old/new predictions, total objective với ; distillation temperature. PDF tr. 5
  • Main result table: Table 1 single-new-task performance trên ImageNet/Places → VOC/CUB/Scenes/MNIST. PDF tr. 7
  • Ablation/design choices: Table 2 warm-up/design alternatives; Figure 7 loss balance and response-preserving losses. PDF tr. 9, PDF tr. 10
  • Limitations: Old task preservation yếu hơn khi new data quá lệch old distribution, ví dụ ImageNet → MNIST. PDF tr. 7

Chỗ cần đọc trước

  • Figure 1-3
  • Section 3 procedure/objective
  • Table 1
  • Figure 4 sequential tasks
  • Table 2(b) warm-up
  • Figure 7 loss balance

Phase 2 — Pass 1 Recall

Closed-book recall của tôi

Problem

Why does it matter?

Research gap

Main idea

Main contribution

Main result

Prompt kiểm tra recall

I have completed the first pass of Learning without Forgetting.
 
Here is my understanding:
 
[PASTE MY NOTES]
 
Compare my understanding against the paper. Identify correct, inaccurate, missing, confusing concepts, and sections/figures/equations to revisit.

Phase 3 — Problem / Motivation / Gap

MụcDiễn giải bằng lời của tôiEvidence / citation
General problemHọc thêm task mới vào network hiện có mà vẫn giữ performance task cũ.PDF tr. 1
Why it mattersOld data có thể không lưu được, quá lớn, proprietary, hoặc retraining joint model quá đắt.PDF tr. 1
What prior work solvesFine-tuning học task mới; feature extraction giữ task cũ; joint training tốt nhưng cần old data.PDF tr. 2
What prior work fails to solveFine-tuning quên task cũ; feature extraction hạn chế task mới; joint training vi phạm giả định không có old data.PDF tr. 2
Exact research gapCần joint-optimize shared parameters cho task mới nhưng không cần old training data.PDF tr. 5
Hypothesis / intuitionOld model responses trên new data có thể đóng vai trò surrogate targets để giữ old behavior.PDF tr. 5
Contribution addressing the gapLwF dùng recorded responses + distillation loss + new task loss để học không quên.PDF tr. 5

Phase 4 — Method / Architecture

Tôi tự vẽ trước

Existing model: shared theta_s + old heads theta_o
New task data X_n, Y_n
-> forward old model on X_n
-> record old-task responses Y_o
-> initialize new head theta_n
-> warm-up theta_n while freezing theta_s/theta_o
-> joint optimize theta_s, theta_o, theta_n
   L_new(Y_n, Yhat_n) + lambda_o L_old(Y_o, Yhat_o) + R
-> evaluate both old and new tasks

Component map

ComponentInputOperationOutputPurposeEvidence
Original networkforward pass before updaterecorded responses create old-task soft targets without old dataPDF tr. 5
New task headshared representationrandom init + warm-uplearn task-specific classifierPDF tr. 5
Old-task prediction branchcurrent model forwardcompare to recorded old responsesPDF tr. 5
New-task prediction branchcurrent model forwardsupervised new task learningPDF tr. 5
Loss balanceold/new losses trade-offchosen optimumcontrol old-new performance trade-offPDF tr. 10

Điều tôi vẫn chưa hiểu

  • Khi không kích hoạt old features, có thể dùng unlabeled old-domain anchors thay thế không?
  • Trong CRE, recorded responses nên là logits, prototype similarities, hay router distribution?

Phase 6 — Equations

Prompt operational equation walkthrough

Walk me through the key equations or formal blocks in this paper.
 
For each equation/block, explain input, output, where it is used, behavior encouraged, failure mode if removed, and supporting table/figure/result.
Focus only on operational understanding.

Equation queue

Eq.Dùng để làm gì?Biến chínhBehavior được khuyến khíchEvidence / ablationStatus
Record old-task responses on new-task data.tạo surrogate old labels không cần old data.Figure 3; PDF tr. 5scaffold
definitionsCurrent model outputs cho old/new tasks.joint model phải vừa học mới vừa giữ cũ.Figure 3; PDF tr. 5scaffold
Total objectiveBalance new CE, old distillation, regularization.preserve old behavior while learning new.Table 1, Figure 7; PDF tr. 7scaffold
Modified CE with temperatureDistillation target smoothing.giữ relative class similarities, not just top class.Figure 7 loss comparison; PDF tr. 10scaffold

Phase 7 — Loss Functions

Total Loss
├── L_new -> supervised loss for the new task
├── lambda_o L_old -> distillation loss preserving old-task responses
└── R -> regularization / weight decay
LossEquationInputsTrains componentBehaviorWeightAblation
cross-entropyshared + new headlearn new task1compared through baselines
modified CE/KDshared + old headpreserve old output behavior default 1Figure 7
weight decayparametersall trainable paramsregularization0.0005 reported in notenot central

Phase 8 — Experiments

Experiment map

ExperimentResearch questionDatasetBaselinesMetricTable/FigureMain resultCaveat
Single new taskLwF giữ old task và học new task không cần old data tốt không?ImageNet/Places → VOC/CUB/Scenes/MNISTFine-tuning, LFL, fine-tune FC, feature extraction, joint trainingmAP/accuracyTable 1LwF tốt trên new task và giữ old task hơn fine-tuning; gần joint training ở nhiều case.ImageNet → MNIST là failure case.
Sequential tasksKhi thêm nhiều task liên tiếp, degradation thế nào?Places → VOC parts; ImageNet → Scenes partssame baselinestask performance over sequenceFigure 4LwF degrade chậm hơn fine-tuning.Số task còn nhỏ.
Data sizeÍt/more new data có thay đổi conclusion không?ImageNet → CUB subsamplingbaselinesaccuracyFigure 5Observation chính vẫn giữ.scatter/runs limited.
Design alternativesLayer choice, network expansion, warm-upselected task pairsvariantsold/new performanceTable 2Warm-up thiết yếu cho fine-tuning; LwF ít nhạy hơn nhưng vẫn dùng.Không phải ablation thuần loss.
Loss balance và response loss shape ảnh hưởng ra sao?Places→VOC, ImageNet→ScenesLwF variantsold/new trade-offFigure 7KD loss hơi tốt hơn L1/L2/CE; tạo trade-off.Advantage không lớn.

Protocol fingerprint

  • Dataset and split: old tasks ImageNet/Places365; new tasks VOC, CUB, Scenes, MNIST; sequential variants chia VOC/Scenes thành parts.
  • Scenario / label space: old task training data unavailable; only new task data used.
  • Backbone: AlexNet main; VGG-16 additional.
  • Frozen/trainable components: warm-up trains new head; joint optimize shared, old, new parameters.
  • Metric and averaging: VOC mAP; others top-1 accuracy.
  • Baseline implementation: fine-tuning, feature extraction, fine-tune FC, LFL, joint training.
  • External data / memory: no old data for LwF; joint training uses old data as upper-bound.
  • Evaluation after each task: old and new task performance; sequential Figure 4.

Phase 9 — Claim → Evidence

ClaimWhere claim appearsExperimentEvidenceMy judgmentCaveat
LwF learns new task while preserving old task without old data.Abstract/IntroTable 1LwF beats fine-tuning on old task, often near joint training.Strong for reported vision tasks.weaker when distributions are very dissimilar.
LwF can improve new task vs fine-tuning.Abstract/DiscussionTable 1In many task pairs LwF new-task result comparable/better.Supported, surprising.depends on task similarity and regularization effect.
Warm-up protects old task in fine-tuning/LwF training.Section 4.2Table 2(b)Fine-tuning no warm-up old-task drop is severe.Strong practical lesson.LwF less sensitive than fine-tuning.
KD response-preserving loss is slightly preferable.Section 4.2Figure 7KD slightly outperforms L1/L2/CE variants.Moderate.authors say advantage not large.

Phase 10 — Ablation Study

ComponentIntended purposeWith componentWithout componentDifferenceConclusion justifiedNot justified
Warm-up in fine-tuningavoid random new head disturbing shared layersImageNet→CUB old 50.9no warm-up old 42.5-8.4warm-up crucial for fine-tuning old-task retentionnot proof warm-up solves forgetting alone
Warm-up in LwFstabilize new head before joint optimizeImageNet→CUB old/new 54.7/57.7no warm-up 53.5/59.9old -1.2, new +2.2LwF robust but trade-off changesnot universally better with warm-up
Network expansion + LwFadd capacity54.4/57.0LwF 54.7/57.7slightly worseextra capacity not necessary herenot a broad capacity conclusion
Loss typeresponse preservationKD line slightly betterL1/L2/CE variantssmallKD is reasonable defaultnot critical determinant

Phase 11 — Critical Reading

Reviewer notes

  • Strongest contribution: simple formulation of distillation-based continual learning without old data.
  • Weakest part: preserving old behavior using only depends on new data covering useful old decision regions.
  • Main assumption: old model responses on new-task inputs contain enough information to constrain old-task behavior.
  • Alternative explanation: performance gains on new task may be regularization more than knowledge preservation.
  • Missing experiment: direct analysis of how representative is for old-task feature activation.
  • Generalization risk: vision classification results do not directly transfer to relation extraction sequence protocols.
  • Reproducibility risk: old datasets, task splits, warm-up schedules, and hyperparameter selection matter.

Phase 12 — Reproduction Check

ItemStatusDetailMissing detail / risk
Dataset and splitClearly specifiedImageNet, Places365, VOC, CUB, Scenes, MNISTexact preprocessing/task part split needs care
PreprocessingPartially specifieddescribed around experimentscode useful for exact replication
Model / backboneClearly specifiedAlexNet/VGGframework versions old
Training procedureClearly specifiedwarm-up then joint optimizeschedule details scattered
Loss functionsClearly specifiednew CE + old KD + regularizationtemperature/lambda tuning
OptimizerPartially specifiedSGD details in canonical note/sectionsexact schedules needed
Evaluation protocolClearly specifiedold/new task metricscompare validation/test carefully
Inference procedureClearly specifiedunified network with task headstask identity assumption matters

Phase 13 — Completeness / Oral Exam

Completion criteria

  • Giải thích LwF không cần old data bằng recorded responses.
  • Viết được objective LwF và vai trò .
  • Phân biệt LwF với fine-tuning, feature extraction, joint training.
  • Nói được failure mode distribution mismatch.
  • Đọc được Table 1 và Figure 7.
  • Chuyển ý tưởng sang CRE mà không overclaim.

Prompt oral exam

Act as my PhD advisor. Quiz me one question at a time about Learning without Forgetting: problem, recorded responses, objective, warm-up, experiments, failure modes, and how it relates to SCKD/Mean Teacher/TAPTA. Wait for my answer before giving feedback.

Final Paper Note Handoff

  • Tóm tắt một câu:
  • Problem:
  • Gap:
  • Method overview:
  • Important equations:
  • Protocol fingerprint:
  • Main results:
  • Ablation:
  • Limitations:
  • Critical judgment:
  • Concepts cần tạo/cập nhật:

Liên kết