2026-09-23 - LwF - Gemini Notebook Workflow
Ranh giới
Đây là working note scaffold từ canonical paper note/PDF để hỗ trợ đọc. Không đánh dấu đã đọc xong; phần recall cá nhân để trống.
Setup
- Paper: Learning without Forgetting
- PDF: Learning without Forgetting.pdf
- Paper note chính: Learning without Forgetting
- Gemini Notebook / NotebookLM URL:
- Mục tiêu buổi đọc: nắm LwF như baseline KD cho continual learning, đặc biệt recorded old responses trên dữ liệu task mới, warm-up, loss balance và failure mode khi task mới lệch phân phối.
- Phần cần đọc trước: Figure 1-3, Section 3, Table 1, Figure 4, Table 2(b), Figure 7, Discussion.
- PDF count đã kiểm tra: 13 trang.
Từ điển khái niệm nhanh
| Khái niệm | Định nghĩa ngắn trong paper này | Vì sao quan trọng | Link |
|---|---|---|---|
| LwF | Học task mới bằng dữ liệu mới, đồng thời dùng output cũ trên dữ liệu mới làm distillation targets cho old tasks. | Baseline kinh điển cho rehearsal-free distillation-based continual learning. | Learning without Forgetting |
| Recorded responses | Output của original model trên dữ liệu task mới trước khi update. | Thay thế old task data bằng soft targets. | Knowledge Distillation |
| Shared parameters | Các layer chung được update khi học task mới. | Nơi xảy ra trade-off old/new tasks. | Continual Learning |
| Old task parameters | Head/parameters cho task cũ. | LwF vẫn joint-optimize old head để match old responses. | Catastrophic Forgetting |
| New task parameters | Head mới random initialized cho task mới. | Cần warm-up để tránh gradient lớn phá shared representation. | Continual Learning |
| Distillation loss | Modified cross-entropy với temperature giữa old responses và current old predictions. | Cơ chế giữ old behavior mà không có old data. | Knowledge Distillation |
| Distribution mismatch | Task mới không kích hoạt feature/domain của old task. | Failure mode chính, quan trọng khi chuyển sang CRE/TAPTA. | Catastrophic Forgetting |
Phase 1 — Paper Map
Paper map — scaffold từ nguồn
- Problem: Thêm task/capability mới vào CNN đã train mà không có dữ liệu huấn luyện của task cũ; cần tránh catastrophic forgetting. PDF tr. 1, PDF tr. 2
- Motivation: Old data có thể quá lớn, không được lưu vì privacy/proprietary, hoặc không thực tế để retrain joint model. PDF tr. 1
- Main idea: Chạy original model trên dữ liệu task mới để ghi lại old-task responses ; train model mới bằng loss cho task mới + distillation loss giữ output cũ trên . PDF tr. 5
- Important figure: Figure 2 so sánh fine-tuning, feature extraction, joint training và LwF; Figure 3 procedure. PDF tr. 3, PDF tr. 5
- Important equations: recorded responses, old/new predictions, total objective với ; distillation temperature. PDF tr. 5
- Main result table: Table 1 single-new-task performance trên ImageNet/Places → VOC/CUB/Scenes/MNIST. PDF tr. 7
- Ablation/design choices: Table 2 warm-up/design alternatives; Figure 7 loss balance and response-preserving losses. PDF tr. 9, PDF tr. 10
- Limitations: Old task preservation yếu hơn khi new data quá lệch old distribution, ví dụ ImageNet → MNIST. PDF tr. 7
Chỗ cần đọc trước
- Figure 1-3
- Section 3 procedure/objective
- Table 1
- Figure 4 sequential tasks
- Table 2(b) warm-up
- Figure 7 loss balance
Phase 2 — Pass 1 Recall
Closed-book recall của tôi
Problem
Why does it matter?
Research gap
Main idea
Main contribution
Main result
Prompt kiểm tra recall
I have completed the first pass of Learning without Forgetting.
Here is my understanding:
[PASTE MY NOTES]
Compare my understanding against the paper. Identify correct, inaccurate, missing, confusing concepts, and sections/figures/equations to revisit.Phase 3 — Problem / Motivation / Gap
| Mục | Diễn giải bằng lời của tôi | Evidence / citation |
|---|---|---|
| General problem | Học thêm task mới vào network hiện có mà vẫn giữ performance task cũ. | PDF tr. 1 |
| Why it matters | Old data có thể không lưu được, quá lớn, proprietary, hoặc retraining joint model quá đắt. | PDF tr. 1 |
| What prior work solves | Fine-tuning học task mới; feature extraction giữ task cũ; joint training tốt nhưng cần old data. | PDF tr. 2 |
| What prior work fails to solve | Fine-tuning quên task cũ; feature extraction hạn chế task mới; joint training vi phạm giả định không có old data. | PDF tr. 2 |
| Exact research gap | Cần joint-optimize shared parameters cho task mới nhưng không cần old training data. | PDF tr. 5 |
| Hypothesis / intuition | Old model responses trên new data có thể đóng vai trò surrogate targets để giữ old behavior. | PDF tr. 5 |
| Contribution addressing the gap | LwF dùng recorded responses + distillation loss + new task loss để học không quên. | PDF tr. 5 |
Phase 4 — Method / Architecture
Tôi tự vẽ trước
Existing model: shared theta_s + old heads theta_o
New task data X_n, Y_n
-> forward old model on X_n
-> record old-task responses Y_o
-> initialize new head theta_n
-> warm-up theta_n while freezing theta_s/theta_o
-> joint optimize theta_s, theta_o, theta_n
L_new(Y_n, Yhat_n) + lambda_o L_old(Y_o, Yhat_o) + R
-> evaluate both old and new tasksComponent map
| Component | Input | Operation | Output | Purpose | Evidence |
|---|---|---|---|---|---|
| Original network | forward pass before update | recorded responses | create old-task soft targets without old data | PDF tr. 5 | |
| New task head | shared representation | random init + warm-up | learn task-specific classifier | PDF tr. 5 | |
| Old-task prediction branch | current model forward | compare to recorded old responses | PDF tr. 5 | ||
| New-task prediction branch | current model forward | supervised new task learning | PDF tr. 5 | ||
| Loss balance | old/new losses | trade-off | chosen optimum | control old-new performance trade-off | PDF tr. 10 |
Điều tôi vẫn chưa hiểu
- Khi không kích hoạt old features, có thể dùng unlabeled old-domain anchors thay thế không?
- Trong CRE, recorded responses nên là logits, prototype similarities, hay router distribution?
Phase 6 — Equations
Prompt operational equation walkthrough
Walk me through the key equations or formal blocks in this paper.
For each equation/block, explain input, output, where it is used, behavior encouraged, failure mode if removed, and supporting table/figure/result.
Focus only on operational understanding.Equation queue
| Eq. | Dùng để làm gì? | Biến chính | Behavior được khuyến khích | Evidence / ablation | Status |
|---|---|---|---|---|---|
| Record old-task responses on new-task data. | tạo surrogate old labels không cần old data. | Figure 3; PDF tr. 5 | scaffold | ||
| definitions | Current model outputs cho old/new tasks. | joint model phải vừa học mới vừa giữ cũ. | Figure 3; PDF tr. 5 | scaffold | |
| Total objective | Balance new CE, old distillation, regularization. | preserve old behavior while learning new. | Table 1, Figure 7; PDF tr. 7 | scaffold | |
| Modified CE with temperature | Distillation target smoothing. | giữ relative class similarities, not just top class. | Figure 7 loss comparison; PDF tr. 10 | scaffold |
Phase 7 — Loss Functions
Total Loss
├── L_new -> supervised loss for the new task
├── lambda_o L_old -> distillation loss preserving old-task responses
└── R -> regularization / weight decay| Loss | Equation | Inputs | Trains component | Behavior | Weight | Ablation |
|---|---|---|---|---|---|---|
| cross-entropy | shared + new head | learn new task | 1 | compared through baselines | ||
| modified CE/KD | shared + old head | preserve old output behavior | default 1 | Figure 7 | ||
| weight decay | parameters | all trainable params | regularization | 0.0005 reported in note | not central |
Phase 8 — Experiments
Experiment map
| Experiment | Research question | Dataset | Baselines | Metric | Table/Figure | Main result | Caveat |
|---|---|---|---|---|---|---|---|
| Single new task | LwF giữ old task và học new task không cần old data tốt không? | ImageNet/Places → VOC/CUB/Scenes/MNIST | Fine-tuning, LFL, fine-tune FC, feature extraction, joint training | mAP/accuracy | Table 1 | LwF tốt trên new task và giữ old task hơn fine-tuning; gần joint training ở nhiều case. | ImageNet → MNIST là failure case. |
| Sequential tasks | Khi thêm nhiều task liên tiếp, degradation thế nào? | Places → VOC parts; ImageNet → Scenes parts | same baselines | task performance over sequence | Figure 4 | LwF degrade chậm hơn fine-tuning. | Số task còn nhỏ. |
| Data size | Ít/more new data có thay đổi conclusion không? | ImageNet → CUB subsampling | baselines | accuracy | Figure 5 | Observation chính vẫn giữ. | scatter/runs limited. |
| Design alternatives | Layer choice, network expansion, warm-up | selected task pairs | variants | old/new performance | Table 2 | Warm-up thiết yếu cho fine-tuning; LwF ít nhạy hơn nhưng vẫn dùng. | Không phải ablation thuần loss. |
| Loss balance | và response loss shape ảnh hưởng ra sao? | Places→VOC, ImageNet→Scenes | LwF variants | old/new trade-off | Figure 7 | KD loss hơi tốt hơn L1/L2/CE; tạo trade-off. | Advantage không lớn. |
Protocol fingerprint
- Dataset and split: old tasks ImageNet/Places365; new tasks VOC, CUB, Scenes, MNIST; sequential variants chia VOC/Scenes thành parts.
- Scenario / label space: old task training data unavailable; only new task data used.
- Backbone: AlexNet main; VGG-16 additional.
- Frozen/trainable components: warm-up trains new head; joint optimize shared, old, new parameters.
- Metric and averaging: VOC mAP; others top-1 accuracy.
- Baseline implementation: fine-tuning, feature extraction, fine-tune FC, LFL, joint training.
- External data / memory: no old data for LwF; joint training uses old data as upper-bound.
- Evaluation after each task: old and new task performance; sequential Figure 4.
Phase 9 — Claim → Evidence
| Claim | Where claim appears | Experiment | Evidence | My judgment | Caveat |
|---|---|---|---|---|---|
| LwF learns new task while preserving old task without old data. | Abstract/Intro | Table 1 | LwF beats fine-tuning on old task, often near joint training. | Strong for reported vision tasks. | weaker when distributions are very dissimilar. |
| LwF can improve new task vs fine-tuning. | Abstract/Discussion | Table 1 | In many task pairs LwF new-task result comparable/better. | Supported, surprising. | depends on task similarity and regularization effect. |
| Warm-up protects old task in fine-tuning/LwF training. | Section 4.2 | Table 2(b) | Fine-tuning no warm-up old-task drop is severe. | Strong practical lesson. | LwF less sensitive than fine-tuning. |
| KD response-preserving loss is slightly preferable. | Section 4.2 | Figure 7 | KD slightly outperforms L1/L2/CE variants. | Moderate. | authors say advantage not large. |
Phase 10 — Ablation Study
| Component | Intended purpose | With component | Without component | Difference | Conclusion justified | Not justified |
|---|---|---|---|---|---|---|
| Warm-up in fine-tuning | avoid random new head disturbing shared layers | ImageNet→CUB old 50.9 | no warm-up old 42.5 | -8.4 | warm-up crucial for fine-tuning old-task retention | not proof warm-up solves forgetting alone |
| Warm-up in LwF | stabilize new head before joint optimize | ImageNet→CUB old/new 54.7/57.7 | no warm-up 53.5/59.9 | old -1.2, new +2.2 | LwF robust but trade-off changes | not universally better with warm-up |
| Network expansion + LwF | add capacity | 54.4/57.0 | LwF 54.7/57.7 | slightly worse | extra capacity not necessary here | not a broad capacity conclusion |
| Loss type | response preservation | KD line slightly better | L1/L2/CE variants | small | KD is reasonable default | not critical determinant |
Phase 11 — Critical Reading
Reviewer notes
- Strongest contribution: simple formulation of distillation-based continual learning without old data.
- Weakest part: preserving old behavior using only depends on new data covering useful old decision regions.
- Main assumption: old model responses on new-task inputs contain enough information to constrain old-task behavior.
- Alternative explanation: performance gains on new task may be regularization more than knowledge preservation.
- Missing experiment: direct analysis of how representative is for old-task feature activation.
- Generalization risk: vision classification results do not directly transfer to relation extraction sequence protocols.
- Reproducibility risk: old datasets, task splits, warm-up schedules, and hyperparameter selection matter.
Phase 12 — Reproduction Check
| Item | Status | Detail | Missing detail / risk |
|---|---|---|---|
| Dataset and split | Clearly specified | ImageNet, Places365, VOC, CUB, Scenes, MNIST | exact preprocessing/task part split needs care |
| Preprocessing | Partially specified | described around experiments | code useful for exact replication |
| Model / backbone | Clearly specified | AlexNet/VGG | framework versions old |
| Training procedure | Clearly specified | warm-up then joint optimize | schedule details scattered |
| Loss functions | Clearly specified | new CE + old KD + regularization | temperature/lambda tuning |
| Optimizer | Partially specified | SGD details in canonical note/sections | exact schedules needed |
| Evaluation protocol | Clearly specified | old/new task metrics | compare validation/test carefully |
| Inference procedure | Clearly specified | unified network with task heads | task identity assumption matters |
Phase 13 — Completeness / Oral Exam
Completion criteria
- Giải thích LwF không cần old data bằng recorded responses.
- Viết được objective LwF và vai trò .
- Phân biệt LwF với fine-tuning, feature extraction, joint training.
- Nói được failure mode distribution mismatch.
- Đọc được Table 1 và Figure 7.
- Chuyển ý tưởng sang CRE mà không overclaim.
Prompt oral exam
Act as my PhD advisor. Quiz me one question at a time about Learning without Forgetting: problem, recorded responses, objective, warm-up, experiments, failure modes, and how it relates to SCKD/Mean Teacher/TAPTA. Wait for my answer before giving feedback.Final Paper Note Handoff
- Tóm tắt một câu:
- Problem:
- Gap:
- Method overview:
- Important equations:
- Protocol fingerprint:
- Main results:
- Ablation:
- Limitations:
- Critical judgment:
- Concepts cần tạo/cập nhật:
Liên kết
- Paper note: Learning without Forgetting
- PDF: Learning without Forgetting.pdf
- Concepts: Continual Learning, Knowledge Distillation, Catastrophic Forgetting, Continual Relation Extraction
- Related papers: Mean Teachers are Better Role Models, Serial Contrastive Knowledge Distillation for Continual Few-shot Relation Extraction