2026-09-23 - SCKD - Gemini Notebook Workflow
Ranh giới
Đây là working note scaffold từ paper note/PDF. Các phần factual đã gắn citation; closed-book recall, oral exam answer và kết luận cá nhân để bạn tự điền sau khi đọc.
Setup
- Paper: Serial Contrastive Knowledge Distillation for Continual Few-shot Relation Extraction
- PDF: Serial Contrastive Knowledge Distillation for Continual Few-shot Relation Extraction.pdf
- Paper note chính: Serial Contrastive Knowledge Distillation for Continual Few-shot Relation Extraction
- Gemini Notebook / NotebookLM URL:
- Mục tiêu buổi đọc: hiểu serial contrastive knowledge distillation, khác LwF/CRL/CRECL ở đâu, và baseline này tác động gì đến thiết kế TAPTA.
- Phần cần đọc trước: Abstract, Section 3, Algorithm 1, Figure 2, Table 1-3, Section 6.
- PDF count đã kiểm tra: 14 trang.
Từ điển khái niệm nhanh
| Khái niệm | Định nghĩa ngắn trong paper này | Vì sao quan trọng | Link |
|---|---|---|---|
| Continual few-shot RE | Task đầu có data nhiều, các task sau chỉ N-way K-shot; sau task đánh giá trên tất cả relation đã thấy. | Là setting chính mà SCKD tối ưu. | Continual Few-Shot Relation Extraction |
| SCKD | Serial Contrastive Knowledge Distillation: feature/prediction/hidden contrastive distillation theo chuỗi model trước-sau. | Đóng góp chính để giảm forgetting và overfitting. | Serial Contrastive Knowledge Distillation for Continual Few-shot Relation Extraction |
| Memory | Bộ nhớ lưu vài typical samples từ các relation cũ và mới. | SCKD là memory-based method; memory size mặc định một sample/relation trong main experiments. | Replay in Continual Learning |
| Prototype | Mean hidden representation của typical samples cho relation . | Dùng để sinh pseudo samples cho contrastive distillation. | Prototype Learning |
| Pseudo samples | Vector giả sinh quanh prototype với Gaussian noise và diagonal covariance. | Tăng tín hiệu contrastive khi few-shot data ít. | Contrastive Learning |
| Feature distillation loss giữa feature của previous model và current model. | Giữ encoder không lệch quá mạnh về relation mới. | Knowledge Distillation | |
| Representation distillation loss trên hidden representations. | Giữ hidden representation hiện tại gần previous model. | Knowledge Distillation | |
| Distillation triplet loss với hard positive/negative từ pseudo + real samples. | Thành phần contrastive quan trọng nhất trong ablation. | Contrastive Learning | |
| Prediction distillation loss trên old relation logits với temperature. | Gần với LwF ở mức khớp soft prediction cho relation cũ. | Knowledge Distillation |
Phase 1 — Paper Map
Paper map — scaffold từ nguồn
- Problem: Continual few-shot RE phải học relation mới với rất ít labeled examples, đồng thời giữ khả năng phân loại relation cũ. PDF tr. 1
- Motivation: Few-shot samples dễ không đại diện cho relation mới, gây overfitting và làm forgetting nặng hơn; relation representations cũng dễ bị confusion khi số relation tăng. PDF tr. 1, PDF tr. 2
- Gap: Existing memory/prototype/contrastive methods vẫn phụ thuộc prototype từ typical samples, chưa giữ đủ khác biệt sample-level giữa relation, hoặc cần extra data/knowledge. PDF tr. 2, PDF tr. 3
- Main idea: Kết hợp memory samples, prototype-based pseudo samples, bidirectional entity-replacement augmentation, và serial distillation gồm feature, hidden contrastive, prediction distillation. PDF tr. 3, PDF tr. 5
- Important figure: Figure 2 mô tả serial contrastive KD qua previous/current model, pseudo samples, feature/hidden/prediction distillation. PDF tr. 5
- Important equations: Eq. 1-2 encoder/classification; Eq. 3 prototype; Eq. 4 ; Eq. 5 pseudo sample; Eq. 6-8 hidden contrastive; Eq. 9-12 prediction/total loss; Eq. 13 BWT. PDF tr. 4, PDF tr. 6, PDF tr. 8
- Main result table: Table 1 FewRel/TACRED 5-shot; SCKD tốt nhất sau final task. PDF tr. 7
- Ablation table: Table 2 module ablation; Table 3 fine-grained distillation loss ablation. PDF tr. 7, PDF tr. 8
- Limitations: Memory-based, tốn storage; mới evaluated trong RE setting, chưa kiểm ở continual few-shot tasks khác. PDF tr. 9
Chỗ cần đọc trước
- Section 3.1 task definition
- Algorithm 1
- Figure 2
- Eq. 4-12
- Table 1-3
- Appendix A hyperparameters
Phase 2 — Pass 1 Recall
Closed-book recall của tôi
Problem
Why does it matter?
Research gap
Main idea
Main contribution
Main result
Prompt kiểm tra recall
I have completed the first pass of SCKD.
Here is my understanding:
[PASTE MY NOTES]
Compare my understanding against the paper. Identify correct, inaccurate, missing, confusing concepts, and sections/equations/tables to revisit.Phase 3 — Problem / Motivation / Gap
| Mục | Diễn giải bằng lời của tôi | Evidence / citation |
|---|---|---|
| General problem | Continual RE phải học relation mới tuần tự và vẫn classify được relation cũ. | PDF tr. 1 |
| Why it matters | Emerging relations thường ít labeled samples; retraining từ đầu không thực tế. | PDF tr. 1 |
| What prior work solves | Memory-based continual RE lưu exemplar/prototype để giảm forgetting. | PDF tr. 1 |
| What prior work fails to solve | Prototype/sample memory có thể không đủ đại diện; contrastive/prototype methods chưa giữ đủ khoảng cách sample-level hoặc cần extra data/knowledge. | PDF tr. 2 |
| Exact research gap | Cần vừa preserve prior model knowledge vừa giữ representations của relation khác nhau tách biệt trong few-shot continual RE. | PDF tr. 2 |
| Hypothesis / intuition | Serial KD + contrastive pseudo samples sẽ giảm forgetting; augmentation giữa memory/current task giảm overfitting. | PDF tr. 2 |
| Contribution addressing the gap | SCKD với feature, hidden contrastive, prediction distillation; entity replacement augmentation; evaluation strict trên FewRel/TACRED. | PDF tr. 6 |
Phase 4 — Method / Architecture
Tôi tự vẽ trước
Task T_j data D_j + previous memory M_{j-1}
-> initialize Phi_j from Phi_{j-1}
-> adapt on D_j
-> select L typical samples/relation by k-means
-> update memory M_j and relation set R_j
-> compute prototypes for observed relations
-> bidirectional entity replacement augmentation
-> generate pseudo samples around prototypes
-> serial contrastive KD on augmented current data
-> serial contrastive KD on augmented memory replay
-> evaluate on all observed tasksComponent map
| Component | Input | Operation | Output | Purpose | Evidence |
|---|---|---|---|---|---|
| BERT encoder + entity markers | sentence with entity mentions | encode [E1]/[E2], concatenate token representations | sample feature | relation-aware text representation | PDF tr. 4 |
| Hidden projection | dropout, linear layer, layer norm | hidden | representation used by classifier/prototype | PDF tr. 4 | |
| Memory selection | current task samples | k-means, closest to centroids | typical samples/relation | compact replay memory | PDF tr. 4 |
| Prototype generation | memory samples | average hidden representations | relation anchor and pseudo sample center | PDF tr. 4 | |
| Data augmentation | replace similar entities bidirectionally | reduce few-shot overfitting | PDF tr. 4 | ||
| Serial KD | previous/current model + data + pseudo samples | updated current model | preserve old knowledge and separate relation representations | PDF tr. 5 |
Điều tôi vẫn chưa hiểu
- Pseudo covariance được lưu/freeze tại thời điểm relation xuất hiện hay recompute qua tasks?
- Distillation được áp dụng trên current augmented data và memory replay khác nhau thế nào trong code?
- SCKD có dùng relation descriptions không? Paper nói không, nhưng cần kiểm code nếu reproduce.
Phase 6 — Equations
Prompt operational equation walkthrough
Walk me through the key equations or formal blocks in this paper.
For each equation/block, explain input, output, pipeline location, behavior encouraged, failure mode if removed, and supporting table/figure/ablation.
Focus only on operational understanding of equations and formal mechanisms.Equation queue
| Eq. | Dùng để làm gì? | Biến chính | Behavior được khuyến khích | Evidence / ablation | Status |
|---|---|---|---|---|---|
| Eq. 1 | Project BERT entity features thành hidden representation. | representation ổn định cho classifier/prototype. | PDF tr. 4 | scaffold | |
| Eq. 2 | Classification loss cho current task. | học relation hiện tại. | Final loss Eq. 12; PDF tr. 4 | scaffold | |
| Eq. 3 | Prototype bằng mean hidden của memory samples. | tạo anchor relation. | Pseudo samples Eq. 5; PDF tr. 4 | scaffold | |
| Eq. 4 | Feature distillation . | encoder current không lệch khỏi previous. | Table 3 w/o ; PDF tr. 8 | scaffold | |
| Eq. 5 | Sinh pseudo sample quanh prototype. | tăng contrastive samples cho few-shot relation. | Figure 2, Table 2/3; PDF tr. 5 | scaffold | |
| Eq. 6 | Representation distillation . | giữ hidden representation giống previous model. | Table 3 w/o ; PDF tr. 8 | scaffold | |
| Eq. 7 | Distillation triplet loss . | kéo hard positive gần, hard negative xa. | Ablation drop lớn nhất; PDF tr. 8 | scaffold | |
| Eq. 9-10 | Prediction distillation . | old/current logits, temperature | khớp soft distribution trên old relations. | Table 3 w/o ; PDF tr. 6 | scaffold |
| Eq. 11-12 | Total distillation/final loss. | cân bằng classification và preservation. | Appendix A hyperparameters; PDF tr. 11 | scaffold | |
| Eq. 13 | Backward transfer. | đo forgetting sau final task. | Figure 3; PDF tr. 8 | scaffold |
Phase 7 — Loss Functions
Total Loss L
├── L_csf -> học current relation classification
└── L_dst
├── L_fd -> feature distillation
├── L_hcd
│ ├── L_rd -> hidden representation distillation
│ └── L_dtr -> contrastive/triplet separation
└── L_pd -> prediction distillation on old relations| Loss | Equation | Inputs | Trains component | Behavior | Weight | Ablation |
|---|---|---|---|---|---|---|
| Eq. 2 | current task samples | encoder/classifier | learn current task | not isolated | ||
| Eq. 4 | features from previous/current model | encoder | reduce feature drift | Table 3 | ||
| Eq. 6 | hidden reps previous/current | projection/dropout | preserve hidden reps | inside | Table 3 | |
| Eq. 7 | real+pseudo samples | hidden space | hard positive/negative separation | inside | largest drop when removed | |
| Eq. 9-10 | old/current logits, | classifier/head | preserve old relation probability behavior | Table 3 | ||
| Eq. 11 | three KD losses | full model | preserve prior knowledge | Table 2 w/o dst. |
Phase 8 — Experiments
Experiment map
| Experiment | Research question | Dataset | Baselines | Metric | Table/Figure | Main result | Caveat |
|---|---|---|---|---|---|---|---|
| Main 5-shot | SCKD có hơn continual/few-shot RE baselines không? | FewRel 10-way-5-shot, TACRED 5-way-5-shot | Finetune, Joint-train, RP-CRE, CRL, CRECL, ERDA | average accuracy | Table 1 | Final task: FewRel 62.98, TACRED 52.11, tốt nhất among practical baselines. | Joint-train không luôn upper bound do imbalance. |
| Module ablation | Distillation và augmentation đóng góp thế nào? | FewRel/TACRED | w/o dst., w/o aug., w/o both | average accuracy | Table 2 | Bỏ distillation giảm mạnh hơn bỏ augmentation. | Augmentation gain nhỏ trong main table. |
| Fine-grained KD ablation | Loss KD nào quan trọng nhất? | FewRel/TACRED | w/o , , , | average accuracy | Table 3 | Bỏ drop rõ nhất. | Need inspect code for exact implementation. |
| Few-shot RE comparison | So với few-shot RE classic khi adapted? | FewRel/TACRED current tasks | GNN, Proto, BERT-PAIR | current task accuracy | Table 4 | SCKD vượt các few-shot RE models. | Setup chuyển đổi support/query hơi đặc thù. |
| BWT / t-SNE | Có giảm forgetting và tách representation không? | FewRel/TACRED | RP-CRE, CRL, CRECL, ERDA | BWT, visualization | Figure 3-4 | BWT ít âm nhất; clusters tách hơn CRECL. | t-SNE chỉ minh họa. |
Protocol fingerprint
- Dataset and split: FewRel 80 relations thành 8 tasks x 10 relations; TACRED bỏ
no_relation, 41 relations thành 8 tasks. - Scenario / label space: first task có 100 samples/relation; later tasks 5-shot hoặc 10-shot.
- Backbone: BERT encoder với entity marker features; hidden dim 768.
- Memory budget: main experiments dùng sample/relation.
- Seeds / number of runs: 6 random seeds, same task sequence as ERDA.
- Metric and averaging: ; means/stds reported.
- Evaluation: strict evaluation với toàn bộ observed relation labels làm negatives.
- Optimizer/hyperparameters: Adam; batch size 16; encoder lr 1e-5; classifier lr 1e-3; dropout 0.5; pseudo samples/relation 10; ; Appendix A.
Phase 9 — Claim → Evidence
| Claim | Where claim appears | Experiment | Evidence | My judgment | Caveat |
|---|---|---|---|---|---|
| SCKD giảm forgetting và overfitting trong continual few-shot RE. | Abstract/Intro | Main results + BWT | Table 1, Figure 3 | Strong trong FewRel/TACRED setup. | Memory-based, vẫn cần exemplar. |
| Serial contrastive KD là module quan trọng nhất. | Section 4.2.2 | module ablation | w/o dst. giảm final FewRel 62.98 → 58.96; TACRED 52.11 → 46.52. | Strong. | Table uses averaged reported values; no reproduction. |
| Distillation triplet loss quan trọng nhất trong KD losses. | Section 4.2.2 | fine-grained ablation | w/o final FewRel 61.18, TACRED 50.94; drop lớn nhất. | Strong relative to other KD loss ablations. | Need verify statistical significance. |
| Pseudo samples giúp representation phân biệt hơn. | Section 4.2.3-4.2.5 | few-shot RE comparison/t-SNE | Table 4, Figure 4 | Plausible. | t-SNE visual không đủ độc lập để prove. |
Phase 10 — Ablation Study
| Component | Intended purpose | With component | Without component | Difference | Conclusion justified | Not justified |
|---|---|---|---|---|---|---|
| Serial KD module | preserve prior knowledge + contrastive separation | FewRel T8 62.98; TACRED T8 52.11 | 58.96; 46.52 | -4.02; -5.59 | KD module cốt lõi. | Không tách được từng subloss ở Table 2. |
| Augmentation | giảm overfitting few-shot | 62.98; 52.11 | 62.51; 51.79 | -0.47; -0.32 | có ích nhưng nhỏ hơn KD. | Không chứng minh augmentation luôn đáng cost. |
| relation separation | 62.98; 52.11 | 61.18; 50.94 | -1.80; -1.17 | triplet contrastive quan trọng nhất trong sublosses. | Không thay thế full contrastive alternatives. | |
| Memory size | kiểm tác động | SCKD maintains best | N/A | Table 5/7 | SCKD tận dụng memory tốt. | Main result vẫn phụ thuộc memory. |
Phase 11 — Critical Reading
Reviewer notes
- Strongest contribution: nối KD với contrastive separation ở cả feature/hidden/prediction level cho CFRE.
- Weakest part: vẫn memory-based; privacy/storage claim yếu nếu bài toán cần rehearsal-free.
- Main assumption: prototype-centered Gaussian pseudo samples phản ánh đủ local distribution của relation.
- Alternative explanation: performance gain có thể đến từ memory usage/hyperparameter matching hơn là bản thân serial KD; cần code-level reproduction.
- Missing experiment: compare trực tiếp với LwF-style prediction-only KD trong cùng CFRE protocol.
- Generalization risk: FewRel/TACRED task construction là variant; không đồng nhất với original FewRel few-shot evaluation.
- Reproducibility risk: code dependencies cũ PyTorch 1.7.1/Transformers 2.11.0; entity replacement details cần kiểm code.
Phase 12 — Reproduction Check
| Item | Status | Detail | Missing detail / risk |
|---|---|---|---|
| Dataset and split | Clearly specified | FewRel/TACRED task construction | cần exact task order seeds/files |
| Preprocessing | Partially specified | entity markers and BERT token reps | tokenization/code details cần kiểm |
| Input representation | Clearly specified | concatenate [E1]/[E2] BERT reps | entity marker insertion exact format |
| Model / backbone | Clearly specified | BERT + dropout/projection/classifier | model checkpoint not explicitly named in text |
| Sampling procedure | Partially specified | k-means typical samples, | k-means seed/details |
| Memory / replay strategy | Clearly specified | one sample/relation main experiments | privacy/storage caveat |
| Loss functions | Clearly specified | Eq. 2,4,6,7,9-12 | implementation details for hard mining |
| Optimizer | Clearly specified | Adam | version sensitivity |
| Learning rate | Clearly specified | encoder/dropout 1e-5, classifier 1e-3 | schedule not deeply detailed |
| Batch size | Clearly specified | 16, grad accumulation 4 | GPU memory dependent |
| Hyperparameters | Clearly specified | Appendix A Table 6 | grid search cost |
| Evaluation protocol | Clearly specified | strict observed labels | compare with loose results carefully |
Phase 13 — Completeness / Oral Exam
Completion criteria
- Giải thích SCKD khác LwF ở feature/hidden/prediction level.
- Vẽ được Algorithm 1 từ input task đến memory update.
- Phân biệt , , , .
- Nói được pseudo samples sinh từ prototype và covariance nào.
- Đọc được Table 1-3 và kết luận ablation.
- Nêu được limitation memory-based và transfer beyond RE.
Prompt oral exam
Act as my PhD advisor. Quiz me one question at a time about SCKD: task definition, Algorithm 1, Eq. 4-12, protocol, ablations, limitations, and relation to LwF/Mean Teacher/TAPTA. Wait for my answer before feedback.Final Paper Note Handoff
- Tóm tắt một câu:
- Problem:
- Gap:
- Method overview:
- Important equations:
- Protocol fingerprint:
- Main results:
- Ablation:
- Limitations:
- Critical judgment:
- Concepts cần tạo/cập nhật:
Liên kết
- Paper note: Serial Contrastive Knowledge Distillation for Continual Few-shot Relation Extraction
- PDF: Serial Contrastive Knowledge Distillation for Continual Few-shot Relation Extraction.pdf
- Concepts: Continual Few-Shot Relation Extraction, Knowledge Distillation, Contrastive Learning, Replay in Continual Learning, Prototype Learning
- Related papers: Learning without Forgetting, Mean Teachers are Better Role Models, Continual Few-shot Relation Learning via Embedding Space Regularization and Data Augmentation