2026-09-23 - SCKD - Gemini Notebook Workflow

Ranh giới

Đây là working note scaffold từ paper note/PDF. Các phần factual đã gắn citation; closed-book recall, oral exam answer và kết luận cá nhân để bạn tự điền sau khi đọc.

Setup

Từ điển khái niệm nhanh

Khái niệmĐịnh nghĩa ngắn trong paper nàyVì sao quan trọngLink
Continual few-shot RETask đầu có data nhiều, các task sau chỉ N-way K-shot; sau task đánh giá trên tất cả relation đã thấy.Là setting chính mà SCKD tối ưu.Continual Few-Shot Relation Extraction
SCKDSerial Contrastive Knowledge Distillation: feature/prediction/hidden contrastive distillation theo chuỗi model trước-sau.Đóng góp chính để giảm forgetting và overfitting.Serial Contrastive Knowledge Distillation for Continual Few-shot Relation Extraction
Memory Bộ nhớ lưu vài typical samples từ các relation cũ và mới.SCKD là memory-based method; memory size mặc định một sample/relation trong main experiments.Replay in Continual Learning
Prototype Mean hidden representation của typical samples cho relation .Dùng để sinh pseudo samples cho contrastive distillation.Prototype Learning
Pseudo samples Vector giả sinh quanh prototype với Gaussian noise và diagonal covariance.Tăng tín hiệu contrastive khi few-shot data ít.Contrastive Learning
Feature distillation loss giữa feature của previous model và current model.Giữ encoder không lệch quá mạnh về relation mới.Knowledge Distillation
Representation distillation loss trên hidden representations.Giữ hidden representation hiện tại gần previous model.Knowledge Distillation
Distillation triplet loss với hard positive/negative từ pseudo + real samples.Thành phần contrastive quan trọng nhất trong ablation.Contrastive Learning
Prediction distillation loss trên old relation logits với temperature.Gần với LwF ở mức khớp soft prediction cho relation cũ.Knowledge Distillation

Phase 1 — Paper Map

Paper map — scaffold từ nguồn

  • Problem: Continual few-shot RE phải học relation mới với rất ít labeled examples, đồng thời giữ khả năng phân loại relation cũ. PDF tr. 1
  • Motivation: Few-shot samples dễ không đại diện cho relation mới, gây overfitting và làm forgetting nặng hơn; relation representations cũng dễ bị confusion khi số relation tăng. PDF tr. 1, PDF tr. 2
  • Gap: Existing memory/prototype/contrastive methods vẫn phụ thuộc prototype từ typical samples, chưa giữ đủ khác biệt sample-level giữa relation, hoặc cần extra data/knowledge. PDF tr. 2, PDF tr. 3
  • Main idea: Kết hợp memory samples, prototype-based pseudo samples, bidirectional entity-replacement augmentation, và serial distillation gồm feature, hidden contrastive, prediction distillation. PDF tr. 3, PDF tr. 5
  • Important figure: Figure 2 mô tả serial contrastive KD qua previous/current model, pseudo samples, feature/hidden/prediction distillation. PDF tr. 5
  • Important equations: Eq. 1-2 encoder/classification; Eq. 3 prototype; Eq. 4 ; Eq. 5 pseudo sample; Eq. 6-8 hidden contrastive; Eq. 9-12 prediction/total loss; Eq. 13 BWT. PDF tr. 4, PDF tr. 6, PDF tr. 8
  • Main result table: Table 1 FewRel/TACRED 5-shot; SCKD tốt nhất sau final task. PDF tr. 7
  • Ablation table: Table 2 module ablation; Table 3 fine-grained distillation loss ablation. PDF tr. 7, PDF tr. 8
  • Limitations: Memory-based, tốn storage; mới evaluated trong RE setting, chưa kiểm ở continual few-shot tasks khác. PDF tr. 9

Chỗ cần đọc trước

  • Section 3.1 task definition
  • Algorithm 1
  • Figure 2
  • Eq. 4-12
  • Table 1-3
  • Appendix A hyperparameters

Phase 2 — Pass 1 Recall

Closed-book recall của tôi

Problem

Why does it matter?

Research gap

Main idea

Main contribution

Main result

Prompt kiểm tra recall

I have completed the first pass of SCKD.
 
Here is my understanding:
 
[PASTE MY NOTES]
 
Compare my understanding against the paper. Identify correct, inaccurate, missing, confusing concepts, and sections/equations/tables to revisit.

Phase 3 — Problem / Motivation / Gap

MụcDiễn giải bằng lời của tôiEvidence / citation
General problemContinual RE phải học relation mới tuần tự và vẫn classify được relation cũ.PDF tr. 1
Why it mattersEmerging relations thường ít labeled samples; retraining từ đầu không thực tế.PDF tr. 1
What prior work solvesMemory-based continual RE lưu exemplar/prototype để giảm forgetting.PDF tr. 1
What prior work fails to solvePrototype/sample memory có thể không đủ đại diện; contrastive/prototype methods chưa giữ đủ khoảng cách sample-level hoặc cần extra data/knowledge.PDF tr. 2
Exact research gapCần vừa preserve prior model knowledge vừa giữ representations của relation khác nhau tách biệt trong few-shot continual RE.PDF tr. 2
Hypothesis / intuitionSerial KD + contrastive pseudo samples sẽ giảm forgetting; augmentation giữa memory/current task giảm overfitting.PDF tr. 2
Contribution addressing the gapSCKD với feature, hidden contrastive, prediction distillation; entity replacement augmentation; evaluation strict trên FewRel/TACRED.PDF tr. 6

Phase 4 — Method / Architecture

Tôi tự vẽ trước

Task T_j data D_j + previous memory M_{j-1}
-> initialize Phi_j from Phi_{j-1}
-> adapt on D_j
-> select L typical samples/relation by k-means
-> update memory M_j and relation set R_j
-> compute prototypes for observed relations
-> bidirectional entity replacement augmentation
-> generate pseudo samples around prototypes
-> serial contrastive KD on augmented current data
-> serial contrastive KD on augmented memory replay
-> evaluate on all observed tasks

Component map

ComponentInputOperationOutputPurposeEvidence
BERT encoder + entity markerssentence with entity mentionsencode [E1]/[E2], concatenate token representationssample feature relation-aware text representationPDF tr. 4
Hidden projectiondropout, linear layer, layer normhidden representation used by classifier/prototypePDF tr. 4
Memory selectioncurrent task samplesk-means, closest to centroids typical samples/relationcompact replay memoryPDF tr. 4
Prototype generationmemory samplesaverage hidden representationsrelation anchor and pseudo sample centerPDF tr. 4
Data augmentationreplace similar entities bidirectionallyreduce few-shot overfittingPDF tr. 4
Serial KDprevious/current model + data + pseudo samplesupdated current modelpreserve old knowledge and separate relation representationsPDF tr. 5

Điều tôi vẫn chưa hiểu

  • Pseudo covariance được lưu/freeze tại thời điểm relation xuất hiện hay recompute qua tasks?
  • Distillation được áp dụng trên current augmented data và memory replay khác nhau thế nào trong code?
  • SCKD có dùng relation descriptions không? Paper nói không, nhưng cần kiểm code nếu reproduce.

Phase 6 — Equations

Prompt operational equation walkthrough

Walk me through the key equations or formal blocks in this paper.
 
For each equation/block, explain input, output, pipeline location, behavior encouraged, failure mode if removed, and supporting table/figure/ablation.
Focus only on operational understanding of equations and formal mechanisms.

Equation queue

Eq.Dùng để làm gì?Biến chínhBehavior được khuyến khíchEvidence / ablationStatus
Eq. 1Project BERT entity features thành hidden representation.representation ổn định cho classifier/prototype.PDF tr. 4scaffold
Eq. 2Classification loss cho current task.học relation hiện tại.Final loss Eq. 12; PDF tr. 4scaffold
Eq. 3Prototype bằng mean hidden của memory samples.tạo anchor relation.Pseudo samples Eq. 5; PDF tr. 4scaffold
Eq. 4Feature distillation .encoder current không lệch khỏi previous.Table 3 w/o ; PDF tr. 8scaffold
Eq. 5Sinh pseudo sample quanh prototype.tăng contrastive samples cho few-shot relation.Figure 2, Table 2/3; PDF tr. 5scaffold
Eq. 6Representation distillation .giữ hidden representation giống previous model.Table 3 w/o ; PDF tr. 8scaffold
Eq. 7Distillation triplet loss .kéo hard positive gần, hard negative xa.Ablation drop lớn nhất; PDF tr. 8scaffold
Eq. 9-10Prediction distillation .old/current logits, temperature khớp soft distribution trên old relations.Table 3 w/o ; PDF tr. 6scaffold
Eq. 11-12Total distillation/final loss.cân bằng classification và preservation.Appendix A hyperparameters; PDF tr. 11scaffold
Eq. 13Backward transfer.đo forgetting sau final task.Figure 3; PDF tr. 8scaffold

Phase 7 — Loss Functions

Total Loss L
├── L_csf -> học current relation classification
└── L_dst
    ├── L_fd -> feature distillation
    ├── L_hcd
    │   ├── L_rd -> hidden representation distillation
    │   └── L_dtr -> contrastive/triplet separation
    └── L_pd -> prediction distillation on old relations
LossEquationInputsTrains componentBehaviorWeightAblation
Eq. 2current task samplesencoder/classifierlearn current tasknot isolated
Eq. 4features from previous/current modelencoderreduce feature driftTable 3
Eq. 6hidden reps previous/currentprojection/dropoutpreserve hidden repsinside Table 3
Eq. 7real+pseudo sampleshidden spacehard positive/negative separationinside largest drop when removed
Eq. 9-10old/current logits, classifier/headpreserve old relation probability behaviorTable 3
Eq. 11three KD lossesfull modelpreserve prior knowledgeTable 2 w/o dst.

Phase 8 — Experiments

Experiment map

ExperimentResearch questionDatasetBaselinesMetricTable/FigureMain resultCaveat
Main 5-shotSCKD có hơn continual/few-shot RE baselines không?FewRel 10-way-5-shot, TACRED 5-way-5-shotFinetune, Joint-train, RP-CRE, CRL, CRECL, ERDAaverage accuracyTable 1Final task: FewRel 62.98, TACRED 52.11, tốt nhất among practical baselines.Joint-train không luôn upper bound do imbalance.
Module ablationDistillation và augmentation đóng góp thế nào?FewRel/TACREDw/o dst., w/o aug., w/o bothaverage accuracyTable 2Bỏ distillation giảm mạnh hơn bỏ augmentation.Augmentation gain nhỏ trong main table.
Fine-grained KD ablationLoss KD nào quan trọng nhất?FewRel/TACREDw/o , , , average accuracyTable 3Bỏ drop rõ nhất.Need inspect code for exact implementation.
Few-shot RE comparisonSo với few-shot RE classic khi adapted?FewRel/TACRED current tasksGNN, Proto, BERT-PAIRcurrent task accuracyTable 4SCKD vượt các few-shot RE models.Setup chuyển đổi support/query hơi đặc thù.
BWT / t-SNECó giảm forgetting và tách representation không?FewRel/TACREDRP-CRE, CRL, CRECL, ERDABWT, visualizationFigure 3-4BWT ít âm nhất; clusters tách hơn CRECL.t-SNE chỉ minh họa.

Protocol fingerprint

  • Dataset and split: FewRel 80 relations thành 8 tasks x 10 relations; TACRED bỏ no_relation, 41 relations thành 8 tasks.
  • Scenario / label space: first task có 100 samples/relation; later tasks 5-shot hoặc 10-shot.
  • Backbone: BERT encoder với entity marker features; hidden dim 768.
  • Memory budget: main experiments dùng sample/relation.
  • Seeds / number of runs: 6 random seeds, same task sequence as ERDA.
  • Metric and averaging: ; means/stds reported.
  • Evaluation: strict evaluation với toàn bộ observed relation labels làm negatives.
  • Optimizer/hyperparameters: Adam; batch size 16; encoder lr 1e-5; classifier lr 1e-3; dropout 0.5; pseudo samples/relation 10; ; Appendix A.

Phase 9 — Claim → Evidence

ClaimWhere claim appearsExperimentEvidenceMy judgmentCaveat
SCKD giảm forgetting và overfitting trong continual few-shot RE.Abstract/IntroMain results + BWTTable 1, Figure 3Strong trong FewRel/TACRED setup.Memory-based, vẫn cần exemplar.
Serial contrastive KD là module quan trọng nhất.Section 4.2.2module ablationw/o dst. giảm final FewRel 62.98 → 58.96; TACRED 52.11 → 46.52.Strong.Table uses averaged reported values; no reproduction.
Distillation triplet loss quan trọng nhất trong KD losses.Section 4.2.2fine-grained ablationw/o final FewRel 61.18, TACRED 50.94; drop lớn nhất.Strong relative to other KD loss ablations.Need verify statistical significance.
Pseudo samples giúp representation phân biệt hơn.Section 4.2.3-4.2.5few-shot RE comparison/t-SNETable 4, Figure 4Plausible.t-SNE visual không đủ độc lập để prove.

Phase 10 — Ablation Study

ComponentIntended purposeWith componentWithout componentDifferenceConclusion justifiedNot justified
Serial KD modulepreserve prior knowledge + contrastive separationFewRel T8 62.98; TACRED T8 52.1158.96; 46.52-4.02; -5.59KD module cốt lõi.Không tách được từng subloss ở Table 2.
Augmentationgiảm overfitting few-shot62.98; 52.1162.51; 51.79-0.47; -0.32có ích nhưng nhỏ hơn KD.Không chứng minh augmentation luôn đáng cost.
relation separation62.98; 52.1161.18; 50.94-1.80; -1.17triplet contrastive quan trọng nhất trong sublosses.Không thay thế full contrastive alternatives.
Memory sizekiểm tác động SCKD maintains bestN/ATable 5/7SCKD tận dụng memory tốt.Main result vẫn phụ thuộc memory.

Phase 11 — Critical Reading

Reviewer notes

  • Strongest contribution: nối KD với contrastive separation ở cả feature/hidden/prediction level cho CFRE.
  • Weakest part: vẫn memory-based; privacy/storage claim yếu nếu bài toán cần rehearsal-free.
  • Main assumption: prototype-centered Gaussian pseudo samples phản ánh đủ local distribution của relation.
  • Alternative explanation: performance gain có thể đến từ memory usage/hyperparameter matching hơn là bản thân serial KD; cần code-level reproduction.
  • Missing experiment: compare trực tiếp với LwF-style prediction-only KD trong cùng CFRE protocol.
  • Generalization risk: FewRel/TACRED task construction là variant; không đồng nhất với original FewRel few-shot evaluation.
  • Reproducibility risk: code dependencies cũ PyTorch 1.7.1/Transformers 2.11.0; entity replacement details cần kiểm code.

Phase 12 — Reproduction Check

ItemStatusDetailMissing detail / risk
Dataset and splitClearly specifiedFewRel/TACRED task constructioncần exact task order seeds/files
PreprocessingPartially specifiedentity markers and BERT token repstokenization/code details cần kiểm
Input representationClearly specifiedconcatenate [E1]/[E2] BERT repsentity marker insertion exact format
Model / backboneClearly specifiedBERT + dropout/projection/classifiermodel checkpoint not explicitly named in text
Sampling procedurePartially specifiedk-means typical samples, k-means seed/details
Memory / replay strategyClearly specifiedone sample/relation main experimentsprivacy/storage caveat
Loss functionsClearly specifiedEq. 2,4,6,7,9-12implementation details for hard mining
OptimizerClearly specifiedAdamversion sensitivity
Learning rateClearly specifiedencoder/dropout 1e-5, classifier 1e-3schedule not deeply detailed
Batch sizeClearly specified16, grad accumulation 4GPU memory dependent
HyperparametersClearly specifiedAppendix A Table 6grid search cost
Evaluation protocolClearly specifiedstrict observed labelscompare with loose results carefully

Phase 13 — Completeness / Oral Exam

Completion criteria

  • Giải thích SCKD khác LwF ở feature/hidden/prediction level.
  • Vẽ được Algorithm 1 từ input task đến memory update.
  • Phân biệt , , , .
  • Nói được pseudo samples sinh từ prototype và covariance nào.
  • Đọc được Table 1-3 và kết luận ablation.
  • Nêu được limitation memory-based và transfer beyond RE.

Prompt oral exam

Act as my PhD advisor. Quiz me one question at a time about SCKD: task definition, Algorithm 1, Eq. 4-12, protocol, ablations, limitations, and relation to LwF/Mean Teacher/TAPTA. Wait for my answer before feedback.

Final Paper Note Handoff

  • Tóm tắt một câu:
  • Problem:
  • Gap:
  • Method overview:
  • Important equations:
  • Protocol fingerprint:
  • Main results:
  • Ablation:
  • Limitations:
  • Critical judgment:
  • Concepts cần tạo/cập nhật:

Liên kết