2026-08-21 - WAVE++ - Gemini Notebook Workflow

Ranh giới

Đây là working note scaffold từ paper note/PDF để hỗ trợ đọc với Gemini Notebook. Các phần “closed-book recall”, “mình tự trả lời”, “oral exam” vẫn để trống vì chưa có câu trả lời cá nhân của bạn; không coi note này là bằng chứng đã đọc xong paper.

Setup

Từ điển khái niệm nhanh

Khái niệmĐịnh nghĩa ngắn trong paper nàyVì sao quan trọngLink
CREContinual Relation Extraction: relation label space mở rộng theo task và model phải nhớ relation cũ.Đây là bài toán nền của WAVE++.Continual Relation Extraction
WAVE++Bản mở rộng của WAVE-CRE với adaptive prompt pool, label-description alignment, cascade voting và latent replay.Là paper cần đọc như extension chứ không phải duplicate của WAVE-CRE.WAVE++ - Capturing Within-Task Variance for Continual Relation Extraction
Within-task varianceSự đa dạng nội bộ của examples trong cùng task/relation group.Là động cơ chính cho nhiều prompt experts thay vì một prompt/task.WAVE++ - Capturing Within-Task Variance for Continual Relation Extraction
Prompt poolTập prompt/prefix experts để route input vào adapter phù hợp.Giữ khả năng thích nghi với nhiều mode trong task.Prompt Pool
Label descriptionMô tả bằng ngôn ngữ tự nhiên của relation label.WAVE++ dùng label semantics làm anchor cho relation/task inference.Relation Extraction
Label-description alignmentCăn representation của input với representation/mô tả label liên quan.Giúp giảm nhầm relation gần nghĩa và tận dụng semantics của label.Embedding Space Regularization
Cascade votingCơ chế quyết định theo nhiều bước/votes để chọn task hoặc relation ổn định hơn.Nhắm vào lỗi task prediction/inference trong CRE.Continual Relation Extraction
Latent generative replayReplay synthetic latent vectors của relation cũ thay vì lưu đầy đủ raw examples.Bảo vệ classifier trước bias task mới với ít raw memory hơn.Replay in Continual Learning
Statistical significanceKiểm định xem improvement có ổn định hay chỉ do variance giữa runs.Appendix WAVE++ dùng để củng cố claim kết quả.Model Benchmarking
Running timeChi phí train/inference khi thêm prompt pool, voting và replay.Cần đọc cùng accuracy để đánh giá trade-off thực tế.Transformer Inference Optimization

Phase 1 - Paper Map

Prompt gửi Gemini

Do not summarize the paper in detail yet. Create a structural map of this paper and identify problem, motivation, gap, contributions, pipeline, components, losses, datasets, baselines, metrics, main experiments, ablations, appendix analyses, and limitations. For every item, point to the relevant section, figure, table, or equation. The purpose is to tell me WHERE to read, not to replace my reading.

Paper map - scaffold từ nguồn

  • Problem: continual relation extraction cần học nhiều task liên tiếp và giữ performance trên relation cũ. PDF tr. 4
  • Motivation: WAVE-CRE nêu within-task variance; WAVE++ thêm label semantics và cascade voting để cải thiện task/relation inference.
  • Gap: prompt pool đơn thuần chưa tận dụng đủ label descriptions và task prediction vẫn là nguồn lỗi.
  • Main idea: kết hợp prompt pool, label-description alignment, cascade voting và latent generative replay.
  • Main contributions: adaptive prompt pool, label description contrastive alignment, cascade voting, stronger experiments/appendix.
  • Important figure/equations: formalization và method ở đầu phần method. PDF tr. 4, PDF tr. 5
  • Main result table: final-stage results. PDF tr. 16
  • Ablation table: prompt pool, label descriptions, generative replay. PDF tr. 18
  • Appendix cần đọc: label descriptions, statistical tests, task prediction, running time. PDF tr. 27, PDF tr. 30

Chỗ cần đọc trước

  • Abstract + Introduction
  • Problem formalization
  • Prompt pool
  • Label-description alignment
  • Cascade voting
  • Generative replay
  • Table main results
  • Ablation + appendix

Phase 2 - Pass 1 Recall

Closed-book recall của tôi

Problem

Why does it matter?

Research gap

Main idea

Main contribution

Main result

Prompt kiểm tra recall

I have completed the first pass of the paper.
 
Here is my understanding:
 
[PASTE MY NOTES]
 
Compare my understanding against the paper. Return what is correct, inaccurate, missing, confusing, and which sections/citations I should revisit. Do not rewrite the entire paper for me.

Phase 3 - Problem / Motivation / Gap

MụcDiễn giải bằng lời của tôiEvidence / citation
General problemCRE là class-incremental RE, nơi label space mở rộng theo task và model phải phân loại trên labels đã thấy.PDF tr. 4
Why it mattersDữ liệu quan hệ cũ khó giữ đầy đủ, còn relation mới xuất hiện liên tục.PDF tr. 2
What prior work solvesWAVE-CRE dùng prompt pools và latent replay để bắt within-task variance và giảm forgetting.PDF tr. 16
What prior work fails to solveTask prediction và label semantics vẫn có thể làm relation gần nghĩa bị nhầm.inferred
Exact research gapCần khai thác label descriptions và cascade decision để chọn task/relation ổn định hơn.PDF tr. 5
Hypothesis / intuitionLabel anchors giúp representation của input có điểm neo ngữ nghĩa; cascade voting giảm lỗi chọn task.inferred
Contribution addressing the gapWAVE++ thêm label-description alignment và cascade voting lên nền WAVE-CRE.PDF tr. 5

Câu hỏi tự kiểm tra

  • WAVE++ thêm gì không có trong WAVE-CRE?
  • Label description giúp relation classifier hay task predictor?
  • Cascade voting có đổi inference assumption không?

Phase 4 - Method / Architecture

Tôi tự vẽ trước

Sentence + entity pair
-> PLM + task-specific prompt pool
-> input representation
-> label description representation
-> contrastive/alignment objective
-> cascade voting for task/relation candidate
-> relation classifier over seen relations
-> latent generative replay for old relation features

Component map

ComponentInputOperationOutputPurposeEvidence
Prompt poolinput + task/prompt keysroute input to prompt expertsprompted representationcapture within-task variancePDF tr. 5
Label descriptionsrelation label text/descriptionsencode label semanticslabel anchorsalign input with relation meaningPDF tr. 7
Contrastive alignmentinput/label representationspull positives, push negativesaligned embedding spacereduce semantic confusionPDF tr. 7
Cascade votingcandidate task/relation scoresstaged votingtask/relation decisionimprove task identity inferencePDF tr. 9
Generative replayold latent distributionssample old featuresreplay dataretain old relationsPDF tr. 11

Điều tôi vẫn chưa hiểu

  • Label descriptions được viết thủ công, lấy từ dataset, hay generated?
  • Cascade voting có thử mọi task/pool không, hay dùng shortlist?
  • Generative replay và label alignment tương tác thế nào trong loss tổng?

Phase 5 - Section Recall

Prompt pool + label descriptions

  • Input: sentence/entity pair và relation label descriptions.
  • Process: prompt pool encode input; label encoder tạo anchors; alignment kéo input đúng về label đúng.
  • Output: representation có cả signal context và label semantics.
  • Purpose: giảm nhầm relation gần nghĩa và làm prompt selection có ngữ nghĩa hơn.
  • Still unclear: chất lượng label description ảnh hưởng bao nhiêu.

Cascade voting + replay

  • Input: scores từ nhiều prompt/task candidates và old latent samples.
  • Process: cascade voting chọn task/relation; replay giữ classifier không trôi về task mới.
  • Output: prediction trên relation đã thấy.
  • Purpose: giảm task identity error và catastrophic forgetting.
  • Still unclear: trade-off latency so với WAVE-CRE.

Phase 6 - Equations

Prompt operational equation walkthrough

Walk me through the key equations or formal blocks in this paper.
 
For each equation/block, explain:
1. Input: what variables or objects go into it.
2. Output: what it produces.
3. Where it is used in the training/inference pipeline.
4. What behavior it encourages.
5. What would likely break or become weaker if removed.
6. Which table, figure, ablation, or result supports its usefulness.
 
Do not summarize the whole paper. Focus only on operational understanding of the equations and formal mechanisms.
Eq/blockDùng để làm gì?Biến chínhBehavior được khuyến khíchEvidence / ablationStatus
CRE formalizationđịnh nghĩa stream/task/relationstasks, relations, examplesevaluate trên labels đã thấyPDF tr. 4source-checked
Prompt pool scoringchọn prompt expertsinput query, expert keysinput-specific adaptation; prompt pool → một prompt/task làm FewRel giảm 1.3 và TACRED giảm 1.4PDF tr. 8, PDF tr. 18source-checked
Label alignment lossalign input với label textinput embedding, label-description embeddingsemantic separation; bỏ label descriptions làm FewRel giảm 1.9 và TACRED giảm 1.8PDF tr. 10, PDF tr. 18source-checked
Cascade votingtask/relation decisionMahalanobis scores từ BERT/prompt poolsrobust task inference; task prediction hơn WAVE-CRE 2.9 trên FewRel và 5.6 trên TACREDPDF tr. 12, PDF tr. 19source-checked
Replay objectiveretain old classeslatent old samples từ Gaussian statisticsgiảm forgetting; bỏ generative replay làm FewRel giảm 25.6 và TACRED giảm 22.2PDF tr. 14, PDF tr. 18source-checked

Phase 7 - Loss Functions

WAVE++ training signal
├── current-task relation classification
├── prompt pool routing / adaptive prompting
├── label-description alignment
├── latent generative replay
└── task/relation decision via cascade voting
Loss/blockInputsTrains componentBehaviorWeightAblation
Relation classificationcurrent examplesclassifier/prompt pool hiện tạiclassify new relationssource-checkedmain result
Label-description contrastive/alignmentinput + label descriptionsrepresentation/label anchorsreduce semantic confusionsource-checkedlabel description ablation
Replay lossgenerated old latent samplesclassifier/shared spaceretain old relationssource-checkedreplay ablation, largest drop
Cascade voting objective/scoretask/relation candidatesvoting score, không phải learned MLP task predictorimprove task identitysource-checkedtask prediction analysis

Phase 8 - Experiments

ExperimentResearch questionDatasetBaselinesMetricTable/FigureMain resultCaveat
Main resultsWAVE++ có hơn WAVE-CRE/SOTA không?FewRel/TACREDWAVE-CRE, EoE, rehearsal baselinesfinal-stage accuracymain tablepaper note ghi FewRel T10 87.7, TACRED T10 82.5protocol-specific
Ablationcomponent nào đóng góp?FewRel/TACREDWAVE++ variantsaccuracy dropablation tablebỏ replay giảm 25.6/22.2; bỏ descriptions giảm 1.9/1.8; bỏ prompt pool giảm 1.3/1.4reported/observed, not reproduced
Task predictioncascade voting có tốt hơn WAVE-CRE predictor không?FewRel/TACREDWAVE-CREtask prediction accuracyanalysisFewRel 88.3 vs 85.4; TACRED 84.8 vs 79.2 ở latency tăng
Label description appendixlabel wording/semantics có ảnh hưởng không?FewRel/TACREDdescription variantsaccuracyappendixappendix kiểm sensitivity của label descriptionsdescription quality sensitive
Running timecost tăng bao nhiêu?FewRel/TACREDWAVE-CRE vs WAVE++ms/sample hoặc timeappendixinference latency tăngtrade-off deployment

Protocol fingerprint

  • Dataset and split: FewRel và TACRED, chi tiết appendix. PDF tr. 30
  • Scenario / label space: continual relation extraction, evaluate sau mỗi learning stage.
  • Backbone: BERT frozen; train prompt pools và classifier.
  • Seeds / number of runs: cần kiểm từ experiment section.
  • Metric and averaging: final-stage accuracy và stage-wise accuracy.
  • Replay/memory: latent generative replay.
  • Task identity at inference: cascade voting/prediction, không nên giả định task oracle.
  • Inference cost: appendix có running time. PDF tr. 30

Phase 9 - Claim to Evidence

ClaimWhere claim appearsExperimentEvidenceMy judgmentCaveat
WAVE++ mở rộng WAVE-CRE bằng label descriptions/cascade voting.MethodarchitecturePDF tr. 5supportedcần tách inherited/new components
Main results tốt hơn WAVE-CRE ở stage cuối.Resultsmain tablePDF tr. 16observedchưa reproduced
Generative replay là component mạnh.Ablationcomponent removalPDF tr. 18strong within frameworkkhông chứng minh Gaussian replay là tối ưu
Task prediction cải thiện so với WAVE-CRE.Analysistask predictionPDF tr. 19importantlatency/cost tăng
Appendix kiểm thêm label descriptions/significance/time.Appendixextra analysesPDF tr. 27, PDF tr. 30usefulcần đọc kỹ trước khi cite

Phase 10 - Ablation Study

ComponentIntended purposeWith componentWithout componentDifferenceConclusion justifiedNot justified
Prompt poolcapture within-task varianceFewRel 87.7 / TACRED 82.5một prompt/task: 86.4 / 81.1+1.3 / +1.4adaptive prompting có íchvariance đã được đo trực tiếp
Label descriptionssemantic anchorsFewRel 87.7 / TACRED 82.5bỏ descriptions: 85.8 / 80.7+1.9 / +1.8label semantics hỗ trợ modeldescription nào cũng tốt
Generative replayretain old relationsFewRel 87.7 / TACRED 82.5bỏ replay: 62.1 / 60.3+25.6 / +22.2replay rất quan trọng trong framework nàyrehearsal-free nghĩa là không lưu gì
Cascade votingtask identity inferenceFewRel 88.3 / TACRED 84.8 task predictionWAVE-CRE 85.4 / 79.2+2.9 / +5.6task prediction cải thiệnkhông có inference overhead

Phase 11 - Critical Reading

  • Strongest contribution: biến WAVE-CRE thành system hoàn chỉnh hơn bằng label semantics và cascade voting.
  • Weakest part cần kiểm: nhiều component cùng thêm vào làm attribution khó.
  • Main assumption: label descriptions có chất lượng tốt và task boundary/training stream rõ.
  • Alternative explanation: phần gain lớn có thể đến từ replay/task prediction hơn là label descriptions.
  • Missing experiment: standardized description generation, cross-domain descriptions, prompt utilization.
  • Generalization risk: label descriptions có thể yếu khi relation labels mơ hồ hoặc domain-specific.
  • Reproducibility risk: cần code/hyperparams cho cascade voting và description construction.

Phase 12 - Reproduction Check

ItemTrạng tháiCần làm
PDF localdoneđã có 30 trang
Codetodotìm official repo nếu cần reproduce
Dataset splittodođối chiếu appendix
Label descriptionstodotrích nguồn/format descriptions
Main tablepartialđã ghi FewRel/TACRED và so với WAVE-CRE/EoE
Ablationdone-for-T10đã điền prompt pool, label descriptions, generative replay
Task predictionpartialđã điền task prediction; latency vẫn cần đọc appendix khi reproduce
Statistical testspartialpaper note ghi Royston tests trên tám class của một FewRel task; cần kiểm setup nếu cite sâu

Phase 13 - Completeness / Oral Exam

Mình tự trả lời sau khi đọc

  1. WAVE++ khác WAVE-CRE ở những component nào?
  2. Label descriptions được đưa vào representation/loss thế nào?
  3. Cascade voting giải quyết task identity bằng cách nào?
  4. Replay trong WAVE++ giữ tri thức cũ ở cấp data, feature hay prototype?
  5. Ablation nào chứng minh prompt pool, label descriptions, replay?
  6. Khi nào WAVE++ không nên so sánh trực tiếp với CPL/ConPL?
  7. Chi phí inference tăng ở đâu?

Prompt oral exam

Quiz me on WAVE++. Ask one question at a time. Focus on differences from WAVE-CRE, label descriptions, cascade voting, generative replay, ablations, task prediction, protocol fingerprint, and latency caveats. Do not give the answer until I respond.

Final Paper Note Handoff

Ý cần chuyển sang paper note

  • Bổ sung exact equations/losses nếu cần.
  • Điền full ablation numbers chính ở từ paper note/PDF.
  • Ghi rõ component nào inherited từ WAVE-CRE, component nào mới.
  • Thêm caveat về inference latency và task prediction.

Liên kết