NLP Transformers - Chapter 09 - Dealing with Few to No Labels
Mục tiêu đọc
- Hiểu các chiến lược khi thiếu nhãn dữ liệu.
- Biết dùng baseline đơn giản trước khi fine-tune Transformer.
- Nắm data augmentation, embeddings lookup, prompts và tận dụng unlabeled data.
Ý chính
- Khi ít nhãn, baseline đơn giản giúp biết Transformer có thật sự cần thiết không.
- Embeddings có thể dùng như lookup table để tìm ví dụ gần nhau.
- Unlabeled data vẫn có giá trị thông qua language model fine-tuning hoặc semi-supervised methods.
Demo thực hành
Zero-shot classification khi chưa có dữ liệu gán nhãn.
from transformers import pipeline
classifier = pipeline("zero-shot-classification", model="facebook/bart-large-mnli")
texts = [
"The app crashes whenever I try to upload a file.",
"Please add dark mode to the dashboard.",
]
labels = ["bug", "feature request", "question", "documentation"]
for text in texts:
result = classifier(text, candidate_labels=labels)
print(text)
print(list(zip(result["labels"], result["scores"])))Khái niệm quan trọng
Active Recall
- Khi ít nhãn, vì sao nên làm baseline trước?
- Zero-shot classification dựa trên giả định gì?
- Data augmentation có thể làm hỏng dữ liệu ra sao?
- Unlabeled data giúp được gì cho classifier?
Checklist
- Đọc xong chapter
- Chạy demo zero-shot classification
- Nghĩ một use case cá nhân có ít nhãn
- Tách concept cần dùng lại
- Cập nhật tiến độ sách