CS224N
Mục tiêu
- Gom các lecture, note phụ trợ, chapter liên quan và paper đọc kèm của CS224N vào một bản đồ học tập.
- Dùng các source note bên dưới làm nơi xử lý tài liệu trước khi rút kiến thức vào concept note.
Lộ trình học sâu
Lecture chính
- CS224N 2026 - Lecture 02 - Word Vectors
- CS224N 2026 - Lecture 03 - Neural Network Foundations
- CS224N 2026 - Lecture 04 - Language Models and Recurrent Neural Networks
- CS224N 2026 - Lecture 05 - Attention and Transformers
- CS224N 2026 - Lecture 06 - Final Projects and Practical Tips
- CS224N 2026 - Lecture 07 - Pretraining
- CS224N 2026 - Lecture 08 - Post-training
- CS224N 2026 - Lecture 09 - Efficient Adaptation
- CS224N 2026 - Lecture 10 - RAG and Language Agents
- CS224N 2026 - Lecture 11 - Evaluation
- CS224N 2026 - Lecture 12 - Reasoning Part 1
- CS224N 2026 - Lecture 13 - Reasoning Part 2
- CS224N 2026 - Lecture 14 - Tokenization and Multilinguality
- CS224N 2026 - Lecture 16 - AIs Impact on Humanity
- CS224N 2026 - Lecture 19 - The Art of Artificial Reasoning for Small Language Models
Course notes
- CS224N - Notes - Backpropagation Old
- CS224N 2017 - Review of Differential Calculus Theory
- CS224N 2019 - Notes - Computing Neural Network Gradients
- CS224N 2019 - Notes 02 - Word Vectors II - GloVe Evaluation and Training
- CS224N 2019 - Notes 03 - Neural Networks and Backpropagation
- CS224N 2019 - Notes 05 - Language Models RNN GRU and LSTM
- CS224N 2023 - Notes 01 - Introduction and Word2Vec - Draft
- CS224N 2023 - Notes 10 - Self-Attention and Transformers - Draft
Chapter đọc kèm
- SLP 2026 - Chapter 02 - Words and Tokens
- SLP 2026 - Chapter 09 - Post-training - Instruction Tuning Alignment and Test-Time
- SLP 2026 - Chapter 10 - Masked Language Models
Papers đọc kèm
- 2011 - Natural Language Processing Almost from Scratch - JMLR
- 2013 - Distributed Representations of Words and Phrases and their Compositionality - NeurIPS
- 2013 - Efficient Estimation of Word Representations in Vector Space - arXiv 1301.3781v3
- 2013 - On the Difficulty of Training Recurrent Neural Networks - arXiv 1211.5063v2
- 2014 - GloVe - Global Vectors for Word Representation
- 2015 - Improving Distributional Similarity with Lessons Learned from Word Embeddings - TACL Q15-1016
- 2016 - A Latent Variable Model Approach to PMI-based Word Embeddings - TACL Q16-1028
- 2016 - Layer Normalization - arXiv 1607.06450v1
- 2016 - Neural Machine Translation of Rare Words with Subword Units - arXiv 1508.07909v5
- 2017 - Attention Is All You Need - arXiv 1706.03762v7
- 2017 - Derivatives Backpropagation and Vectorization - Justin Johnson
- 2018 - BERT - Pre-training of Deep Bidirectional Transformers for Language Understanding - arXiv 1810.04805v2
- 2018 - Image Transformer - arXiv 1802.05751v3
- 2018 - Music Transformer - Generating Music with Long-Term Structure - arXiv 1809.04281v3
- 2018 - On the Dimensionality of Word Embedding - NeurIPS
- 2019 - Parameter-Efficient Transfer Learning for NLP - arXiv 1902.00751v2
- 2019 - The Lottery Ticket Hypothesis - Finding Sparse Trainable Neural Networks - arXiv 1803.03635v5
- 2020 - Contextual Word Representations - A Contextual Introduction - arXiv 1902.06006v3
- 2020 - Language Models are Few-Shot Learners - arXiv 2005.14165v4
- 2020 - Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks - arXiv 2005.11401v4
- 2020 - Unsupervised Cross-lingual Representation Learning at Scale - ACL Main 747
- 2021 - LoRA - Low-Rank Adaptation of Large Language Models - arXiv 2106.09685v2
- 2021 - Measuring Massive Multitask Language Understanding - arXiv 2009.03300v3
- 2021 - RoFormer - Enhanced Transformer with Rotary Position Embedding - arXiv 2104.09864v5
- 2022 - Chain-of-Thought Prompting Elicits Reasoning in Large Language Models - arXiv 2201.11903v6
- 2022 - Fast Inference from Transformers via Speculative Decoding - arXiv 2211.17192v2
- 2022 - Retrieval-Augmented Multimodal Language Modeling - arXiv 2211.12561v2
- 2022 - Scaling Instruction-Finetuned Language Models - arXiv 2210.11416v5
- 2023 - AlpacaFarm - A Simulation Framework for Methods that Learn from Human Feedback - arXiv 2305.14387v4
- 2023 - Direct Preference Optimization - Your Language Model is Secretly a Reward Model - arXiv 2305.18290v3
- 2023 - Do All Languages Cost the Same - Tokenization in the Era of Commercial Language Models - EMNLP Main 614
- 2023 - Holistic Evaluation of Language Models - TMLR 2023
- 2023 - How Far Can Camels Go - Exploring Instruction Tuning on Open Resources - arXiv 2306.04751v2
- 2023 - Lets Verify Step by Step - arXiv 2305.20050v1
- 2023 - ReAct - Synergizing Reasoning and Acting in Language Models - arXiv 2210.03629v3
- 2023 - Scaling Autoregressive Multi-Modal Models - Pretraining and Instruction Tuning - arXiv 2309.02591v1
- 2023 - Scaling Laws for Generative Mixed-Modal Language Models - arXiv 2301.03728v1
- 2023 - Self-Consistency Improves Chain of Thought Reasoning in Language Models - ICLR 2023
- 2023 - Toolformer - Language Models Can Teach Themselves to Use Tools - arXiv 2302.04761v1
- 2024 - Chameleon - Mixed-Modal Early-Fusion Foundation Models - arXiv 2405.09818v2
- 2024 - LMFusion - Adapting Pretrained Language Models for Multimodal Generation - arXiv 2412.15188v4
- 2024 - Scaling LLM Test-Time Compute Optimally Can Be More Effective Than Scaling Model Parameters - arXiv 2408.03314v1
- 2024 - The Llama 3 Herd of Models - arXiv 2407.21783v3
- 2024 - Transfusion - Predict the Next Token and Diffuse Images with One Multi-Modal Model - arXiv 2408.11039v1
- 2024 - Visual Sketchpad - Sketching as a Visual Chain of Thought for Multimodal Language Models - arXiv 2406.09403v3
- 2025 - Agentic Interpretability - Because We Have LLMs We Can and Should Pursue It - arXiv 2506.12152v1
- 2025 - Bridging the Human-AI Knowledge Gap Through Concept Discovery and Transfer in AlphaZero - PNAS
- 2025 - DAPO - An Open-Source LLM Reinforcement Learning System at Scale - arXiv 2503.14476v2
- 2025 - DeepSeek-R1 - Incentivizing Reasoning Capability in LLMs via Reinforcement Learning - arXiv 2501.12948v2
- 2025 - Mixture-of-Transformers - A Sparse and Scalable Architecture for Multi-Modal Foundation Models - TMLR 2025
- 2025 - Multimodal RewardBench - Holistic Evaluation of Reward Models for Vision Language Models - arXiv 2502.14191v1
- 2025 - Neologism Learning for Controllability and Self-Verbalization - arXiv 2510.08506v1
- 2025 - OneFlow - Concurrent Mixed-Modal and Interleaved Generation with Edit Flows - arXiv 2510.03506v3
- 2025 - Reconstruction Alignment Improves Unified Multimodal Models - arXiv 2509.07295v3
- 2025 - We Cant Understand AI Using Our Existing Vocabulary - arXiv 2502.07586v1
Concept trung tâm
- Tokenization
- BPE
- Embedding
- Word2Vec
- GloVe
- Neural NLP
- GRU
- LSTM
- Self-Attention
- Multi-Head Attention
- Transformer
- Masked Language Modeling
- Rotary Positional Embeddings
- Large Language Model
- Instruction Fine-Tuning
- Retrieval-Augmented Generation
- LLM Agent
- Tool Use
- Speculative Decoding
- Chain-of-Thought Prompting
- Self-Consistency Decoding
- Test-Time Compute
- RLHF
- DPO
- Reward Model
- Parameter-Efficient Fine-Tuning
- Adapter
- LoRA
- QLoRA
- Multimodal LLM
- AI Hallucination