Reading List
2017arXiv ↗
Attention Is All You Need
Vaswani et al. · The Transformer paper. Know every component cold.
2020arXiv ↗
Scaling Laws for Neural Language Models
Kaplan et al. (OpenAI) · Defines how model size, data, compute relate to loss.
2022arXiv ↗
Training Compute-Optimal LLMs (Chinchilla)
Hoffmann et al. (DeepMind) · Corrects scaling law - 20 tokens per parameter optimal.
2017arXiv ↗
Deep Residual Learning for Image Recognition
He et al. (MSRA) · ResNets. Why skip connections enable very deep networks.
2012arXiv ↗
ImageNet Classification with Deep CNNs (AlexNet)
Krizhevsky et al. · Started deep learning revolution. Know the architecture.
2015arXiv ↗
Unreasonable Effectiveness of RNNs
Karpathy · Intuition building. Read before understanding LSTMs.
2015arXiv ↗
Understanding LSTM Networks
Olah (colah) · Best visual explanation of LSTMs. Required reading.
2021arXiv ↗
An Image is Worth 16x16 Words (ViT)
Dosovitskiy et al. · Transformers for vision - architecture and pretraining insights.
2022arXiv ↗
FlashAttention: Fast & Memory-Efficient Exact Attention
Dao et al. · IO-aware attention. How to actually implement Transformers efficiently.
2021arXiv ↗
LoRA: Low-Rank Adaptation of LLMs
Hu et al. · Parameter-efficient fine-tuning. Ubiquitous in practice.
2022arXiv ↗
Training LLMs to Follow Instructions (InstructGPT)
Ouyang et al. (OpenAI) · Original RLHF paper. How ChatGPT was made.
2023arXiv ↗
Direct Preference Optimization (DPO)
Rafailov et al. · Simpler RLHF alternative. Used in most modern models.
2022arXiv ↗
Constitutional AI
Bai et al. (Anthropic) · How Claude was trained. Key for Anthropic interviews.
2021arXiv ↗
CLIP: Connecting Text and Images
Radford et al. (OpenAI) · Contrastive multimodal learning foundation. Zero-shot vision.
2022arXiv ↗
Whisper: Robust Speech Recognition
Radford et al. (OpenAI) · Encoder-decoder Transformer for speech. Audio chunk processing.
2023arXiv ↗
Segment Anything Model (SAM)
Kirillov et al. (Meta) · Promptable segmentation. Key CV model of 2023.
2020arXiv ↗
Denoising Diffusion Probabilistic Models (DDPM)
Ho et al. · Diffusion model foundations. Forward/reverse process.
2020arXiv ↗
Language Models are Few-Shot Learners (GPT-3)
Brown et al. (OpenAI) · In-context learning, scaling. The paper that changed everything.
2023arXiv ↗
LLaMA 2: Open Foundation and Fine-Tuned Chat Models
Touvron et al. (Meta) · Open source LLM architecture. GQA, RLHF at scale.
2019arXiv ↗
GPipe: Efficient Training using Pipeline Parallelism
Huang et al. (Google) · Ilya's reading list. Pipeline parallelism foundations.
2014arXiv ↗
Neural Turing Machines
Graves et al. · Ilya's reading list. Memory + attention before Transformers.
2013arXiv ↗
Auto-Encoding Variational Bayes (VAE)
Kingma & Welling · VAE foundations. ELBO, reparameterization trick.
2023arXiv ↗
QLoRA: Efficient Finetuning of Quantized LLMs
Dettmers et al. · Fine-tune 65B model on single GPU. Critical for startup ML.
2024arXiv ↗
DeepSeek-V2
DeepSeek AI · MLA attention, fine-grained MoE, GRPO training. Modern architecture.