Self-paced course
Paper Breakdowns
The “What the Paper Actually Says” series — close readings of the systems and ML papers that matter, focused on the mechanism, the numbers, and when not to use the idea.
0 / 41 lessons complete
Curriculum
- Cassandra: What the Paper Actually Says
- EAGLE: Speculative Decoding with Feature-Level Prediction — What the Paper Actually Says
- The Llama 3 Herd of Models: What the Paper Actually Says
- LLM.int8(): What the 8-bit Matrix Multiplication Paper Actually Says
- Mooncake: What the KV-Cache-Centric Disaggregated Serving Paper Actually Says
- Pregel: What the Large-Scale Graph Processing Paper Actually Says
- Switch Transformers: What the Sparse MoE Scaling Paper Actually Says
- Titans: What the Test-Time Memorization Paper Actually Says
- T5: What the Text-to-Text Paper Actually Says
- SGLang and RadixAttention: What the Paper Actually Says
- SARATHI: What the Chunked-Prefill Paper Actually Says
- Mixture of Depths: What the Paper Actually Says
- MapReduce: What the Google Paper Actually Says
- Mamba: What the Selective State Space Paper Actually Says
- Kafka: What the Original Paper Actually Says
- H2O: Heavy-Hitter Oracle for KV Cache Eviction — What the Paper Actually Says
- DeepSeek-V3: What the Frontier-on-a-Budget Paper Actually Says
- Dapper: What Google's Distributed Tracing Paper Actually Says
- Toolformer: What the Paper Actually Says
- MegaScale: What ByteDance's 12,288-GPU Training Paper Actually Says
- LoRA: What the Low-Rank Adaptation Paper Actually Says
- HNSW: What the Vector Similarity Search Paper Actually Says
- GFS: What the Google File System Paper Actually Says
- Generative Agents: What the Paper Actually Says
- FrugalGPT: What the Paper Actually Says
- DistServe: Disaggregating Prefill and Decoding for Goodput-Optimized LLM Serving
- Bigtable: What the Distributed Storage Paper Actually Says
- Zanzibar: What Google's Authorization Paper Actually Says
- StreamingLLM: What the Attention Sink Paper Actually Says
- Spanner: What Google's Globally-Distributed Database Paper Actually Says
- Raft: What the Understandable Consensus Algorithm Paper Actually Says
- Multi-Token Prediction: What the Meta FAIR Paper Actually Says
- MapReduce: What the Paper Actually Says
- Llama 2: What the Open-Source RLHF Paper Actually Says
- BitNet b1.58: What the 1-bit LLM Paper Actually Says
- Tree of Thoughts: What the Paper Actually Says About LLM Search
- Self-RAG: What Adaptive Retrieval Actually Means in Production
- ReAct: What the Reasoning + Acting Paper Actually Says
- QLoRA: What the Paper Actually Says About Fine-Tuning on Consumer Hardware
- MLA: What the Multi-Head Latent Attention Paper Actually Says
- Mistral 7B: What the Sliding Window Attention Paper Actually Says
Prefer it in your inbox?
Get this track as a free daily email course — one short lesson a day.