← All tracks

Self-paced course

LLM Inference & Serving

How large language models actually run in production — KV-cache management, batching, speculative decoding, quantization, and the systems work that turns a model into a serveable endpoint.

18 lessons~3h totalFree
0 / 18 lessons complete

Curriculum

  1. EAGLE: Speculative Decoding with Feature-Level Prediction — What the Paper Actually Says12 min read
  2. LLM.int8(): What the 8-bit Matrix Multiplication Paper Actually Says10 min read
  3. Mooncake: What the KV-Cache-Centric Disaggregated Serving Paper Actually Says14 min read
  4. Titans: What the Test-Time Memorization Paper Actually Says9 min read
  5. SGLang and RadixAttention: What the Paper Actually Says11 min read
  6. SARATHI: What the Chunked-Prefill Paper Actually Says11 min read
  7. Mixture of Depths: What the Paper Actually Says14 min read
  8. Mamba: What the Selective State Space Paper Actually Says11 min read
  9. Mixture of Experts: What the Architecture Actually Does to Your Inference Budget9 min read
  10. H2O: Heavy-Hitter Oracle for KV Cache Eviction — What the Paper Actually Says14 min read
  11. Ring Attention: What the Near-Infinite Context Paper Actually Says15 min read
  12. DistServe: Disaggregating Prefill and Decoding for Goodput-Optimized LLM Serving13 min read
  13. StreamingLLM: What the Attention Sink Paper Actually Says8 min read
  14. S-LoRA: What the Paper Actually Says About Serving Thousands of LoRA Adapters11 min read
  15. Multi-Token Prediction: What the Meta FAIR Paper Actually Says9 min read
  16. BitNet b1.58: What the 1-bit LLM Paper Actually Says10 min read
  17. MLA: What the Multi-Head Latent Attention Paper Actually Says11 min read
  18. Mistral 7B: What the Sliding Window Attention Paper Actually Says9 min read