← All tracks

Self-paced course

Fine-Tuning & Adaptation

Adapting pre-trained models to your task without burning a cluster — LoRA and its variants, quantized training, and the tradeoffs that decide quality.

8 lessons~2h totalFree
0 / 8 lessons complete

Curriculum

  1. The Llama 3 Herd of Models: What the Paper Actually Says11 min read
  2. Switch Transformers: What the Sparse MoE Scaling Paper Actually Says13 min read
  3. T5: What the Text-to-Text Paper Actually Says13 min read
  4. DeepSeek-V3: What the Frontier-on-a-Budget Paper Actually Says13 min read
  5. Toolformer: What the Paper Actually Says11 min read
  6. Ring Attention: What the Near-Infinite Context Paper Actually Says15 min read
  7. MegaScale: What ByteDance's 12,288-GPU Training Paper Actually Says14 min read
  8. LoRA: What the Low-Rank Adaptation Paper Actually Says10 min read