Self-paced course
Fine-Tuning & Adaptation
Adapting pre-trained models to your task without burning a cluster — LoRA and its variants, quantized training, and the tradeoffs that decide quality.
0 / 12 lessons complete
Curriculum
- The Llama 3 Herd of Models: What the Paper Actually Says
- Switch Transformers: What the Sparse MoE Scaling Paper Actually Says
- T5: What the Text-to-Text Paper Actually Says
- DeepSeek-V3: What the Frontier-on-a-Budget Paper Actually Says
- Toolformer: What the Paper Actually Says
- Ring Attention: What the Near-Infinite Context Paper Actually Says
- MegaScale: What ByteDance's 12,288-GPU Training Paper Actually Says
- LoRA: What the Low-Rank Adaptation Paper Actually Says
- S-LoRA: What the Paper Actually Says About Serving Thousands of LoRA Adapters
- Llama 2: What the Open-Source RLHF Paper Actually Says
- BitNet b1.58: What the 1-bit LLM Paper Actually Says
- QLoRA: What the Paper Actually Says About Fine-Tuning on Consumer Hardware
Prefer it in your inbox?
Get this track as a free daily email course — one short lesson a day.