Tag
Fine-tuning
7 posts tagged Fine-tuning.
LIMA: What the Superficial Alignment Paper Actually Says
A 1,000-example SFT run on LLaMA-65B outperforms models trained with 52,000 examples and beats text-davinci-003 in human evals. LIMA's core claim: your model already knows how to be helpful — alignment is just format learning.
QLoRA: What the Paper Actually Says About Fine-Tuning on Consumer Hardware
Full fine-tuning a 65B model requires ~780 GB of GPU memory. QLoRA gets it to 48 GB — a single A100 — without meaningfully degrading quality. The three mechanisms that make this possible are more interesting than the headline number.
Llama 2: What the Open-Source RLHF Paper Actually Says
By turn 15, your agent is ignoring its system prompt. Llama 2 documented the fix — Ghost Attention — but also the full iterative RLHF pipeline with rejection sampling that outperforms PPO alone. Here's what the paper actually says.
S-LoRA: What the Paper Actually Says About Serving Thousands of LoRA Adapters
LoRA makes fine-tuning cheap. Serving thousands of LoRA adapters from a single GPU cluster is a different problem entirely — one that requires rethinking memory management, batching, and tensor parallelism from scratch.
LoRA: What the Low-Rank Adaptation Paper Actually Says
Full fine-tuning GPT-3 requires roughly 1.4 TB of optimizer state. LoRA gets trainable parameters down to ~4.7M with comparable quality — by exploiting a property of pre-trained models that most engineers know about but don't fully reason from.
Toolformer: What the Paper Actually Says
A 6.7B model beats GPT-3 175B on math by learning to use a calculator. Toolformer's self-supervised training pipeline is the interesting part — and it's more constrained than the demos imply.
T5: What the Text-to-Text Paper Actually Says
Every instruction-tuned model today owes something to T5's core idea: every NLP task is just sequence-to-sequence. But the paper's real contribution is a systematic ablation of what actually helps in transfer learning — and several of the answers are counterintuitive.