← All writing

Tag

Fine-tuning

7 posts tagged Fine-tuning.

LIMA: What the Superficial Alignment Paper Actually Says

A 1,000-example SFT run on LLaMA-65B outperforms models trained with 52,000 examples and beats text-davinci-003 in human evals. LIMA's core claim: your model already knows how to be helpful — alignment is just format learning.

QLoRA: What the Paper Actually Says About Fine-Tuning on Consumer Hardware

Full fine-tuning a 65B model requires ~780 GB of GPU memory. QLoRA gets it to 48 GB — a single A100 — without meaningfully degrading quality. The three mechanisms that make this possible are more interesting than the headline number.

Llama 2: What the Open-Source RLHF Paper Actually Says

By turn 15, your agent is ignoring its system prompt. Llama 2 documented the fix — Ghost Attention — but also the full iterative RLHF pipeline with rejection sampling that outperforms PPO alone. Here's what the paper actually says.

S-LoRA: What the Paper Actually Says About Serving Thousands of LoRA Adapters

LoRA makes fine-tuning cheap. Serving thousands of LoRA adapters from a single GPU cluster is a different problem entirely — one that requires rethinking memory management, batching, and tensor parallelism from scratch.

LoRA: What the Low-Rank Adaptation Paper Actually Says

Full fine-tuning GPT-3 requires roughly 1.4 TB of optimizer state. LoRA gets trainable parameters down to ~4.7M with comparable quality — by exploiting a property of pre-trained models that most engineers know about but don't fully reason from.

Toolformer: What the Paper Actually Says

A 6.7B model beats GPT-3 175B on math by learning to use a calculator. Toolformer's self-supervised training pipeline is the interesting part — and it's more constrained than the demos imply.

T5: What the Text-to-Text Paper Actually Says

Every instruction-tuned model today owes something to T5's core idea: every NLP task is just sequence-to-sequence. But the paper's real contribution is a systematic ablation of what actually helps in transfer learning — and several of the answers are counterintuitive.