← All writing
Paper Breakdown

LIMA: what the superficial alignment paper actually says

Reading Zhou et al. (Meta AI, NeurIPS 2023) after watching a team spend three months and $200K on RLHF infrastructure for a model that still answered questions badly.

The RLHF pipeline had everything: 50,000 preference pairs collected from contractors, a reward model trained on them, PPO with a KL penalty tuned across six hyperparameter sweeps. The resulting model was technically "aligned" — it wouldn't produce slurs, it had preambles on factual questions, it refused requests it shouldn't have. But ask it to explain a database join to a new engineer, and it gave you three paragraphs of hedging and a mediocre explanation. Ask the unaligned SFT model, and it often did better.

The team concluded they needed more preference data. Someone bought more annotation hours. The model got marginally more polite and marginally less useful.

LIMA — "LIMA: Less Is More for Alignment", Chunting Zhou et al., Meta AI Research, NeurIPS 2023 — is the paper that explains what was happening. The core finding: 1,000 carefully curated supervised fine-tuning examples produced a model preferred by human raters over text-davinci-003 (53% win rate) and Bard (65% win rate). No reinforcement learning. No reward model. No PPO. Just one thousand examples, chosen carefully.

This isn't a paper about saving money. It's a paper about what alignment is actually doing.

The Superficial Alignment Hypothesis

The paper's central claim is stated plainly:

"A model's knowledge and capabilities are learnt almost entirely during pretraining, while alignment teaches it which subdistribution of formats should be used when interacting with users."

This is the Superficial Alignment Hypothesis. If it's right, the implications are sharp. Alignment fine-tuning doesn't teach the model new knowledge. It doesn't improve reasoning or expand capabilities. It teaches the model how to present its existing knowledge in a format appropriate for user interaction — the difference between a model that says "The capital of France is Paris" and one that says "The capital of France is Paris. Is there anything else about European geography you'd like to know?"

The hypothesis predicts three things:

  • A small number of diverse, high-quality examples should be sufficient for alignment
  • Adding more low-quality examples should not improve and might degrade performance
  • The base model's pretraining quality sets the ceiling on what alignment can achieve

The paper tests all three empirically, and the results are cleaner than you'd expect.

What 1,000 examples actually looks like

The LIMA dataset is 1,000 prompt-response pairs. This is the entire training set. No augmentation, no curriculum, no staged training.

The composition:

  • Stack Exchange (200 examples): Questions with score ≥ 20 from the top 75 domains, filtered to topics that generalize. Answers were rewritten to remove Stack Exchange formatting artifacts.
  • wikiHow (200 examples): High-quality "how to" articles converted into instruction-response pairs.
  • Pushshift Reddit (150 examples): Top-voted posts from subreddits known for substantive, helpful responses — r/askscience, r/changemyview, r/explainlikeimfive, and similar.
  • Super-Natural Instructions (200 examples): Diverse NLP tasks to prevent mode collapse on conversational formats.
  • Manually written (250 examples): Meta team members wrote these themselves to fill gaps — math problems, coding, multi-turn conversation, topics the other sources didn't cover well.

The curation criteria matter more than the sources. The team filtered aggressively on three dimensions:

Quality: Every example must demonstrate what a good answer actually looks like. If the response would pass review from a knowledgeable editor, it goes in. If not, it doesn't, regardless of upvotes.

Diversity: The dataset needs to cover the space of tasks and topics the model will encounter at inference time. The paper explicitly tracks topic and format diversity and discards duplicates.

Style consistency: All responses are written or rewritten in a consistent voice. Informal questions get direct answers. Technical questions get structured responses. Requests for lists get lists. The model is learning a format policy, and inconsistent format demonstrations confuse that learning.

The dataset took several weeks of careful human curation. That's not coincidental — the paper's implicit argument is that curation time substitutes for data volume.

The quantity ablation: where the plateau is

To test whether 1,000 examples is actually the right number, the paper runs a controlled ablation. They train the same LLaMA-65B base model on 32, 128, 256, 512, and 1,000 examples, then extend to 2,000 examples by adding lower-quality examples to the original curated set, and finally compare against Alpaca's 52,000 GPT-3-generated examples.

The performance curve is instructive. The model improves meaningfully from 32 to 512 examples. From 512 to 1,000, improvement continues but slows. At 2,000 examples — the 1,000 curated plus 1,000 lower-quality additions — performance drops relative to the 1,000-example model.

The Alpaca comparison is the headline. Alpaca uses 52,000 examples generated by GPT-3 self-instruct. The LIMA-trained model beats the Alpaca-trained model in human preference evals with a 72% win rate. More data at lower quality is strictly worse than less data at higher quality, by a substantial margin.

This is exactly what the Superficial Alignment Hypothesis predicts: once the model has seen enough diverse, high-quality examples to learn the format distribution, more examples don't help. You're teaching a format policy, not capabilities, and a format policy can be learned from a modest number of examples.

The quality ablation: what actually drives degradation

The paper tests quality directly by taking the 1,000-example dataset and mixing in batches of lower-quality examples scraped from the same sources without curation. Adding 32 low-quality examples has minimal effect. Adding 256 starts to degrade outputs. Adding 2,632 produces measurably worse human preference ratings.

The mechanism appears to be format pollution. Lower-quality examples have inconsistent response styles — hedging where confidence is warranted, padding and verbosity where directness is better, incomplete answers that stop at a summary instead of explaining. The model learns this distribution and applies it at inference time.

The implication for practitioners: if your alignment fine-tuning made your model worse than the base SFT model, the most likely cause is data quality, not quantity. The instinct to add more data is wrong. You want to remove data.

The human evaluation results

LIMA's published results come from human preference comparisons where annotators saw two model outputs side-by-side and chose which they preferred. LIMA (LLaMA-65B, 1K examples) against various baselines:

| Opponent | LIMA wins | Tie | LIMA loses | |----------|-----------|-----|------------| | Alpaca (LLaMA-65B, 52K examples) | 72% | — | 28% | | text-davinci-003 | 53% | 17% | 30% | | Bard | 65% | — | 35% | | Claude | 38% | 5% | 57% | | GPT-4 | 13% | 5% | 82% |

A few things to read carefully:

First, the gap against Claude and GPT-4 is substantial. LIMA is not competing with frontier models — it's competitive with first-generation RLHF models. That's the actual claim, and it's still a meaningful one.

Second, the LLaMA-65B base model is doing the heavy lifting. LIMA on LLaMA-7B produces noticeably weaker results. The Superficial Alignment Hypothesis implies that alignment performance is bounded by pretraining quality, and that bound is visible at 7B.

Third, the test set skews toward instruction following, explanation, and coding. LIMA does not perform well on tasks requiring current information, nuanced refusals, or multi-turn correction. Those aren't in the paper's headline numbers.

The safety experiment: a production warning

The paper includes 30 examples of harmful queries paired with appropriate refusal responses. This gives the model basic safety behavior — it refuses clearly harmful requests at roughly the rate of first-generation RLHF models on standard eval sets.

Then the paper tests what happens when you add 30 different examples that don't reinforce safety. The safety behavior degrades.

This is the most important result for production decisions. LIMA's safety properties are learned from a vanishingly small fraction of the training distribution. They're fragile by construction. Constitutional AI and the full RLHF pipeline exist partly because safety needs to be robust under adversarial prompting. Thirty curated safety examples won't achieve that, and the paper doesn't claim otherwise — but teams reading only the headline results sometimes miss this.

The multi-turn problem

LIMA's evaluation is almost entirely single-turn. The paper includes a small multi-turn analysis using 10 dialogues, and the results are mixed. LIMA maintains coherence across turns but struggles with correction. If a user says "that's wrong, try again," the model often produces a rephrasing of the original answer rather than a substantively different one.

Multi-turn repair, clarification, and iterative refinement are skills that need examples in the training data. LIMA's 1,000 examples don't include enough of them. You'd need to add those deliberately — which is exactly what the paper predicts: format policies need examples of the formats you want.

When not to use this approach

Safety-critical consumer products. LIMA's safety is fragile by design. If you need robust refusals across adversarial prompting, you need Constitutional AI, RLHF with safety-specific preference data, or both. Thirty curated safety examples will not protect you under adversarial load.

Tasks requiring preference learning. LIMA teaches format, not preferences. If you need the model to prefer concise over verbose, to ask clarifying questions before attempting complex tasks, or to prioritize certain values in edge cases — that's preference learning. You need preference data and a mechanism to optimize it. DPO costs much less than PPO and is worth considering here before 1000-example SFT.

Domain knowledge not in pretraining. If you're building a model for a specialized domain your base LLM never saw — proprietary codebases, internal processes, post-cutoff events — format learning won't help. The knowledge isn't there to unlock. You need continued pretraining, retrieval augmentation, or both.

Models smaller than ~30B parameters. The Superficial Alignment Hypothesis holds most cleanly for large base models. The LIMA results with LLaMA-7B are noticeably weaker. If your deployment target is a small model, the 1,000-example approach may not transfer.

Consistent behavior across a long tail of inputs. LIMA is optimized for human preference evals, which are good at catching obviously bad responses but miss subtle inconsistency. If your deployment requires high consistency across thousands of diverse prompts — not just good performance on 252 test cases — you'll want more coverage in your training data and more systematic evaluation.

What this means for fine-tuning pipelines

The practical implication isn't "replace your RLHF pipeline with 1,000 examples." It's "audit your data before you build infrastructure."

Most teams building instruction fine-tuning pipelines start with data volume as the primary variable. Collect more prompts, collect more responses, generate more with self-instruct. LIMA's evidence suggests the leverage is in the opposite direction: take your existing data, discard the bottom 80% by quality, ensure the remaining examples are diverse, rewrite responses to be consistently high-quality, and measure whether that closes your gap before investing in preference data collection and RL training.

For many business-specific fine-tuning tasks — turning a general model into a domain assistant for internal tooling, customer support, code review — this approach is likely sufficient. The base LLM already understands the domain from pretraining. You're teaching it which format to apply. A thousand high-quality examples from domain experts will outperform ten thousand examples from contractors unfamiliar with the domain, for exactly the reason LIMA describes.

The prerequisite is honesty about what alignment is doing in your specific case. If you're unlocking knowledge and capabilities that exist in pretraining, 1,000 good examples is probably enough. If you're trying to teach the model something it never learned — whether that's a specialized domain, robust safety properties, or nuanced preference patterns — the Superficial Alignment Hypothesis doesn't apply, and neither does the shortcut.

The deeper lesson is about what fine-tuning actually is. It's not capability transfer. It's style transfer applied to an existing capability distribution. That's a narrower thing than most fine-tuning pipelines assume, and treating it as such — small, curated, specific — tends to produce better results than treating it as another pretraining run with smaller compute.