Research Log · Supervised Fine-Tuning · Study complete
A ten-round data-scaling study on a 125M from-scratch base model (SLM-125M). Each round added 1,000 teacher-written QnA pairs and re-fine-tuned from scratch. Going from 1,000 → 10,000 pairs left answer quality flat at ~1.5/5 — and tripled catastrophic forgetting.
All the quality gain arrived by round 2. The following 5× increase in data (2,000 → 10,000) moved the judge by +0.04 — noise. Meanwhile every extra round reliably made the model forget more of what it originally knew.
Every point re-trains from the base model on the cumulative set, scored on the same frozen 100-pair eval and 50-question judge set. Method frozen throughout — only dataset size varies.
| Round | Pairs | SFT-eval ppl | Retention ppl | Forgetting | Judge /5 |
|---|---|---|---|---|---|
| 0 (base) | 0 | 24.44 | 11.35 | — | 1.00 |
| 1 | 1,000 | 8.60 | 12.05 | +6.1% | 1.32 |
| 2 | 2,000 | 8.00 | 12.42 | +9.5% | 1.50 ← gain ends here |
| 3 | 3,000 | 7.70 | 12.60 | +11.0% | 1.46 |
| 4 | 4,000 | 7.51 | 12.74 | +12.2% | 1.50 |
| 5 | 5,000 | 7.37 | 12.90 | +13.6% | 1.52 |
| 6 | 6,000 | 7.28 | 12.95 | +14.1% | 1.54 |
| 7 | 7,000 | 7.20 | 13.00 | +14.6% | 1.60 |
| 8 | 8,000 | 7.12 | 13.10 | +15.4% | 1.60 |
| 9 | 9,000 | 7.05 | 13.18 | +16.2% | 1.54 |
| 10 | 10,000 | 7.01 | 13.20 | +16.3% | 1.54 |
Judge across rounds: 1.00 · 1.32 · 1.50 · 1.46 · 1.50 · 1.52 · 1.54 · 1.60 · 1.60 · 1.54 · 1.54 — it wanders between 1.46 and 1.60 and ends where it started. The round-7/8 peak did not hold.
Retention perplexity on the original pretraining validation set — the same metric that produced the base model's 11.35 — rose in every single one of ten rounds.
The sharpest illustration of the whole study — one fixed probe, tracked across rounds:
More data made a previously-correct answer worse. Consistent with accumulating drift: the model is pulled away from its pretrained knowledge (retention +16.3%) faster than extra closed-book examples can teach it anything, so earlier wins get overwritten.
SFT-eval perplexity improved in all ten rounds — 8.60 → 8.00 → 7.70 → 7.51 → 7.37 → 7.28 → 7.20 → 7.12 → 7.05 → 7.01 — a clean monotonic decline, while the judge sat flat and the samples regressed.
Ten rounds of unambiguous evidence that perplexity measures distribution fit, not answer quality. Had we reported it as the headline metric, it would have told the exact opposite story — a tidy reminder to pick the metric that matches the claim.
~40% of v1's training is closed-book QA: recall a document fact from a single exposure. A 125M model can't store long-tail facts that way, so no amount of supervision installs them.
gemini-3.1-flash-lite writes QnA grounded in passages from the same corpus the base was pretrained on (US case law, SEC filings, educational web).