Chinchilla finds most large language models are undertrained
DeepMind posted “Training Compute-Optimal Large Language Models” on 29 March 2022. Its 70-billion-parameter Chinchilla, trained on four times the data used for the 280-billion-parameter Gopher at equal compute cost, outperformed Gopher and GPT-3 across a broad set of benchmarks.
Why it mattered Labs had been spending compute on size. The paper argued that data volume had been under-weighted at the same cost, and it changed how later models were sized and trained.