Pretraining progress is mostly coming from data
Breaking down 6 years of pretraining progress into data vs model improvements

TL;DR
- From 2019 to 2025, AI pretraining saw 3.24x more compute efficiency gains from data improvements compared to model improvements.
- Model improvements were crucial for enabling larger-scale training by overcoming constraints, rather than solely for compute efficiency.
- Gains from data and model improvements are mostly independent, with additive effects explaining 88% of performance variance.
- Data improvements, such as better filtering and larger scrapes, were key, while model progress included architectural and optimizer tweaks.
- The effectiveness of data improvements may decrease for very large models with high capacity.