Pretraining progress is mostly coming from data

Breaking down 6 years of pretraining progress into data vs model improvements

Pretraining progress is mostly coming from data

TL;DR

  • From 2019 to 2025, AI pretraining saw 3.24x more compute efficiency gains from data improvements compared to model improvements.
  • Model improvements were crucial for enabling larger-scale training by overcoming constraints, rather than solely for compute efficiency.
  • Gains from data and model improvements are mostly independent, with additive effects explaining 88% of performance variance.
  • Data improvements, such as better filtering and larger scrapes, were key, while model progress included architectural and optimizer tweaks.
  • The effectiveness of data improvements may decrease for very large models with high capacity.