AI training efficiency: From Throughput to Goodput

Pretraining a modern large language model (LLM), often with ~100B parameters or more, typically involves thousands of accelerators and massive token corpora, running for days to months. At that scale, success is commonly reduced to two headline outcomes:

AI training efficiency: From Throughput to Goodput

TL;DR

  • Pretraining large language models (LLMs) involves thousands of accelerators.
  • Massive token corpora are used in LLM pretraining.
  • The training process can span from days to months.
  • Success is typically measured by two main outcomes.