tech

Today's Vergecast: How to train your data.

Training data is the raw material of the AI industry. Claude, ChatGPT, Gemini, and the rest are built on top of oceans of stuff. What is that stuff? Books. Blog posts. YouTube videos. News articles. All of it and more, in virtually incomprehensible quantities. Alex Reisner, a staff writer at The Atlantic who has been investigating training data, explains how AI companies get all this data, why they’d really prefer you not know what’s in it, and whether training data could ever be a fair trade.

Today's Vergecast: How to train your data.

TL;DR

  • Training data is fundamental to the AI industry.
  • Sources of training data include books, blog posts, YouTube videos, and news articles.
  • Alex Reisner is investigating how AI companies obtain this data.
  • AI companies may prefer that the contents of their training data remain unknown.