A Pangram Labs study of 800 pre-trained language models reveals that "wild" AI-generated web text, which now comprises over 31% of recent web data, negatively impacts model performance when the goal is to process human text. The research demonstrates that while small amounts of AI tokens may initially help data-starved models, they quickly reverse into harm as their proportion increases.

  • For models trained on high budgets of human text, adding AI tokens raises loss almost immediately, whereas fresh human tokens continue to lower it.
  • Existing scaling laws fail to predict this behavior; the authors propose a new law with separate benefit and harm terms that reduces prediction error by 41% compared to existing methods.
  • The study recommends filtering AI text for human-targeted models, repeating human text before adding AI data, and reporting validation loss on human and AI text separately.

The findings suggest that AI text remains valuable only when the target is AI text, but should be filtered or handled carefully when training for human-centric tasks. The authors release WildAI, an 83B-token corpus with labels, along with code and model weights.