A Pangram Labs study of "wild" AI-generated web text reveals that adding such tokens to pretraining data initially lowers loss on human text but quickly reverses into harm as the ratio increases. The researchers pretrained 800 language models to derive a new scaling law that accounts for this dynamic, which existing laws like Chinchilla fail to predict.

  • Analysis of FineWeb data shows AI-generated tokens rose from 27.5% in June 2026 to 31.1% by August.
  • For data-starved models, added AI tokens lower human text loss initially but saturate and then increase it.
  • Models trained on high budgets of human text see their loss rise almost immediately upon adding AI tokens.
  • The new scaling law predicts held-out human-text loss for models up to 3.6x larger with 41% lower error than existing laws.

The authors recommend filtering AI text when the target is human text, repeating human text before expanding datasets with AI-generated content, and reporting validation loss on human and AI text separately.