A significant portion of the text available on the internet is now generated by artificial intelligence, a trend that is increasing the cost of training language models. Researchers found that by August 2026, AI-generated tokens accounted for 31.1% of a filtered web crawl, up from 27.5% in June of the same year. This "wild" AI text, created by various models and intended for human consumption, is now a substantial component of the data used to train artificial intelligence.
The study, conducted by researchers from Pangram Labs and the University of Maryland, analyzed web data using the FineWeb quality filtering method. They applied the Pangram AI detector to web crawls, observing the increasing prevalence of AI-generated content. To understand the impact of this synthetic text on language model pretraining, the researchers trained 800 language models. These models varied in size and were trained with different ratios of AI-generated to human-written text.
The findings indicate that while adding AI-generated tokens can initially lower the loss on human text for smaller models, this benefit plateaus and can even reverse under certain conditions. Crucially, the researchers determined that achieving the same level of performance as a model trained solely on human text, at a standard of 20 tokens per parameter, requires 1.6 times more computational power when using data with the August 2026 AI token share. This means that training models on current web data, which includes a substantial amount of AI-generated text, incurs a significant "compute tax."
Existing scaling laws, such as those proposed in the Chinchilla paper, do not fully account for this phenomenon. The new research introduces a revised scaling law with separate components for the benefits and harms of AI text. This model reportedly reduces prediction errors by 41% on held-out model sizes compared to previous methods.
The implications of this research are substantial for any organization involved in collecting pretraining data from web crawls. The study quantifies the increased computational resources necessary to achieve desired model performance in the face of rising AI-generated content. The researchers have also released the WildAI corpus, a dataset of 83 billion tokens labeled for AI origin, topic, and other attributes, to aid further research in this area. The trend of increasing AI-generated text shows no immediate signs of slowing, suggesting that compute requirements for AI development will continue to be influenced by this factor.
