OpenAI and Microsoft Confirm LLMs Were Trained on Stolen Web Content and Are Now Destroying It

OpenAI and Microsoft trained their language models on billions of web pages taken without permission. Those models now generate content that appears on the web. Search engines and users treat generated text as real information. The original web gets diluted with synthetic material. The next training run will use this contaminated web as source material.
This is a data laundering cycle. The stolen goods become the infrastructure. Nobody had to file paperwork because the entire process operates at scale before anyone notices the loop. The damage is baked into the next version before the current version ships.
The web becomes progressively less useful as training data. The models become progressively less accurate on downstream tasks. Eventually the systems train on their own errors. This continues until something changes. Nothing will change.