We Didn’t Just Train AI on the Internet. We Started Training It on Itself.

TL;DR AI
2 min readKey summary
The web is shifting from mostly human-made content to AI-generated or AI-rewritten material, creating a feedback loop for model training.
This raises concerns that future models may learn from increasingly synthetic data, making them more uniform and less original.
The article argues that more compute alone will not fix the loss of high-quality human signals and content diversity.
As a result, AI labs are turning to proprietary human datasets from sources like Reddit, GitHub, Stack Overflow, and publisher archives.
