It’s hard to put this into words succinctly. But when I was a kid before the Internet, the library was the main source of human knowledge, and it was very organized through the Dewey Decimal System which imposed a kind of tree-like hierarchy across a wide spectrum of topics. And while I imagined that one day this might all wind up served by computers, I thought this organization would survive the transition. But it didn’t.
I suppose in the early days of the Internet, they tried? Yahoo was kind of an Internet directory in the beginning, and the Whole Internet catalog was another such effort. But these eventually fell apart and we had to resort to search engines to find anything. Frankly, it’s a bit like the early days of personal computing when we used flat file systems instead of hierarchical ones with proper directory trees. When you were limited to what could fit on a floppy, this wasn’t such a big burden. But the Internet as it stands might as well be a giant flat file system.
Today, even the search engines are failing us, and we are turning to AI. In the pre-Internet era, we had sort of an equivalent to this also. They were called librarians. But a librarian’s job was made easier by the fact that all the books were carefully organized by topic. For AI, it’s as though the library were just a massive pile of books tossed around haphazardly and the librarian had to make sense of it enough to pull whatever you’re looking for out of the chaos. Is it at all surprising, then, that it takes a huge amount of energy and computing resources to make any of this work?


I do some freelancing that involves online desk research, and for some topics, it’s like pulling teeth to find a source written by a real human. Even non-Google search engines can’t help but give me dozens of obviously AI-written pages. Which is why I as a human am getting paid to do real research. It’s literally wading through slop to find anything a human did.
The thing is, AI companies are terrified of model collapse because if they train AI systems on AI output, the model becomes useless. But the rate at which websites on any topic are just slop being scraped to train a new model seems to be either unsustainable for using the open internet for that training, or an assured way to reach model collapase quickly.
That’s actually not a fact, more of a possibility. As far as am I aware, reinforcement learning from AI feedback (RLAIF) is currently employed by some of the most capable frontier AI companies; Anthropic openly admits to it, and their models are arguably the best.
A term that is often thrown around nowadays is recursive self-improvement (RSI), and it seems that there are at this time more knowledgeable proponents of RSI for intelligence development than opponents who raise the risk of model collapse.
But yes, training on today’s sloppified public internet would probably only have downsides.
There are different theories around model collapse, so there might be a way to avoid it even if most of your training data is AI generated. The jury is still out though
That is just so depressing. It’s eating its own garbage to create more content. Like some image that’s been run through multiple passes of lossy compression with different algorithms.