It’s hard to put this into words succinctly. But when I was a kid before the Internet, the library was the main source of human knowledge, and it was very organized through the Dewey Decimal System which imposed a kind of tree-like hierarchy across a wide spectrum of topics. And while I imagined that one day this might all wind up served by computers, I thought this organization would survive the transition. But it didn’t.
I suppose in the early days of the Internet, they tried? Yahoo was kind of an Internet directory in the beginning, and the Whole Internet catalog was another such effort. But these eventually fell apart and we had to resort to search engines to find anything. Frankly, it’s a bit like the early days of personal computing when we used flat file systems instead of hierarchical ones with proper directory trees. When you were limited to what could fit on a floppy, this wasn’t such a big burden. But the Internet as it stands might as well be a giant flat file system.
Today, even the search engines are failing us, and we are turning to AI. In the pre-Internet era, we had sort of an equivalent to this also. They were called librarians. But a librarian’s job was made easier by the fact that all the books were carefully organized by topic. For AI, it’s as though the library were just a massive pile of books tossed around haphazardly and the librarian had to make sense of it enough to pull whatever you’re looking for out of the chaos. Is it at all surprising, then, that it takes a huge amount of energy and computing resources to make any of this work?
Wow, there is a lot to think about in this comment thread! It’s going to take a lot more showers.
I’m not an information theorist, but I do wonder sometimes if nature prefers a more chaotic data organization? Even when you look at lower levels than the Internet. Hash tables seem to be winning over tree structures. Randomized network protocols are winning over orderly ones. And even if you look at AI, this used to be a broader term covering other topics such as expert systems which tried to condense human knowledge and a very low entropy way.
If there’s an inherent bias towards disorganization, where does it stem from? I guess there can be a somewhat better average performance, even if the worst case can be really, really bad. There may also be better resilience? Organized systems tend to fall apart very quickly when something goes wrong.
search engines are failing us
But this isn’t a problem with the internet content or search engine technology. This is because every search engine company just wants to show you ads and paid content rather than actually search for things anymore. Which 20 years ago, it used to do quite well.
It’s gotta be partly to do with content though, right? Are there any search engines that effectively sort through the slop?
yep
Search engines were, at one point, pretty fuckin good at finding what you needed, and would put it at the top of the first page of results
but then some companies coughgooglecough realized that users would be okay going to page 2 or 3 for their results, so they intentionally undid a decade+ of search engine improvements, and started forcefully putting relevant shit 2-3-4 pages in, so they could harvest more clicks and shove more ads in your face.
and thus enshitification began.
you’re completely right. there’s tons of problems with the way that we organize information today. another big issue is that a lot of publicly funded scientific research is behind paywalls of journals, and the journals don’t even have a good indexing system where you could easily search for stuff! it’s infuriating.
It seems like this problem stands to get worse too, as researchers see a need to protect their work before an LLM comes along and makes the big discovery on their backs.
Apologies if I’m reading too much into what you’ve said here, but it reads to me as if you’re implying search engines couldn’t keep up with the volume of the internet and AI is the logical next progression.
Search engines actually worked pretty amazingly at one point, but then were kneecapped because Google et al were more interested in advertising money than useful search results. Then when they needed an excuse to justify ludicrous spending on AI, they shoehorned AI into web search where it is very much the wrong tool for the job if you want accurate, high relevance results.
For AI, it’s as though the library were just a massive pile of books tossed around haphazardly and the librarian had to make sense of it enough to pull whatever you’re looking for out of the chaos. Is it at all surprising, then, that it takes a huge amount of energy and computing resources to make any of this work?
It’s more like the librarian has taken it upon themself (for some reason) to write a book that is the mathematical average of every other book in the library. Then when any patron comes looking for a specific book, the librarian just hands them this slag instead hoping they won’t notice the difference.
Now that said, I 100% agree that running a good “old-style” search engine does still take a huge amount of computing resources, but I’d be surprised if the requirements were anywhere near what is being thrown at training LLMs.
It’s more like the librarian has taken it upon themself (for some reason) to write a book that is the mathematical average of every other book in the library. Then when any patron comes looking for a specific book, the librarian just hands them this slag instead hoping they won’t notice the difference.
Honestly, best description of LLMs for search I’ve seen so far.
Ah, so you’re saying it’s more of an enshitification thing that killed the search engines? I can believe that to an extent.
But as the amount of content on the Internet has increased, it’s also become more difficult to narrow the queries to get what you’re really looking for. The main thing for me with AI is it allows you to refine the search without discarding the context of what you have already tried.
If the content were more organized to begin with, though, I suspect you could drill a bit deeper into the knowledge tree and then apply your conventional (non-AI) search over a more relevant branch?
It’s not just increase in content. Search engines don’t even do searches based on keywords or logic operators anymore. They run some NLP on the query input and do who knows what to give you results. If we could still make searcher that contain specifically certain terms or exclude others, it would be still possible to navigate into the mess
Google has a financial incentive for you to spend more time on their site. If you are gone in one click, you see less ads. Same problem with dating platforms. The user’s interest is in conflict with the platform’s.
There are leaks/interviews how their search and ads teams were at a conflict internally, ofc money won.
Does the ad revenue account for people who hardly realize the ads are even there?
Let’s not forget Search Engine Optimization (SEO) and content farms which were actively corrupting (adding entropy to) search engine results even before LLMs, yay capitalism. They’ve dialed it to twelve with LLMs.
Heh, I had actually forgotten about that. It seems almost quaint compared to where we’re at now.
Search algorithms are meant to work regardless of the amount of content or sorting. Page rank used to be a game changer back then to sort websites by relevance.
AI has no search algorithms. It’s about 1000% less efficient, because its algorithms are built for token prediction, not sifting through websites.
So yeah, my money is on enshitification.
Then when they needed an excuse to justify ludicrous spending on AI, they shoehorned AI into web search where it is very much the wrong tool for the job if you want accurate, high relevance results.
AI is very much the right tool for sorting through the high volume of garbage search results looking for nuggets of relevance. Google built it into search because if they didn’t, they would be cut out of search other than as an MCP provider. I’m sure they’d rather not spend compute running an AI summary of every single query, but they would be out of business the moment someone else provided that service.
Now is it perfect? Of course not. Verify from sources, but you get to those relevant sources faster and easier with AI doing the sifting because Google search has grown so fucking worthless.
AI isn’t better than Google was, but it’s better than Google is.
Google is not at fault for search becoming so terrible. Everyone who engaged in SEO is. The internet would have sloppified much sooner if Google and other search engines had not relentlessly pushed back against SEO and continuously modified their algorithms. Lots of people definitely underestimate the extreme negative impact of SEO, and the lengths that search engine operators have to go to to have their service be at least somewhat useful. As soon as you have a search algorithm, people will attempt to manipulate it.
One missing piece is that librarians are acting as curators of the data. Not anyone can come to the library, put books around, and game/pay the book finding process to make sure their books come first. Having human curators behind an internet search engine is not realistic.
Unless the library gets bought by some oil company and they force the librarian to add a lot of pro-oil books.
Those libraries exist and that’s why we publicly fund libraries.
Not anyone can come to the library, put books around, and game/pay the book finding process to make sure their books come first.
Well, as I recall, things weren’t quite idyllic in the conventional library scene either. People were always moving books around, making librarians furious. Every field trip to the library always began with that lecture. “For the love of all that is holy, do not put books back on the shelf. Let us librarians take care of that.” And once I got to universities with their reference collections, I found that people were purposefully hiding texts in other sections so that they would have exclusive access when they returned. Fun stuff.
I think you’re writing stories for librarians. Their whole existence relies on people using the books. They don’t want you to put the book away because, A- you’ll do it wrong and B- they count them for their general statistics which helps with their funding. They don’t want you being a pretend librarian, they went to school for it and are better at it than you.
People hide books, forget to return books and damage and destroy books. Librarians moniter and curate the collection. Libraries are not museums, the catalog ebs and flows and is a constant state of flux. I can tell you that me wife is happiest when she comes home from working a busy shift where the library was full of people.
Sure, there are the librarians that yell at kids and give you the stink eye for asking questions, but they’re just everyday-sadists and they’re everywhere.
Yes, the internet is decentralised in every way, not only its infrastructure. Even unorganised. If that’s your big discovery, I respect that. But…
the library was the main source of human knowledge, and it was very organized through the Dewey Decimal System
Not all of humanity uses the DD system.
Today, even the search engines are failing us, and we are turning to AI.
This simply isn’t true, and everything you write after that. It’s an insult to actual librarians.
There is no “Artifical Intelligence”.Indexing the internet is a big job. We had data centers even before A-not-I. And people were complaining about how much energy they use. Had we known then how much worse this is going to get very soon…
Artificial Unintelligence does not do that job better btw. It doesn’t do the indexing at all. It just wastes energy. And, as another commenter pointed out, it actively makes the situation you describe worse.
Yes, the internet is decentralised in every way, not only its infrastructure. Even unorganised.
What’s intriguing to me about the fediverse is it’s trying to layer some order over a decentralized physical infrastructure. I think that’s so cool and really want this sort of idea to work!
Not all of humanity uses the DD system.
Yeah that’s fair. I think even my university was using a different system, and of course there’s a lot more out there than book knowledge.
Indexing the internet is a big job. We had data centers even before A-not-I. And people were complaining about how much energy they use. Had we known then how much worse this is going to get very soon….
Artificial Unintelligence does not do that job better btw. It doesn’t do the indexing at all. It just wastes energy. And, as another commenter pointed out, it actively makes the situation you describe worse.
I may not be quite as cynical about AI in its present state, as I feel it has helped me find solutions that had been evading me. I do think it is massively overhyped, however, and am concerned about the directions it’s heading. Concerned a lot, actually.
fediverse is it’s trying to layer some order over a decentralized physical infrastructure.
I’m not sure I follow you. The Fediverse is about social networking which is not all, not even most of the internet.
And the internet itself is a layer of order over its decentralised physical infrastructure.
And the internet itself is a layer of order over its decentralised physical infrastructure.
That’s fair. But you can have different layers of order that manifest in various human pursuits, and the Fediverse seems a fertile ground for exploring such ideas.
In the early days, the Internet was just a bunch of institutional networks all linked together, with little islands of data that may have had some internal organization. But no big picture view existed, and it never really emerged.
It seems to me that what’s happened in recent years is individual companies that offer some sort of information service have been centralizing and gatekeeping their product more and more, which does bring some order but at the expense of what makes a decentralized network powerful in the first place.
I do some freelancing that involves online desk research, and for some topics, it’s like pulling teeth to find a source written by a real human. Even non-Google search engines can’t help but give me dozens of obviously AI-written pages. Which is why I as a human am getting paid to do real research. It’s literally wading through slop to find anything a human did.
The thing is, AI companies are terrified of model collapse because if they train AI systems on AI output, the model becomes useless. But the rate at which websites on any topic are just slop being scraped to train a new model seems to be either unsustainable for using the open internet for that training, or an assured way to reach model collapase quickly.
The thing is, AI companies are terrified of model collapse because if they train AI systems on AI output, the model becomes useless.
That’s actually not a fact, more of a possibility. As far as am I aware, reinforcement learning from AI feedback (RLAIF) is currently employed by some of the most capable frontier AI companies; Anthropic openly admits to it, and their models are arguably the best.
A term that is often thrown around nowadays is recursive self-improvement (RSI), and it seems that there are at this time more knowledgeable proponents of RSI for intelligence development than opponents who raise the risk of model collapse.
But yes, training on today’s sloppified public internet would probably only have downsides.
There are different theories around model collapse, so there might be a way to avoid it even if most of your training data is AI generated. The jury is still out though
That is just so depressing. It’s eating its own garbage to create more content. Like some image that’s been run through multiple passes of lossy compression with different algorithms.
In the early days of the internet there was Gopher which had a strong hierarchy for sorting and returning docs: https://en.wikipedia.org/wiki/Gopher_(protocol)
We could resurrect it, but sadly the AI Slop will still be there I imagine.








