It’s hard to put this into words succinctly. But when I was a kid before the Internet, the library was the main source of human knowledge, and it was very organized through the Dewey Decimal System which imposed a kind of tree-like hierarchy across a wide spectrum of topics. And while I imagined that one day this might all wind up served by computers, I thought this organization would survive the transition. But it didn’t.

I suppose in the early days of the Internet, they tried? Yahoo was kind of an Internet directory in the beginning, and the Whole Internet catalog was another such effort. But these eventually fell apart and we had to resort to search engines to find anything. Frankly, it’s a bit like the early days of personal computing when we used flat file systems instead of hierarchical ones with proper directory trees. When you were limited to what could fit on a floppy, this wasn’t such a big burden. But the Internet as it stands might as well be a giant flat file system.

Today, even the search engines are failing us, and we are turning to AI. In the pre-Internet era, we had sort of an equivalent to this also. They were called librarians. But a librarian’s job was made easier by the fact that all the books were carefully organized by topic. For AI, it’s as though the library were just a massive pile of books tossed around haphazardly and the librarian had to make sense of it enough to pull whatever you’re looking for out of the chaos. Is it at all surprising, then, that it takes a huge amount of energy and computing resources to make any of this work?

  • Corngood@lemmy.ml
    link
    fedilink
    arrow-up
    5
    ·
    3 days ago

    It’s gotta be partly to do with content though, right? Are there any search engines that effectively sort through the slop?

    • jaybone@lemmy.zip
      link
      fedilink
      English
      arrow-up
      2
      ·
      2 days ago

      I meant to go back and edit or reply to another comment to walk back my statement a bit. There is a problem with content. The walled gardens hide content from search engines. So you can only find their content by being logged in to their service and using their search tools (which are also shit.) So yeah, there’s that too.

    • boonhet@sopuli.xyz
      link
      fedilink
      arrow-up
      2
      ·
      2 days ago

      Many people have noticed a decline going back to 2020 when Google decided they need to show more ads, which means you need to search more times or check more pages than the first one rather than getting the desired result right away.

      Kagi seems to be better at this since it’s paid and doesn’t need to show you ads.

      But of course you’re right that content itself has sloppified too.