I really hope they die soon, this is unbearable…

  • Ephera@lemmy.ml
    link
    fedilink
    English
    arrow-up
    7
    ·
    4 hours ago

    My best guess is that they don’t just index things, but rather download straight from the internet when they need fresh training data. They can’t really cache the whole internet after all…

    • Techlos@lemmy.dbzer0.com
      link
      fedilink
      English
      arrow-up
      6
      ·
      4 hours ago

      Bingo, modern datasets are a list of URL’s with metadata rather than the files themselves. Every new team/individual wanting to work with the dataset becomes another DDoS participant.

    • Spice Hoarder@lemmy.zip
      link
      fedilink
      English
      arrow-up
      2
      ·
      3 hours ago

      The sad thing is that they could cache the whole internet if there was a checksum protocol.

      Now that I’m thinking about it, I actually hate the idea that there are several companies out there with graph databases of the entire internet.