25 points

Well, you’ve got a timestamped copy of much of the Web that existed up until latent-diffusion models at archive.org. That may not give you access to newer information, but it’s a pretty whopping big chunk of data to work with.

permalink
report
reply
20 points

Hopefully archive.org have measures in place to stop people from yanking all their data too quickly. As least not without a hefty donation or something. As a user it can chug a bit, and I’m hoping that’s the rate-limiting I’m talking about and not that they’re swamped.

permalink
report
parent
reply
7 points
*

That would go against the principal of the archive imo but regardless, if you take away all means of acquiring data freely, you are just giving companies like OpenAI and Google who already have copies of it an insane advantage.

AI isn’t going away, we need to make sure we have free access to it as to not give our whole economy to a handful of companies.

permalink
report
parent
reply
81 points

As junk web pages written by AI proliferate, the models that rely on that data will suffer.

Good.

permalink
report
reply
23 points

AI making itself sick and worthless after flooding the internet with trash just gives me a warm glow.

permalink
report
reply
1 point

interdasting

permalink
report
reply
57 points

Garbage in; Garbage out.

permalink
report
reply
19 points

Shit-fueled ouroboros

permalink
report
parent
reply
4 points

You can’t explain it!

permalink
report
parent
reply
2 points

Recycle the garbage that comes out… Still more garbage out.

permalink
report
parent
reply

Technology

!technology@lemmy.world

Create post

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related content.
  3. Be excellent to each another!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, to ask if your bot can be added please contact us.
  9. Check for duplicates before posting, duplicates may be removed

Approved Bots


Community stats

  • 17K

    Monthly active users

  • 6.1K

    Posts

  • 132K

    Comments