AI trained on AI garbage spits out AI garbage.(www.technologyreview.com)

posted 4 months ago

ModerateImprovement@sh.itjust.works

technology@lemmy.world

58 commentshide report

Sort:

Hot Top Controversial New Old

[ - ]

Admiral Patrick@dubvee.org

81 points

4 months ago

As junk web pages written by AI proliferate, the models that rely on that data will suffer.

Good.

permalink

report

[ - ]

Catoblepas@lemmy.blahaj.zone

23 points

4 months ago

AI making itself sick and worthless after flooding the internet with trash just gives me a warm glow.

permalink

report

[ - ]

Lvxferre@mander.xyz

36 points

4 months ago

Model degeneration is an already well-known phenomenon. The article already explains well what’s going on so I won’t go into details, but note how this happens because the model does not understand what it is outputting - it’s looking for patterns, not for the meaning conveyed by said patterns.

Frankly at this rate might as well go with a neuro-symbolic approach.

permalink

report

[ - ]

CeeBee_Eh@lemmy.world

-6 points

4 months ago

The issue with your assertion is that people don’t actually work a similar way. Have you ever met someone who was clearly taught "garbage’?

permalink

report

parent

[ - ]

Lvxferre@mander.xyz

11 points

4 months ago

The issue with your assertion is that people don’t actually work a similar way.

I’m talking about LLMs, not about people.

permalink

report

parent

[ - ]

CeeBee_Eh@lemmy.world

-11 points

4 months ago

I know you are, but the argument that an LLM doesn’t understand context is incorrect. It’s not human level understanding, but it’s been demonstrated that they do have a level of understanding.

And to be clear, I’m not talking about consciousness or sapience.

report

[ - ]

14 points

4 months ago

I’d be very wary of extrapolating too much from this paper.

The past research along these lines found that a mix of synthetic and organic data was better than organic alone, and a caveat for all the research to date is that they are using shitty cheap models where there’s a significant performance degrading in the synthetic data as compared to SotA models, where other research has found notable improvements to smaller models from synthetic data from the SotA.

Basically this is only really saying that AI models across multiple types from a year or two ago in capabilities recursively trained with no additional organic data will collapse.

It’s not representative of real world or emerging conditions.

permalink

report

[ - ]

_haha_oh_wow_@sh.itjust.works

1 point

4 months ago

interdasting

permalink

report

Technology

!technology@lemmy.world

Create post

This is a most excellent place for technology news and articles.

Our Rules

Follow the lemmy.world rules.
Only tech related content.
Be excellent to each another!
Mod approved content bots can post up to 10 articles per day.
Threads asking for personal tech support may be deleted.
Politics threads may be removed.
No memes allowed as posts, OK to post as comments.
Only approved bots from the list below, to ask if your bot can be added please contact us.
Check for duplicates before posting, duplicates may be removed

Approved Bots

Community stats

17K
Monthly active users
6.1K
Posts
132K
Comments

Our Rules

Approved Bots

Community stats

Community moderators