this post was submitted on 12 Jul 2023
1 points (100.0% liked)
Technology
73792 readers
3270 users here now
This is a most excellent place for technology news and articles.
Our Rules
- Follow the lemmy.world rules.
- Only tech related news or articles.
- Be excellent to each other!
- Mod approved content bots can post up to 10 articles per day.
- Threads asking for personal tech support may be deleted.
- Politics threads may be removed.
- No memes allowed as posts, OK to post as comments.
- Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
- Check for duplicates before posting, duplicates may be removed
- Accounts 7 days and younger will have their posts automatically removed.
Approved Bots
founded 2 years ago
MODERATORS
The model has become inbred because it’s now impossible to scrape the web without AI content getting ingested, which is full of “hallucinations” and other weird artifacts. The last opportunity to get “uncontaminated” training data was sometime in mid 2022.
Not to say that it’s causing this particular problem, but this issue will emerge eventually. Garbage in = garbage out. Eventually GPT-19 will grow a mighty Habsburg chin.
Maybe not yet, but...
- Spez will turn Reddit into a bot farm and sell this as training data
- Musk turns Twitter into a bigoted cesspool and will sell this as training data, which will subsequently be flagged for low quality (also: a botfarm)
- Threads is a corporate ad dashboard (and we already know how easy it is to GPT copy) and Zuck will sell this as training data
- Facebook is either dead or only good for boomers and Poles
- blogs are dead
- Fediverse is out there waiting to be scraped but possibly too small to sustain a big model
We'te getting there, hopefully.