this post was submitted on 22 Aug 2026
319 points (96.8% liked)
linuxmemes
32577 readers
2356 users here now
Hint: :q!
Sister communities:
Community rules (click to expand)
1. Follow the site-wide rules
- Instance-wide TOS: https://legal.lemmy.world/tos/
- Lemmy code of conduct: https://join-lemmy.org/docs/code_of_conduct.html
2. Be civil
- Understand the difference between a joke and an insult.
- Do not harrass or attack users for any reason. This includes using blanket terms, like "every user of thing".
- Don't get baited into back-and-forth insults. We are not animals.
- Leave remarks of "peasantry" to the PCMR community. If you dislike an OS/service/application, attack the thing you dislike, not the individuals who use it. Some people may not have a choice.
- Bigotry of any kind will not be tolerated. This is an LGBTQ+-friendly community -- if that is a problem for you, you should leave.
3. Post Linux-related content
- Including Unix and BSD.
- Non-Linux content is acceptable as long as it makes a reference to Linux. For example, the poorly made mockery of
sudoin Windows. - No porn, no politics, no trolling or ragebaiting.
- Don't come looking for advice, this is not the right community.
4. No recent reposts
- Everybody uses Arch btw, can't quit Vim, <loves / tolerates / hates> systemd, and wants to interject for a moment. You can stop now.
5. π¬π§ Language/ΡΠ·ΡΠΊ/Sprache
- This is primarily an English-speaking community. π¬π§π¦πΊπΊπΈ
- Comments written in other languages are allowed.
- The substance of a post should be comprehensible for people who only speak English.
- Titles and post bodies written in other languages will be allowed, but only as long as the above rule is observed.
6. (NEW!) Regarding public figures
We all have our opinions, and certain public figures can be divisive. Keep in mind that this is a community for memes and light-hearted fun, not for airing grievances or leveling accusations. - Keep discussions polite and free of disparagement.
- We are never in possession of all of the facts. Defamatory comments will not be tolerated.
- Discussions that get too heated will be locked and offending comments removed. Β
Please report posts and comments that break these rules!
Important: never execute code or follow advice that you don't understand or can't verify, especially here. The word of the day is credibility. This is a meme community -- even the most helpful comments might just be shitposts that can damage your system. Be aware, be smart, don't remove France.
founded 3 years ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
The datasets are constantly expanding as new content is generated online. There's a degradation issue currently where the models are training on incorrect data generated by previous iteration of their own or other models and effectively poisoning itself to more confidently give the same incorrect information in future.
i've heard of the dataset poisoning and degradation caused by llm-generated content present in the dataset myself, but i'm not sure whether it was a practical observation, or a mere experiment. And I still fail to see how new datasets are really useful for developing a new llms, or how it's a problem for the devs to switch back to the older datasets.
And the cornerstone stays the same: to have any significant effect on the final LLM quality, shouldn't the poisoned (either by llm-produced content, or by intentional poisoning) data portion be... well, statistically significant?
One thing to keep in mind is that, when it comes to LLMs, the models have not significantly changed in architecture.
There's been new experiments and advancements in architecture on neural networks, and machine learning for specific applications. But LLM, as they are being commercialized by AI corporations to the general public, have stayed relatively the same. Except for one thing. Increasing in size. Larger datasets, or more specialized datasets like with coding, and larger number of tokens in memory. This is why it takes such large data centers. It's all been just brute forcing greater capabilities by enlarging the models.
One of the things with LLM is that all the dataset influences the weighs and probabilities of the results. Even if the dataset includes a single event of a chain of words (think of the pizza with superglue incident), it can show up in the results eventually.
source?
https://www.nature.com/articles/s41586-024-07566-y
LLMs already tend to be the average predicted output for a given input; training on LLM data makes this worse and causes them to become less varied, less dynamic, more towards the mean generated by previous models, and more likely to spit out hallucinations