this post was submitted on 22 Aug 2026
319 points (96.8% liked)

linuxmemes

32577 readers
2188 users here now

Hint: :q!


Sister communities:


Community rules (click to expand)

1. Follow the site-wide rules

2. Be civil
  • Understand the difference between a joke and an insult.
  • Do not harrass or attack users for any reason. This includes using blanket terms, like "every user of thing".
  • Don't get baited into back-and-forth insults. We are not animals.
  • Leave remarks of "peasantry" to the PCMR community. If you dislike an OS/service/application, attack the thing you dislike, not the individuals who use it. Some people may not have a choice.
  • Bigotry of any kind will not be tolerated. This is an LGBTQ+-friendly community -- if that is a problem for you, you should leave.
  • 3. Post Linux-related content
  • Including Unix and BSD.
  • Non-Linux content is acceptable as long as it makes a reference to Linux. For example, the poorly made mockery of sudo in Windows.
  • No porn, no politics, no trolling or ragebaiting.
  • Don't come looking for advice, this is not the right community.
  • 4. No recent reposts
  • Everybody uses Arch btw, can't quit Vim, <loves / tolerates / hates> systemd, and wants to interject for a moment. You can stop now.
  • 5. πŸ‡¬πŸ‡§ Language/язык/Sprache
  • This is primarily an English-speaking community. πŸ‡¬πŸ‡§πŸ‡¦πŸ‡ΊπŸ‡ΊπŸ‡Έ
  • Comments written in other languages are allowed.
  • The substance of a post should be comprehensible for people who only speak English.
  • Titles and post bodies written in other languages will be allowed, but only as long as the above rule is observed.
  • 6. (NEW!) Regarding public figuresWe all have our opinions, and certain public figures can be divisive. Keep in mind that this is a community for memes and light-hearted fun, not for airing grievances or leveling accusations.
  • Keep discussions polite and free of disparagement.
  • We are never in possession of all of the facts. Defamatory comments will not be tolerated.
  • Discussions that get too heated will be locked and offending comments removed.
  • Β 

    Please report posts and comments that break these rules!


    Important: never execute code or follow advice that you don't understand or can't verify, especially here. The word of the day is credibility. This is a meme community -- even the most helpful comments might just be shitposts that can damage your system. Be aware, be smart, don't remove France.

    founded 3 years ago
    MODERATORS
     
    top 50 comments
    sorted by: hot top controversial new old
    [–] ICastFist@programming.dev 9 points 10 hours ago (1 children)

    Time to get back to Assembly, I guess.

    But do you trust the CPU?

    Well, fuck

    [–] AnUnusualRelic@lemmy.world 2 points 7 hours ago

    The CPU runs Minix, are you going to tell me you don't trust Minix now?

    [–] demizerone@lemmy.world 6 points 1 day ago

    I think about this every day.

    [–] rizzothesmall@sh.itjust.works 17 points 1 day ago (3 children)

    Maybe? If you poison the prompt then there's evidence and it can be undone. Poison fragments of the source training data, however, and that's some KT shit right there. Enterprise foundation models cost bonkers money to train and pretty much slurp up all the data on the internet for mostly automated annotation. Stick something in an obscure part of the internet which becomes part of the training and produces the malicious response and it's going to be both hard and expensive to detect or correct.

    [–] ICastFist@programming.dev 1 points 1 hour ago

    Should be relatively easy, with the amount of once trusted packages that become attack vectors

    [–] CheesyFox@lemmy.sdf.org 7 points 1 day ago (1 children)

    except for poison to take in, it should be a pretty significant part of the dataset. Also, ngl, i'm not much informed on the topic, but aren't all the datasets, if we're talking about generic diffusion models and LLMs, already been formed? From what i gather, the innovation in AI mainly comes from utilizing new architectures, rather than training a model on something unique.

    [–] rizzothesmall@sh.itjust.works 8 points 1 day ago (3 children)

    The datasets are constantly expanding as new content is generated online. There's a degradation issue currently where the models are training on incorrect data generated by previous iteration of their own or other models and effectively poisoning itself to more confidently give the same incorrect information in future.

    [–] CheesyFox@lemmy.sdf.org 2 points 1 day ago* (last edited 1 day ago) (1 children)

    i've heard of the dataset poisoning and degradation caused by llm-generated content present in the dataset myself, but i'm not sure whether it was a practical observation, or a mere experiment. And I still fail to see how new datasets are really useful for developing a new llms, or how it's a problem for the devs to switch back to the older datasets.

    And the cornerstone stays the same: to have any significant effect on the final LLM quality, shouldn't the poisoned (either by llm-produced content, or by intentional poisoning) data portion be... well, statistically significant?

    [–] dustyData@lemmy.world 1 points 7 hours ago* (last edited 6 hours ago)

    One thing to keep in mind is that, when it comes to LLMs, the models have not significantly changed in architecture.

    There's been new experiments and advancements in architecture on neural networks, and machine learning for specific applications. But LLM, as they are being commercialized by AI corporations to the general public, have stayed relatively the same. Except for one thing. Increasing in size. Larger datasets, or more specialized datasets like with coding, and larger number of tokens in memory. This is why it takes such large data centers. It's all been just brute forcing greater capabilities by enlarging the models.

    One of the things with LLM is that all the dataset influences the weighs and probabilities of the results. Even if the dataset includes a single event of a chain of words (think of the pizza with superglue incident), it can show up in the results eventually.

    load more comments (2 replies)
    [–] Onomatopoeia@lemmy.cafe 86 points 2 days ago (1 children)
    [–] SpaceNoodle@lemmy.world 30 points 2 days ago

    Dev is screwed.

    [–] rtxn@lemmy.world 65 points 2 days ago (3 children)
    [–] Valmond@lemmy.dbzer0.com 6 points 1 day ago

    Quis custodiet ipsos custodes?

    [–] Redjard@reddthat.com 19 points 2 days ago (1 children)

    I've had it on my todo for years to work through ddc and trusting trust.
    Which is a method to verify a compiler is matching its source and thus trustworthy.

    An orthogonal approach is reproducible builds, which among many benefits can make sure a few people verifying things benefit everyone who can then see they have the same verified binaries.

    [–] eah@programming.dev 4 points 23 hours ago
    load more comments (1 replies)
    [–] not@lemmy.dbzer0.com 32 points 2 days ago (3 children)
    [–] Natanox@discuss.tchncs.de 7 points 1 day ago (1 children)

    The ending is rather unsatisfactory.

    [–] einfach_orangensaft@sh.itjust.works 3 points 1 day ago (1 children)

    spoilerSome how like all sophisticated technical stories about ai it ends with the acceptance that there is nothing we can do because the alternative would be going back to analog.

    [–] Natanox@discuss.tchncs.de 4 points 23 hours ago

    Sounds like quitter talk to me. It's not like AI is an evolutionary process that just happens somewhere, it's a localized tool - and even if it spreads itself, that would be by far the most humonguous virus ever. I think those authors just write down their own fears.

    load more comments (2 replies)
    [–] abbadon420@sh.itjust.works 42 points 2 days ago (1 children)

    The very fact that Anthropic is now injecting a kind of watermark into every output, is solid proof that such a Ken Thompson hack is a inevetable risk

    [–] AudaciousArmadillo@piefed.blahaj.zone 15 points 1 day ago (1 children)

    Ugh. Fuck "AI" and fuck Anthropic. But please read how the "watermarks" work. TL;TR its like a seeded run in a video game. With the seed and pseudo rng, you get the outcome i.e. the extruded text. In the watermark its the reverse, outcome + prng = seed. The result will be the same "quality" extruded garbage as before.

    [–] douglasg14b@lemmy.world 10 points 1 day ago (1 children)

    Good luck.

    Lemmy is damn near the when it comes to wanting to hold an opinion on a topic without having first understood that topic.

    [–] Axolotl_cpp@feddit.it 12 points 1 day ago

    That's just a reflection of the real life and not unique to Lemmy, dw

    [–] irotsoma@piefed.blahaj.zone 22 points 2 days ago (1 children)

    Devs should be "dev managers and executives". Real developers know LLMs are basically just a tool for finding examples and helping with syntax. Sure they're useful, but I'd never let them write code, much less compile it. Who knows what they'd inject into a build.

    [–] Natanox@discuss.tchncs.de 4 points 1 day ago

    Real developers know LLMs are basically just a tool for finding examples and helping with syntax.

    Your words on gods ears. I also found it to be reasonably useful when you're stuck in the documentation of a library clearly written by and for people who already know it.

    Unfortunately people who do not understand code are blissfully ignorant at what garbage they have the machine spit out. Of course until the LLM nukes their whole project folder or even disk because the probability engine unfortunately picked "clean slate" as the most probable next thing. Not as if that ever happened at companies, lol.

    load more comments
    view more: next β€Ί