this post was submitted on 03 Sep 2026
252 points (97.7% liked)

News

39156 readers
2495 users here now

Welcome to the News community!

Rules:

1. Be civil


Attack the argument, not the person. No racism/sexism/bigotry. Good faith argumentation only. This includes accusing another user of being a bot or paid actor. Trolling is uncivil and is grounds for removal and/or a community ban. Do not respond to rule-breaking content; report it and move on.


2. All posts should contain a source (url) that is as reliable and unbiased as possible and must only contain one link.


Obvious biased sources will be removed at the mods’ discretion. Supporting links can be added in comments or posted separately but not to the post body. Sources may be checked for reliability using Wikipedia, MBFC, AdFontes, GroundNews, etc.


3. No bots, spam or self-promotion.


Only approved bots, which follow the guidelines for bots set by the instance, are allowed.


4. Post titles should be the same as the article used as source. Clickbait titles may be removed.


Posts which titles don’t match the source may be removed. If the site changed their headline, we may ask you to update the post title. Clickbait titles use hyperbolic language and do not accurately describe the article content. When necessary, post titles may be edited, clearly marked with [brackets], but may never be used to editorialize or comment on the content.


5. Only recent news is allowed.


Posts must be news from the most recent 30 days.


6. All posts must be news articles.


No opinion pieces, Listicles, editorials, videos, press releases, or celebrity gossip will be allowed. All posts will be judged on a case-by-case basis. Mods may use discretion to pre-approve videos or press releases from highly credible sources that provide unique, newsworthy content not available or possible in another format.


7. No duplicate posts.


If an article has already been posted, it will be removed. Different articles reporting on the same subject are permitted. If the post that matches your post is very old, we refer you to rule 5.


8. Misinformation is prohibited.


Misinformation / propaganda is strictly prohibited. Any comment or post containing or linking to misinformation will be removed. If you feel that your post has been removed in error, credible sources must be provided.


9. No link shorteners or news aggregators.


All posts must link to original article sources. You may include archival links in the post description. News aggregators such as Yahoo, Google, Hacker News, etc. should be avoided in favor of the original source link. Newswire services such as AP, Reuters, or AFP, are frequently republished and may be shared from other credible sources.


10. Don't copy entire article in your post body


For copyright reasons, you are not allowed to copy an entire article into your post body. This is an instance wide rule, that is strictly enforced in this community.

founded 3 years ago
MODERATORS
 

Musk denied he was aware Grok ever produced ‘any naked underage images’

A survivor of child sexual abuse has sued Elon Musk’s artificial intelligence company, alleging that its chatbot used pictures of her abuse to generate new illegal pornographic images that depict her.

“Using real images of Plaintiff and class members, Grok generated child pornography depicting Plaintiff and class members,” states the complaint, which was filed last week in a US district court in California.

Attorneys for the plaintiff, who is listed as Jane Doe in the case to protect her identity, accuse xAI of both generating CSAM of their client and of ingesting child sexual abuse images depicting her into the company’s datasets after new images were publicly posted.

top 13 comments
sorted by: hot top controversial new old
[–] Digestive_Biscuit@feddit.uk 6 points 6 days ago (3 children)

How does their AI engines (whatever they are called) access such material? I'm assuming its from the dark web or a government library or something and if that's the case then somebody made a decision to include such questionable sources.

[–] IAmYouButYouDontKnowYet@reddthat.com 1 points 3 days ago* (last edited 3 days ago)

I don't think it's too far fetched to say our governments traffick kids.

[–] Ithral@lemmy.blahaj.zone 6 points 6 days ago

Model weights is what they are called. They didnt access them exactly, if the allegations are true, what hapoened was during training those images were downloaded and used to train the image generation weights. So the problem here is the training data was not carefully sanitized and was just randomly downloaded from everywhere resulting in illegal content slipping in.

[–] partofthevoice@lemmy.zip 4 points 6 days ago

It’s more like if you kept the recipe to the photo, rather than the photo itself. The AI model, when trained, is having its “weights” updated — which means they’re tuning a very long set of lists of numbers.

For example, imagine:

(
  [0.028474, 0.274729, …],
  [0.827482, 0.283759, …],
  …*billions of lists
)

The numbers are actually random generated at first. Training can involve tricks like cutting out pieces of an image, then telling an AI to predict what goes in the empty space. Or reducing a photos quality, then telling the AI to increase its quality. The important part is that you distort the image while retaining the original as the answer key.

For each answer, you measure how correct/incorrect the model was. If it’s correct, you don’t do anything. If it’s incorrect, you do mathematical tricks to update those numbers.

Those numbers bias the model, all the way from random gibberish output to coherent output. So they go through the process I mentioned before billions of times, each time with different pictures / distortions, each time ever so slightly nudging those numbers until the model spits out coherent outputs.

Those numbers wind up being something like a recipe. Like if you’d stored the exact pixel configuration to a photo, but didn’t actually have the photo. Except this recipe tries to measure the general semantic relationships between words and images, such that it can generate images when given prompts. This recipe is less deterministic, but nonetheless it’s tainted with CP shit.

[–] zecg@lemmy.world 26 points 1 week ago* (last edited 1 week ago) (2 children)

The distinction of using pre-existing CSAM is notable because law enforcement and child protection organizations frequently give those illegal materials a kind of digital fingerprint – known as a hash – which allows them to track images and videos when they appear online. In this case, attorneys for the plaintiff stated that the Canadian Centre for Child Protection used images’ fingerprints to identify AI-generated CSAM on X that depicted their client.

How does that work? Are there hashes of child sex abuse images somehow reproduced by LLM? Surely the new images don't have the same hashes. Guardian is not really being clear here.

edit: Ars' is less badly written: "This is the first case to accuse xAI of training on CSAM, and the complaint does not go into great detail on that claim. Previously, Ars reported on a controversial dataset that was later scrubbed after researchers found CSAM in the training data, but there’s no indication xAI trained on that data. In a press release from lawyers representing Doe, it explained that Doe’s images were included in a CSAM Hash List maintained by NCMEC, and “that same material” allegedly “was part of the dataset xAI used to build Grok’s image and video generating capabilities.” The complaint similarly only alleged that “CSAM depicting Plaintiff with its longstanding well-known hash values has been used as a part of the dataset used by xAI.”"

[–] cantstopthesignal@sh.itjust.works 7 points 1 week ago* (last edited 1 week ago) (3 children)

From my understanding a hash completely changes when a single part of it is changed. That's the point of a hash. Sounds like they identified the hashed images from the training set.

[–] Clent@lemmy.dbzer0.com 18 points 1 week ago (1 children)

Nope. They use perceptual hashing. That way water marking and cropped images still trigger against the casm scanners.

The known casm results in a specific hash. Modified versions of that will result in a different hash but such that the "distance" is meaningful. The distance is how many bits need to flip to match the known casm hash.

It can lead to false positives so that's why manually checking occurs. It's also why people got very upset that Apple and perhaps others were going to automatically do these checks during upload to the cloud. When performed against billions of images, the false positive rate becomes huge and manually checking becomes too much overhead so people would be falsely accused.

It's very likely what they found in groks models is a false positive but in this case they can probably compel twitter to prove it's a false positive.

[–] Cypher@aussie.zone 2 points 6 days ago

It's also why people got very upset that Apple and perhaps others were going to automatically do these checks during upload to the cloud.

They were planning on doing checks locally, potentially impacting device performance, battery life and privacy as false positives would still be reviewed by humans. Potentially leading to individuals intimate photos being viewed without consent.

you would think, but then it would be lost when any part of the image changed (like adding a watermark or jpeg compression) instead it is a reversable algorithm https://www.researchgate.net/figure/mage-reconstructions-from-PhotoDNA-hashes-4_fig4_374838370

[–] FiskFisk33@startrek.website 1 points 1 week ago

there are other, more robust, ways of fingerprinting images than hashing the data. I'm not sure how it works exactly though, but I can imagine it is not completely different from how shazam can pick out a song in a noisy environment.

[–] Grimy@lemmy.world 3 points 1 week ago* (last edited 1 week ago)

Ya it doesn't make sense. I think it's probably facial recognition and since she's in the training dataset, her face pops up when the model gets pulled in that direction.

They wouldn't have access to the dataset, I don't see how they could know for the Arts explanation. Maybe it learned the hash and reproduces it but that doesn't seem likely. Even the part about adding hashs to track the images doesn't make much sense. That would imply they are distributing it somehow.

I don't get why X isn't running age detection on the output though.

[–] RagingRobot@lemmy.world 17 points 1 week ago (1 children)

If he was unaware why is he sewing to block a law restricting this? How could he possibly be unaware of what's going on inside his own and the whole world knows?

[–] Tollana1234567@lemmy.today 1 points 1 week ago

his whole userbase on x are nazis+CSAM users.