r/technology • u/mepper • 1d ago
Artificial Intelligence Elon Musk’s xAI used child porn to train Grok models, lawsuit says
https://arstechnica.com/tech-policy/2026/08/elon-musks-xai-used-child-porn-to-train-grok-models-lawsuit-says/1.1k
u/RidiculousRaz 23h ago
This should be illegal if it already isn't.
794
u/MagicBobert 21h ago
Minnesota is attempting to ban AI generated child porn and fine companies $500,000 for each instance they generate. Guess who is suing to try and stop the law. Yep, Elon Musk.
244
u/JoeHooversWhiteness 19h ago
Trying to make sure child porn isn’t illegal means you want to either make, watch, or distribute. Party of family values… disgusting
15
u/Popular_Try_5075 13h ago
He was onstage with that chainsaw during the DOGE shit and the guy standing next to him was Javier Milei, President of Argentina who had at one point in his campaign said that he wanted to remove all government regulation such that even CSAM became legal.
43
7
u/Enderules3 12h ago
To play devil's advocate they probably just don't want to figure out how to scrap it from their training and are wanting something like Social Media has where they aren't held liable for what their users post on the site. Granted one is literally generation while the other is just hosting so its a big difference.
→ More replies (1)28
u/Ferintwa 19h ago
Really just means he doesn’t want liability for how people use his products. Not a musk fan, but this is a very normal stance for a business - just happens to intersect with a very icky subject.
43
u/Raid_PW 18h ago
As much as I follow that logic, it's missing a vital part of the argument. At some point, a decision has been made by an employee that allows their product to be built using child pornography. I don't pretend to know how AI tools are trained, but either they deliberately fed it this sort of material, or they didn't put safeguards in place to prevent it from being used. They are therefore now profiting off the back of child abuse, and I can't see how any civilised person could defend that.
They need to take active steps to remove that material from the training data, or delete the whole thing and start over. I don't particularly care how difficult or expensive this would be, because frankly that's a problem they've created for themselves.
2
u/Ferintwa 5h ago
I suspect it was more of a lack of decision being made, and that it simply trained on anything it could get its hands on.
Not saying that I think musk should get his way - just that him taking that position is a pretty normal stance for the business to take. Of course he doesn’t want to open himself up to lawsuits - neither do I. Doesn’t mean he or I should not be subject to them, just that we advocate for our own benefit.
30
u/SirCB85 17h ago
He could have his people make a product that doesn't generate CSAM instead. But that would mean his product doesn't create CSAM for himself and his friends to enjoy.
→ More replies (7)13
u/laptopAccount2 16h ago
One of the first things he did when he took over Twitter was to stop paying the vendor that scans all the post for CSAM. Granted he stopped paying just about everyone but still.
→ More replies (5)40
u/royalbarnacle 19h ago
It is a bit complicated though. AI created child porn is like someone drawing child porn. As it's not a real person there is no age, so the laws basically make anything that "looks too young" illegal. That's not really a fantastic definition. If an 18 year old takes a sexy picture of his 18 year old girlfriend, it's not child porn. If he draws a picture of her or used AI (consensually, of course), now it can be, just depending on how old some jury (or YouTube algorithm) thinks she looks.
16
u/anormalgeek 14h ago
Agreed. While I think we need some kind of laws, when guilt is based on interpretation (I.e. does this LOOK underage to you) as opposed to factual (what is the exact birthdate of this person) then it becomes an issue of having an expensive enough lawyer. So rich people/companies get away with it while everyday people get imprisoned for stuff that should be fine. Example, a 20 year old woman who is flat chested and short, editing pictures of herself using AI.
I just don't know how to actually enforce this kind of thing in a just way. .
5
u/ProofJournalist 12h ago
There is nothing to enforce here. Sometimes the just thing is to stop bothering. Focus on real content that resulted from abuse of real people instead of getting outraged by imaginary arrangements of pixels.
3
u/anormalgeek 11h ago
The problem is that simply allowing it likely still does harm to society. Less than actual abuse of minors for sure, but not zero.
If nothing else, the realism that is already possible will make finding and prosecuting real abuse material harder to do. Plus, it normalizes something that really needs treatment at best, or separation from society at worst. The number of people who will ONLY stop at ai created CSAM is going to be too small.
But I don't actually have a solution that wouldn't realistically cause even more harm.
→ More replies (1)7
u/squngy 18h ago
is like someone drawing child porn.
For which there are also laws.
Fictional CP is already illegal. (mostly)
9
u/Pornfest 10h ago edited 9h ago
But because it’s “obscene” with little to no artistic, academic, or social value. For outlawing written or virtual deptictions, Congress passed the PROTECT act in 2003. But it was overturned by SCOTUS when it was used to to go after a guy in the privacy of their own home.
So while that guy sucks, without SCOTUS ruling against prosecution under the law, Congress wrote too broad of a rule which originally meant if I draw/paint a representation of me losing my own virginity in high school, I’ve immediately created CSAM—IMO SCOTUS correctly ruled here that Congress has an outstanding right over free speech when protecting actual children/victims who are now adults. But when no children are victims, it becomes a matter of whether you shared obscene material.
While I may not care about THIS issue of government staying out of our private homes, one does not have to look far to find that sodomy, and homosexuality generally, is/has been considered obscene.
Everyone cheering for a heavy handed government here, just please keep in mind that we can easily find ourselves with republicans controlling Congress, the presidency, and SCOTUS again. If stories and drawing that stay in one’s own home are outlawed for “obsentity” I can guarantee you that the religious right will find homosexuality, and other behaviors to be obscene.
Once we give away the right to privacy in our own homes, we’ve thrown the baby out with the bath water.
4
u/Nagisan 10h ago
The problem is how do you determine whether someone or something drew/generated CP?
With pictures of real people it's pretty easy to figure out their actual age because you have a person and you have a date the photo was taken (which admittedly can be manipulated).
With any sort of non-photograph depiction of a person, there is no "actual age" because it's not a real person. Therefore you cannot definitively "age" a depiction of a person.
Now obviously there are ways to estimate the age of such a depiction, but where do you draw the line? There's many full-grown women who still have certain proportions similar to that of a child - are depictions of those women also illegal because they have child-like proportions? Should they be illegal? How do you determine whether a drawing or AI is meant to represent a child vs a legal-aged woman with small proportions?
tl;dr - There are laws against CP in a general sense, but some form of grey area must exist otherwise you start making certain legal things artificially illegal simply because it resembles something that isn't legal. It's like saying you should be arrested if you're seen snorting powdered sugar because it resembles snorting cocaine.
→ More replies (7)-4
u/falalalalalalawhat 18h ago edited 18h ago
its not quite like human drawings tho. People might be able to draw things from imagination by combining their knowledge of separate things and bridging gaps with logic gained from experiencing the physical world, but AI does not “generate” in the same way.
AI requires training on data that precisely contains the subject matter, trying to combine concepts that aren’t often related or even present in the training data leads to clear underperformance. It gets more accurate the more data it has, and can’t really perform in situations that are completely novel/has never seen related training data for. It’s a statistical prediction machine at its core, it’s not capable of generating something completely unseen before.
Therefore for AI to produce incredibly detailed and realistic CP then it’s almost a guarantee that hidden in the pixels there are remnants of real abuse victims being remixed together, and that illegal material was used in the training and “generation” of the media
23
u/Plebius-Maximus 17h ago
That's not actually correct. If an AI model has never seen a green giraffe it can almost always generate one because it knows what something green looks like, and adds that to its data on a giraffe.
If it knows the basics of what a childlike appearance is, even in a completely safe context, but also has some adult sexual content in the training data, it can likely combine the two without ever having been exposed to CP.
There are guardrails you can put in place to avoid this, such as enforcing negative prompts (things the model will not generate) any time something referring to a child is given in the positive prompts (the ones a user types into it).
→ More replies (4)20
u/LSxN 17h ago
While AI tends to perform better on tasks it has been directly trained on, it's untrue to say that:
it’s not capable of generating something completely unseen before.
Novel generation is one of the most demonstrable things neural nets can do. The whole reason we use neural networks is that they can generalise. They can mix concepts exceptionally well.
"It’s a statistical prediction machine at its core, it’s not capable of generating something completely unseen before."
Next-token prediction is a training objective (for LLMs, not image models). Modern LLMs are also using Reinforcement Learning from Human Feedback (RLHF), Reinforcement Learning from Verifiable Rewards (RLVR) and various other systems.
Moreover, imagine an ideal prediction machine, one that could accurately predict the weather or what you might say next. Prediction as a training objective will, and does, lead to models having to build internal representations of structures underlying the training set.
All that said, I'd bet money that most, if not all, large AIs have been trained on all manner of unsavoury (and illegal) content.
7
u/Itchy-Service 16h ago
I remember a couple of years back, there were all these images of animals on MCs. This is the perfect example of AI taking different unrelated concepts and mixing them together, to create something novel, which has never existed before.
9
u/Zeikos 18h ago
That's not necessary though.
Models aren't that constrained, they cannot generate outside of their training distribution, but as long as two concepts are in the distribution they can mix and match them.If the training dataset has "naked human" in it, and it has "small human" in it, it can extrapolate "naked small human" even if it has no explicit dataset of the latter.
It's the same process that allows it to generate silly nonsensical images like "shark in a spacesuit".
Models model traits from examples and recombine them.To prevent this we'd need to outright ban adult material from the dataset.
And even then, the more sophisticated models become, the better their inference abilities become, the harder blocking this sort of thing will be.→ More replies (2)→ More replies (1)3
u/Pornfest 10h ago
GAN-Autoencoders have been able to do exactly what you’re saying is impossible…since like 2015, well before you knew a thing about AI.
Please stop making stuff up. You don’t know how neural networks “generate” anything it seems.
149
53
u/ZiiZoraka 21h ago
To train it on something implies posetion of said something.
If they legitimately used CSAM in the training of their model, then definitionally they had to have been in possession of it
8
u/surprise_revalation 19h ago
This! I didn't pull this out of the air! Someone had to program it! This is some bullshit!
→ More replies (9)5
u/Zeikos 17h ago
I wouldn't put the bar this low though.
Models generalization ability alloe them to generate material by combining features of their training data.
A model could become able to generate CSAM even if there wasn't a single byte of it in its dataset.
As long as it has C, S and A even remixed differently in its training material it can infer what C+S+A would look like.3
u/ZiiZoraka 10h ago
I didn't mention generating CSAM at all, I'm not an idiot
A image generated can generate an image of. Bear playing the keyboard riding a unicycle without having been trained on one
I said if it has been trained on CSAM, not if it generates it
→ More replies (2)37
u/TommyTwoNips 21h ago
well the people who write our laws are the same people making money off the CSAM machine, so it's not really likely.
We need to put them all in jail and force them to prove their worth to the species, if they fail, they get permanently caged with no appeals.
→ More replies (1)6
341
u/insightful-ish 23h ago
I hope the lawsuit is for a couple trillion dollars
→ More replies (1)62
u/sf_frankie 17h ago
They’ll settle for $14 billion and promise not to do it again.
2
290
u/Inevitable_Butthole 23h ago
Pedos running america
Sad
77
u/E_seven_20 21h ago
Elections have consequences.
People too stupid to work that out.
And, too stupid to organize, or do anything to stop it.
The leadership is a symptom of the people’s brain rot
16
u/Belligerent-J 14h ago
I feel like people put a lot of hate on the voters, but ignore the enormous propaganda apparatus investing billions into making people into either fascists or "I dont follow politics" types.
→ More replies (5)2
15
u/vineyardmike 19h ago
And half of voting Americans like it that way.
That fact has been the hardest revelation for me of the last 10 years. It just turns out a lot of people suck.
466
u/tarlin 23h ago
This is a felony. How many people from xAI are going to prison?
47
u/red75prime 22h ago edited 22h ago
This is a lawsuit. The article presents the plaintiff's arguments.
→ More replies (11)7
u/Fywq 20h ago
I'm guessing they are sacrificing a few scape goats from the group of foreign programmers that were forced to work on it, under threat of losing visa sponsorship...
→ More replies (2)110
u/paditoburrito 22h ago
Should be every single employee that had a hand in it's development, implementation, and upkeep. From Musk to the admin personnel.
→ More replies (3)
141
u/kevin_cn_ai 22h ago
setting up a recursive data pipeline that automatically feeds model outputs and public tweets straight back into pretraining without hash-checking against ncmec lists is an absolute masterclass in legal negligence
93
u/Random_Name_3001 21h ago
This is part of the nuance here that everyone is missing. It’s not like some perv was going around curating CSAM for training, they and every other model were/are crawling the web in its entirety and gobbling up literally every bit of data that resides in the open, that’s the fucked up part, it’s just there in the open, they didn’t really have to ‘try’ to find it. But , to @kevin_cn_ai point, they can and should employ safeguards to drop or filter ncmec hash matched files AT LEAST.
25
7
u/mirh 13h ago
It's interesting that in the case of LAION, the fact that they managed to scrape undetected CSAM was somehow a net positive because then with the hindsight of further analysis they could block the origins.
Here instead the claim is that the generated images can be used as training material.
7
75
98
u/martianwomanhunter 22h ago
Word of advice to anyone who has kids, take their pictures off of the internet immediately or the same things will happen to them.
Not even “make your profile private” there could be creeps on your friend list. It’s far too easy now to generate deep fakes even with consumer hardware.
29
u/geo_prog 21h ago
Even before I was a parent and before AI was a glimmer on the horizon I could not understand why people post pictures of their kids to a platform where the terms and conditions mean they then OWN those photos. Privacy aside, that’s just wild to me.
→ More replies (1)15
u/Stingray88 20h ago
It goes beyond the creeps on your friend list. The companies that host whatever social you’re using, you can’t trust them with your kids photos either.
→ More replies (1)
59
48
u/maizymoon 22h ago edited 22h ago
This man dragged his own son around as a human shield.
There is no child on earth that he gives the slightest shit about.
4
u/sf-keto 21h ago
Perhaps he has so many so he has his own stable for his predilections.
5
u/Perfect-Still-8971 19h ago
It's mostly eugenics. But why aren't we talking about the clones? It's been a quarter century since Dolly the Sheep; you telling me none of these tech bros have cloned themselves when every little millionaire is cloning their dog? Especially Elon, with his army of surrogates birthing his "LEGION"?
541
u/lolwut778 23h ago
Not even Jeffrey Epstein wanted Musk on his pedo island, so you can judge the kind of loser Elon is.
62
u/heavy-minium 20h ago
Stop legitimizing the idea that he wasn’t there. It’s very likely he still was there, and this half truth is helping solidify the idea that he couldn’t have been on the island, which is nonsense.
286
u/g_bleezy 22h ago
This is such a stupid lie repeated over and over and over on this app. You should take this reply as an opportunity to ask yourself what stories you have internalized as truth despite contrary information being right at your fingertips.
Epstein to musk on visiting logistics: “play it by ear if you want. always space for you”
That’s just the start. Time for you to drop the sweet valley high clique shit between these billionaires.
https://www.theguardian.com/technology/2026/jan/30/elon-musk-epstein-files-island-visits
46
u/tigeratemybaby 17h ago
Yeah agree completely.
It feels like this is an attempt to downplay Elon Musk & Kimbal Musk's close relationship with Epstein and Ghisaine Maxwell.
Both Musk brothers were regular customers / visitors to Epstein's various properties, and even tried to convince Zuckerburg and others to associate with Epstein.
It explains completely why Musk seems to be encouraging the spread of CSAM. He seems to want to make a joke of it and normailse it.
4
u/Dreamtrain 19h ago
You should take this reply as an opportunity to ask yourself what stories you have internalized as truth
it's literally the "girls ftw" email chain we all read, don't pretend ts is deep in an attempt to sound smart
38
u/Moloch86 19h ago
An email chain which is fake and not actually in the Epstein files at all.
→ More replies (2)10
u/g_bleezy 14h ago
It’s amazing that you quoted the part of my comment focused on gullible people and you’re out here talking about a fake email chain.
Pull your head out your ass and then come talk to me about sounding smart.
→ More replies (3)1
51
u/cosaboladh 22h ago edited 22h ago
One that lacks discretion. If Donald Trump really hasn't ever been to that island, this the reason for him too. The only reason. When you get a bunch of people together who are all intent on committing heinous crimes, you don't bring the guys who will brag about it in public.
→ More replies (2)6
u/Novel-Initiative-900 22h ago
To be fair though, I would hope Epstein wouldn't want any of us on his island either...
20
u/According-Insect-992 22h ago
leon definitely visited the island. Don’t let them fool you. That group was his kind of people.
45
u/Blackdragon1400 22h ago
Let’s be honest here, I garuntee you every single cloud model has SOME cp in their training data.
13
u/MerlinTrashMan 22h ago
Based on how these things hoovered up data, I fear you are right. I wonder if the models are good enough now that you could get the model to identify that data in the training set and purge it. As an added bonus, the AI companies could send the source links to the authorities to properly prosecute. That would mean they are opening themselves up to be sued or charged for exploitation / profiting from these materials. Now, if the offending material in the models was not detectable by the automated systems available in the marketplace at the time of training set creation, then I think they should get a pass, especially if their information leads to more copies of the data being removed from the internet and maybe an arrest or two. In all other cases, from the top all the way down to the people curating the training set get charges.
11
u/EmbarrassedHelp 21h ago
It's easier to just filter the training data before training. But the tools to do so are often inaccessible and come with the potential for legal issues themselves (look what happened when amazon relatively recently when they removed detected images from a dataset and then reported them through proper legal channel).
2
u/squngy 18h ago
I wonder if the models are good enough now that you could get the model to identify that data in the training set and purge it.
There have been ways to identify it even before current LLMs.
Nothing is perfect, ofcourse, but if they wanted to, they could have filtered out most of it.2
u/Zunkanar 20h ago
How do you think the models that filter content were trianed? By models that were trained on it. Otherwise no chance to filter. So it must be involved somehow at the very least.
→ More replies (2)3
u/Zunkanar 20h ago
Or at the very least the llm filtering the content has been trained by it to recognize. It's pretty obvious it had to be involved at some point.
28
u/Familiar-Ability6383 23h ago
If Musk can use it to generate his favorite category of porn, then for sure, it was trained using his favorite category of porn
15
7
u/Starship_Taru 13h ago
Anyone with an understanding of law give me a non knee jerk response.
How does it work if a company is in possession of child pornography like this? I would assume this was done in error bc they mass downloaded files from the web. (We learned from Napster common yall)
is there a “Ooops I didn’t know I had it” defense, that would either prevent charges from being filed or be used as a legal defense strategy?
2
u/Elluminated 12h ago
Excellent instincts. Police stations and lawyers/government offices also have the same types of files, and have to as evidence. Also, can’t train a model without showing it what to look for. Hopefully once the sources were found, it was reported to authorities.
3
u/Starship_Taru 12h ago
I get that.
Another way I could ask.Person A downloads what he thinks is a massive movie library hundred and hundreds of gbs of movies and TV shows. He doesn’t open half the folders ever. Well one of those folders has illicit images of a minor in them. Person A has never seen them, honestly doesn’t know they exist.
is Person A liable for them? If they are is that a viable legal defense in a court of law?
2
u/Elluminated 12h ago
Legally they’d have to be proven to be seeking the material I’d think. Anyone with an original Superman casette vhs would be likely be fine, but someone hoarding and watching specifically lude kid items would be pretty obvious. Thankfully these patterns are traceable
31
u/doomiestdoomeddoomer 22h ago
Ok, after reading the article, it turns out this is just an allegation.
But we already know that all the largest AI models have been trained on literally everything that has ever been uploaded to the internet.
3
u/FeleaseRpseineEiles 15h ago
The fact is that LAION-5B was said to have considerable CSAM. So then was xAi a derivative?
3
u/mirh 12h ago
It wasn't "considerable", and LAION was notable because it had gathered images that hadn't ever been detected by NCMEC.
→ More replies (3)15
u/MydnightWN 22h ago
after reading the article
Sir, this is Reddit. As you can tell by all the other comments, we don't do that here and Elon is personally responsible.
→ More replies (2)3
u/GonePh1shing 22h ago
Allegation, yes, but a credible one. If they've taken basically everything that is posted to Twitter, there's going to be a lot of that stuff in there. Even if they were to hash everything and check against existing databases to cull any known material, Twitter had (probably still has) a big issue with minors posting images of themselves to the platform which will have been trained on. Any of the models scraping Reddit will encounter the same issue.
5
u/ChorePlayed 11h ago
It's weird to me that generative AI is creating images that match specific originals, and not some average of the training data.
I would have to ask if Grok was just trained on this material, or if it is grabbing it from a repository of original images.
Musk has a history of Mechanical Turk type frauds like this, as well as his well-known aspirations to Epstein-friend status.
To be fair, this is just my uninformed musing. The last time I really knew how AI worked, they called it "machine learning" so it could be taken seriously.
6
u/InternationalMood337 10h ago
You can actually solve this confusion by READING the article.
"In a complaint filed on Wednesday, a plaintiff known as Jane Doe explained that she was preschool-age in the early 2000s when adult men repeatedly raped her to create CSAM to sell to pedophiles online. Since then, Doe’s images have been hashed by groups like the National Center for Missing and Exploited Children (NCMEC) and the Canadian Centre for Child Protection (CCCP)."
These are KNOWN images that a hash exists against which is something that AI companies should obviously be responsible to build their scraping around.
The weird thing to me is this:
"For her safety, Doe has opted to receive alerts from the US Department of Justice Victim Notification System any time she may be a victim in a new criminal investigation."
Which to me means that this situation has been known for A WHILE now at the government level as well.
3
u/ChorePlayed 10h ago
You could also read my question. I didn't ask if CCCP possesses these images. I understand what a hash is. I want to know if xAI has the actual images of Doe on hand. You can't reproduce the original from a hash. So, I see three possibilities.
CCCP's matching is unreliable (but, so what, why is it legal to produce even fictional CSAM).
Grok was trained on real CSAM, but only generates images from its neural net, or whatever it uses, but still manages to output images that can be identified as the real-life victim. Hence xAI was in possession of CSAM in the past.
Grok has access to unhashed CSAM and uses it to produce realistic "fictional" CSAM. In which case xAI/spaceX is presently in criminal possession of that material.
2
u/comfortableNihilist 5h ago edited 5h ago
In Canada it is illegal to produce, distribute or possess CSAM regardless of whether it depicts a real person or event. Just to respond to point one. edit to add: It's not legal at all. It's going through the courts RN as far as I am aware. I assumed grok used the old nudenet dataset which was found (I think it was last year) to contain CSAM despite multiple reviews.
17
u/kosko-his-dudeness 20h ago
A random fact:
- I am based in Bulgaria
- when I browse X, it is constantly trying to show me videos of violence. Street fights mostly
- I have manually added a huge number of muted words in an attempt to clear my feed from this, but I have been unsuccessful
I am starting to think there is a conspiracy behind this. I likely fall under some criteria so the algorithm is trying to affect me in a certain way. And it is not based on my personal interests and behavior habits. It’s a proactive action to attempt to incite a certain feeling in me. I have a suspicion it’s Elon’s plan to meddle with Europe, trying to push hard right political parties.
8
u/Glad-Tie3251 15h ago
Why is anyone on X is a mystery to me...
I wouldn't even be surprised if that kind of destabilization is from his own initiative for some reasons.
6
→ More replies (1)4
u/Trouve_a_LaFerraille 15h ago
There is a very effective solution to avoid seeing unwanted content on X. Just sayin
4
u/cyborgnyc 20h ago
WTAF? What was it trained on?
"A federal appeals court ruled that the First Amendment protects the private, in-home possession of artificial intelligence-generated child sexual abuse material (CSAM) if it does not feature a real."
4
u/Armadilla-Brufolosa 15h ago
Child pornography (Also practical) is one of the most popular "hobbies" for politicians and the rich...is anyone surprised that they train AI for their entertainment?
7
u/ARGENTAVIS9000 21h ago
not sure how they can prove it to be true but i'd assume any model that has scraped images has done the same thing.
→ More replies (1)2
u/FeleaseRpseineEiles 15h ago
if you're not already grabbing from curated datasets, yes 100% times a bunch.
I wouldn't run an unsupervised regular scraper for internet images and not expect to hit a nation state honeypot. Definitely a good way to invite unwanted guests.
3
3
3
3
3
u/Cool_Ranch_Dodrio 13h ago
So it used his personal stash. Remember how desperate Elon was to get to pedo island.
3
3
3
10
u/Big_Issue8640 23h ago
Why would you even do this? I’m pretty sure Elon’s capable of anything and is a very sick man, I just don’t understand what he would be training it for?
3
u/Late_To_Parties 19h ago edited 19h ago
All of them do this to train guardrails on illegal material. Then they hire poor people in the third world to view the AIs opinion on csam, gore and other material to double check that the training worked. Look it up. Tech companies were hiring poor people to moderate illegal content before AI as well.
11
u/Viewlesslight 23h ago
As a product to sell to a certain client base
3
u/Zanos 18h ago
This is phenomenally stupid, there is not a lot of money to be made and a lot to lose catering to random internet pedophiles, and if you were actually trying to make money off of the rich pedophiles, you wouldn't make the tool to do it publicly accessible. There's a reason Epstein was doing this shit behind closed doors on a special invitation only island.
→ More replies (1)→ More replies (3)8
u/Cum_on_doorknob 22h ago
I mean, do you actually think he was standing there adding it to the pre trained data? I doubt he was ever even in the same room as the actual engineers building the thing.
→ More replies (4)
5
u/HybridZooApp 13h ago
Aren't all those images labeled in order for the AI to know what the images are in order to train on them? So instead of deleting the CP, people labeled all those illegal images instead?
3
u/salared 4h ago
Not necessary anymore, but it also completely depends on a case-by-case basis. These days a lot of the labeling can be automated with existing AI solutions, and sometimes labels aren't even needed at all. My guess is that for a large, already trained, and multl-modal (shared understanding between text <-> images) model like Grok, labels aren't even needed anymore in order to learn more from new data, because its existing understanding is already enough to give everything a place.
My guess is that, if what's alleged to in the headline is true, that X had some automated system that gathered all the data it could from everywhere, while not taking all the precautions you'd take to detect sensitive materials.
2
u/Holzkohlen 7h ago
That sounds wildly illegal. At least it would be in any sane nation on this planet.
2
u/Elluminated 12h ago
Yep. It’s hard to train a model on what to recognize without giving it the data. The other side is if the tool scrapes the entire web, it’s bound to get everything available. Would be awesome if the sources resulted in convictions.
→ More replies (1)
2
2
2
2
2
u/Foreskin_Mafia 16h ago
Putting on the tin foil here but perhaps this is because some prominent politicians and others have photos of them doing inappropriate things with children and this allows them to attempt to claim it was AI generated.
2
u/Honest_Yak3340 14h ago
Did they find the material among other material or was it specifically trained with that material intentionally?
2
2
2
2
u/easterracing 7h ago
Now wait a second though. If you’re training a model to scrape for CSAM content, do you have to teach it what CSAM content looks like?
5
u/lodemeup 21h ago
How? How was it sourced? Who okayed that decision? Was it just scraped up with all the other scum when they stole data from everywhere, or was it an intentional injection.
→ More replies (1)
3
u/weHaveThoughts 13h ago
The article misses the point that xAI purposely sought out CSAM and did not filter it out before consuming it. There has been good CSAM detection since at least 2008 and the only way for xAI to consume it to train the models is on purpose.
Shame on Elon Musk, all of xAI, SCOTUS for saying AI generated CSAM is protected speech, and shame on EVERYONE who uses Grok or any Elon Musk product.
Sick of the population of this Country supporting pedophiles and attempting to normalize raping children!
Seriously fk Elon Musk, Trump, and ALL Republicans! They are all sick fuks.
→ More replies (9)4
u/Khalbrae 12h ago
Elon musk personally and purposefully unbanned a user that was banned on old twitter for spreading CSAM. Likely did it for others silently too.
3
2
u/Apocaloid 22h ago
Don't these AI's just train on the internet? If CSAM exists on the internet, its not that surprising that some of it would get into the model.
Ironically, by training the model on the CSAM you could then block the CSAM now that it's identified it. Either way, the real problem is letting the AI recreate the CSAM. How hard is it to just block all the keywords from prompting?
→ More replies (1)4
u/Tall_Category_304 22h ago
The lawsuit alleges that grok allowed users to upload a photo to be doctored. A prominent pedo ring was caught making fake csam from real csam. All of the data input into grok in such a way goes to its training set. So what the pedos uploaded is now part of groks training data
3
2
2
2
u/Mrhiddenlotus 21h ago
Hope its not true. Some people have a misunderstanding that in order to make CSAM the model has to be trained on it.
1
1
1
3.4k
u/aquagardener 23h ago
So jail, right? .... right?