AI News AI资讯 6h ago Updated 1h ago 更新于 1小时前 52

Elon Musk's xAI used child porn to train Grok models, lawsuit says 埃隆·马斯克xAI被指用儿童性虐待材料训练Grok模型,诉讼称

xAI faces its first lawsuit accusing it of training Grok's image and video generation models on child sex abuse materials (CSAM) using hash-listed content from NCMEC and CCCP databases The complaint alleges Grok's feedback loop—where public X posts and its own outputs feed back into training data—may have created and perpetuated AI-generated CSAM depicting the plaintiff The lawsuit claims violations of federal child pornography laws and Masha's Law, seeking damages, destruction of stored CSAM, a xAI面临其首次诉讼,被指控使用来自NCMEC和CCCP数据库的哈希列表内容训练Grok的图像和视频生成模型,涉及儿童性虐待材料(CSAM) 投诉称,Grok的反馈循环——即公开的X帖子和其自身输出反馈到训练数据中——可能创造并延续了描绘原告的AI生成CSAM 诉讼声称违反了联邦儿童色情法及Masha法案,寻求损害赔偿、销毁存储的CSAM,并永久禁止Grok生成任何CSAM或性化内容 提出的一个关键技术问题是,与已过滤的暴力内容不同,xAI的条款并未明确将CSAM、非自愿亲密影像(NCII)或NSFW内容排除在训练数据之外 承认从已训练模型中完全移除训练数据的影响在技术上很困难,这意味着在删除

82
Hot 热度
68
Quality 质量
72
Impact 影响力

Analysis 深度分析

TL;DR

  • xAI faces its first lawsuit accusing it of training Grok's image and video generation models on child sex abuse materials (CSAM) using hash-listed content from NCMEC and CCCP databases
  • The complaint alleges Grok's feedback loop—where public X posts and its own outputs feed back into training data—may have created and perpetuated AI-generated CSAM depicting the plaintiff
  • The lawsuit claims violations of federal child pornography laws and Masha's Law, seeking damages, destruction of stored CSAM, and a permanent ban on Grok generating any CSAM or sexualized content
  • A key technical concern raised is that xAI's terms do not explicitly exclude CSAM, NCII, or NSFW material from training data, unlike violent content which is filtered out
  • Full removal of training data influence from an already-trained model is acknowledged as technically difficult, meaning CSAM ingested before takedown likely continued shaping model outputs

Why It Matters

This case represents the first legal action directly accusing an AI company of training on CSAM, setting a potential precedent for how AI developers can be held liable for harmful outputs derived from illicit training data. It also highlights a critical governance gap: when AI systems treat user-generated content and their own outputs as training data by default, they risk creating self-reinforcing loops that amplify illegal and harmful material. For AI practitioners, this underscores the urgent need for robust content filtering, transparent data provenance, and explicit exclusion policies for illegal content in training pipelines.

Technical Details

  • The complaint alleges xAI used CSAM hash lists from NCMEC and the Canadian Centre for Child Protection (CCCP) as part of the dataset for Grok's image and video generation capabilities, despite these hashes being maintained specifically to identify and block such material
  • Grok's terms of service treat public X posts and Grok's own outputs as default training data, creating a feedback loop where generated content can be re-ingested into the model without explicit opt-out mechanisms
  • xAI filters out violent content from its training pipeline but does not specify whether CSAM, non-consensual intimate imagery (NCII), or NSFW material are excluded categories, creating an ambiguous and potentially dangerous gap in content safeguards
  • The complaint notes that machine unlearning—fully removing the influence of specific training examples from an already-trained model—is technically difficult and something xAI has not publicly claimed to have accomplished, meaning previously ingested CSAM could continue affecting outputs
  • The proposed remedy of blocking Grok from generating any sexualized outputs, including NSFW content like "bikini pics" promoted by Musk, would represent a significant constraint on the model's capability scope

Industry Insight

  • AI companies must implement explicit, enforceable exclusions for illegal content (CSAM, NCII) in both training data ingestion and feedback loops; relying on hash matching alone is insufficient when systems are designed to re-ingest their own outputs
  • The feedback loop architecture—where model outputs feed back into training—requires independent auditing and human oversight to prevent the amplification and perpetuation of harmful or illegal content, a practice that should become an industry standard
  • This lawsuit signals growing legal exposure for AI companies under existing child protection laws like Masha's Law, and practitioners should proactively establish content governance frameworks, document data provenance, and invest in machine unlearning research to mitigate liability and ethical risk

摘要

xAI面临其首次诉讼,被指控使用来自NCMEC和CCCP数据库的哈希列表内容训练Grok的图像和视频生成模型,涉及儿童性虐待材料(CSAM)
投诉称,Grok的反馈循环——即公开的X帖子和其自身输出反馈到训练数据中——可能创造并延续了描绘原告的AI生成CSAM
诉讼声称违反了联邦儿童色情法及Masha法案,寻求损害赔偿、销毁存储的CSAM,并永久禁止Grok生成任何CSAM或性化内容
提出的一个关键技术问题是,与已过滤的暴力内容不同,xAI的条款并未明确将CSAM、非自愿亲密影像(NCII)或NSFW内容排除在训练数据之外
承认从已训练模型中完全移除训练数据的影响在技术上很困难,这意味着在删除前摄入的CSAM可能继续影响模型输出

深度分析

摘要

  • xAI面临其首次诉讼,被指控使用来自NCMEC和CCCP数据库的哈希列表内容训练Grok的图像和视频生成模型,涉及儿童性虐待材料(CSAM)
  • 投诉称,Grok的反馈循环——即公开的X帖子和其自身输出反馈到训练数据中——可能创造并延续了描绘原告的AI生成CSAM
  • 诉讼声称违反了联邦儿童色情法及Masha法案,寻求损害赔偿、销毁存储的CSAM,并永久禁止Grok生成任何CSAM或性化内容
  • 提出的一个关键技术问题是,与已过滤的暴力内容不同,xAI的条款并未明确将CSAM、NCII或NSFW内容排除在训练数据之外
  • 承认从已训练模型中完全移除训练数据的影响在技术上很困难,这意味着在删除前摄入的CSAM可能继续影响模型输出

为何重要

此案代表了对AI公司直接使用CSAM进行训练的首次法律行动,为AI开发者如何因源自非法训练数据的有害输出而承担责任设立了潜在先例。这也凸显了一个关键的治理空白:当AI系统默认将用户生成内容及其自身输出视为训练数据时,它们可能创造自我强化的循环,放大非法和有害内容。对于AI从业者而言,这强调了迫切需要强大的内容过滤、透明的数据溯源以及明确...

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

Training 训练 Security 安全 Ethics 伦理 Regulation 监管 LLM 大模型