AI News AI资讯 4h ago Updated 1h ago 更新于 1小时前 49

Artists Built a Site to Escape AI. Scrapers Are Coming for It Anyway 艺术家建立网站逃离AI,但爬虫仍在逼近

Artist social network Cara, which explicitly prohibits AI data scraping in its terms of service, was scraped three separate times within ten days, with datasets appearing on Reddit, Hugging Face, and Academic Torrents Founder Jingna Zhang characterized the repeated intrusions as "targeted attacks" designed to inflict harm on artists and undermine their right to consent regarding how their work is used The conflict highlights a fundamental ideological divide: artists view consent as non-negotiabl 艺术家社交网络Cara在10天内遭遇3次数据抓取攻击,近1200万张受版权保护的图片及元数据被非法采集 抓取者将数据先后发布于Reddit、Hugging Face和Academic Torrents等平台,引发关于AI训练数据获取方式的激烈争议 Cara创始人Jingna Zhang将此次事件定性为"针对艺术家的定向攻击",并呼吁社区通过GoFundMe筹集法律辩护资金 Hugging Face以"未托管作品副本、仅存储URL"为由拒绝删除数据集,凸显平台责任与版权保护的边界模糊 首次抓取者事后删除数据并与Cara合作开发开源检测工具,展现冲突双方从对抗走向协作的可能性

72
Hot 热度
68
Quality 质量
70
Impact 影响力

Analysis 深度分析

TL;DR

  • Artist social network Cara, which explicitly prohibits AI data scraping in its terms of service, was scraped three separate times within ten days, with datasets appearing on Reddit, Hugging Face, and Academic Torrents
  • Founder Jingna Zhang characterized the repeated intrusions as "targeted attacks" designed to inflict harm on artists and undermine their right to consent regarding how their work is used
  • The conflict highlights a fundamental ideological divide: artists view consent as non-negotiable, while some scrapers justify data extraction as necessary for AI advancement, comparing it to "building a highway"
  • Cara operates as a volunteer-run, shoestring-budget public benefit corporation serving 1.5 million users, making it financially vulnerable to sustained legal and technical defense costs
  • The first scraper eventually took down his Reddit post and is now co-creating an open source tool with Cara to help artists detect if their work appears in new datasets

Why It Matters

This incident represents a microcosm of the broader, escalating conflict between AI developers who prioritize unrestricted data access and creators who demand consent and control over their intellectual property. For AI practitioners, it underscores the growing legal and ethical risks of scraping publicly available data without considering the explicit wishes of content creators, as well as the potential for reputational damage and platform liability.

Technical Details

  • Cara is a public benefit corporation and artist-focused social network launched in late 2022, with over 1.5 million users, that implements explicit terms of service prohibiting AI access and incorporates technological and legal countermeasures to protect artist consent
  • The first scrape involved approximately 12 million images and metadata posted to the subreddit r/DefendingAIArt by user "MandarinDrawnPoppy994," followed by a second scrape on Hugging Face by "Ioannis/Captive Dreamer," and a third on Academic Torrents on August 22
  • Hugging Face's Trust and Safety team declined to remove the second dataset, arguing that no actual copies of the artworks were hosted on their servers and that the URLs merely pointed to content on Cara, effectively treating the platform as a neutral infrastructure provider
  • The scraped data includes both copyrighted images and associated metadata, which together form a valuable training dataset for generative AI models
  • Cara is developing an open source detection tool in collaboration with the first scraper to help artists identify whether their work has been included in new AI datasets

Industry Insight

  • AI companies and developers should anticipate increasing legal exposure and reputational risk from scraping datasets that contain explicit opt-out terms, as courts and public opinion may increasingly recognize consent-based frameworks over blanket "fair use" arguments for training data
  • Platform intermediaries like Hugging Face face growing pressure to establish clearer content policies around scraped datasets, as their current "neutral infrastructure" stance is being challenged by creators who view it as complicity in rights violations
  • The emergence of open source tools for detecting dataset inclusion signals a growing ecosystem of artist empowerment technologies, which AI developers should monitor closely as these tools may become standard for compliance auditing and licensing verification

TL;DR

  • 艺术家社交网络Cara在10天内遭遇3次数据抓取攻击,近1200万张受版权保护的图片及元数据被非法采集
  • 抓取者将数据先后发布于Reddit、Hugging Face和Academic Torrents等平台,引发关于AI训练数据获取方式的激烈争议
  • Cara创始人Jingna Zhang将此次事件定性为"针对艺术家的定向攻击",并呼吁社区通过GoFundMe筹集法律辩护资金
  • Hugging Face以"未托管作品副本、仅存储URL"为由拒绝删除数据集,凸显平台责任与版权保护的边界模糊
  • 首次抓取者事后删除数据并与Cara合作开发开源检测工具,展现冲突双方从对抗走向协作的可能性

为什么值得看

本文揭示了AI产业快速发展背景下创作者权益保护的核心矛盾,为从业者提供了关于数据伦理、平台责任和法律边界的现实案例。Cara事件反映了当前AI训练数据获取模式的系统性缺陷,对内容平台、AI开发者和政策制定者均具有重要参考价值。

技术解析

  • Cara平台采用明确的服务条款禁止AI访问,并部署技术与法律双重防护措施,但作为志愿者运营的免费项目,缺乏应对大规模数据抓取的资源和能力
  • 抓取行为通过自动化爬虫技术实现,将1200万张图片及其元数据打包为数据集,分别上传至Reddit、Hugging Face和Academic Torrents等开放平台
  • Hugging Face的版权处理机制仅针对平台托管内容,对指向外部URL的元数据索引不承担删除义务,形成监管漏洞
  • 首次抓取者与Cara达成和解并合作开发开源检测工具,为创作者提供验证作品是否被纳入新数据集的技术方案

行业启示

  • AI数据获取的"先抓取后争议"模式正面临越来越强的法律与伦理挑战,企业需重新评估训练数据来源的合规风险
  • 平台责任边界模糊化趋势明显,Hugging Face等开源社区的版权处理标准可能成为行业争议的焦点
  • 创作者权益保护与AI技术发展之间的张力将持续升级,推动"同意优先"的数据伦理框架和配套技术工具的需求增长

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

Creative AI 创意AI Dataset 数据集 Security 安全 Ethics 伦理 Policy 政策