Artists Built a Site to Escape AI. Scrapers Are Coming for It Anyway
Artist social network Cara, which explicitly prohibits AI data scraping in its terms of service, was scraped three separate times within ten days, with datasets appearing on Reddit, Hugging Face, and Academic Torrents Founder Jingna Zhang characterized the repeated intrusions as "targeted attacks" designed to inflict harm on artists and undermine their right to consent regarding how their work is used The conflict highlights a fundamental ideological divide: artists view consent as non-negotiabl
Analysis
TL;DR
- Artist social network Cara, which explicitly prohibits AI data scraping in its terms of service, was scraped three separate times within ten days, with datasets appearing on Reddit, Hugging Face, and Academic Torrents
- Founder Jingna Zhang characterized the repeated intrusions as "targeted attacks" designed to inflict harm on artists and undermine their right to consent regarding how their work is used
- The conflict highlights a fundamental ideological divide: artists view consent as non-negotiable, while some scrapers justify data extraction as necessary for AI advancement, comparing it to "building a highway"
- Cara operates as a volunteer-run, shoestring-budget public benefit corporation serving 1.5 million users, making it financially vulnerable to sustained legal and technical defense costs
- The first scraper eventually took down his Reddit post and is now co-creating an open source tool with Cara to help artists detect if their work appears in new datasets
Why It Matters
This incident represents a microcosm of the broader, escalating conflict between AI developers who prioritize unrestricted data access and creators who demand consent and control over their intellectual property. For AI practitioners, it underscores the growing legal and ethical risks of scraping publicly available data without considering the explicit wishes of content creators, as well as the potential for reputational damage and platform liability.
Technical Details
- Cara is a public benefit corporation and artist-focused social network launched in late 2022, with over 1.5 million users, that implements explicit terms of service prohibiting AI access and incorporates technological and legal countermeasures to protect artist consent
- The first scrape involved approximately 12 million images and metadata posted to the subreddit r/DefendingAIArt by user "MandarinDrawnPoppy994," followed by a second scrape on Hugging Face by "Ioannis/Captive Dreamer," and a third on Academic Torrents on August 22
- Hugging Face's Trust and Safety team declined to remove the second dataset, arguing that no actual copies of the artworks were hosted on their servers and that the URLs merely pointed to content on Cara, effectively treating the platform as a neutral infrastructure provider
- The scraped data includes both copyrighted images and associated metadata, which together form a valuable training dataset for generative AI models
- Cara is developing an open source detection tool in collaboration with the first scraper to help artists identify whether their work has been included in new AI datasets
Industry Insight
- AI companies and developers should anticipate increasing legal exposure and reputational risk from scraping datasets that contain explicit opt-out terms, as courts and public opinion may increasingly recognize consent-based frameworks over blanket "fair use" arguments for training data
- Platform intermediaries like Hugging Face face growing pressure to establish clearer content policies around scraped datasets, as their current "neutral infrastructure" stance is being challenged by creators who view it as complicity in rights violations
- The emergence of open source tools for detecting dataset inclusion signals a growing ecosystem of artist empowerment technologies, which AI developers should monitor closely as these tools may become standard for compliance auditing and licensing verification
Disclaimer: The above content is generated by AI and is for reference only.