"Google and Reddit do not own the Internet," web scraper says after court win
Google sued SerpApi for circumventing anti-scraping technology and selling unauthorized access to search results via a "Google Search API," invoking the Digital Millennium Copyright Act (DMCA). A court dismissed Google’s DMCA claim early, ruling that Google lacks standing because it does not own or license the content in search results. Reddit filed a similar lawsuit against SerpApi and Perplexity for scraping Reddit content appearing in Google search results, but its case faces similar legal hu
Analysis
TL;DR
- Google sued SerpApi for circumventing anti-scraping technology and selling unauthorized access to search results via a "Google Search API," invoking the Digital Millennium Copyright Act (DMCA).
- A court dismissed Google’s DMCA claim early, ruling that Google lacks standing because it does not own or license the content in search results.
- Reddit filed a similar lawsuit against SerpApi and Perplexity for scraping Reddit content appearing in Google search results, but its case faces similar legal hurdles due to lack of copyright ownership.
- Google plans to amend its complaint to focus on copyrighted content in “knowledge panels,” though this strategy risks exposing Google to potential infringement claims if it admits to using unlicensed material.
- Legal experts argue both companies are misusing the DMCA to control content they don’t own, potentially undermining open web principles.
Why It Matters
This case highlights the growing tension between AI-driven data scraping and intellectual property rights as large language models increasingly rely on publicly available web content. The outcome could set a precedent for how platforms like Google and Reddit legally respond to automated extraction tools, influencing future AI training practices and shaping the boundaries of fair use under copyright law.
Technical Details
- Google’s anti-scraping measures include technical barriers designed to prevent bots from accessing search results at scale, which SerpApi allegedly bypassed through reverse engineering or other circumvention methods.
- The DMCA prohibits trafficking in technologies that circumvent digital rights management (DRM) or access controls, but its application here is contested since Google does not claim ownership over most indexed content.
- “Knowledge panels” are algorithmically generated summaries about entities that may include licensed text, images, or data from third-party rights holders—this narrow category forms the basis of Google’s revised legal argument.
- SerpApi operates as a proxy service allowing developers to programmatically retrieve search engine results without triggering rate limits or detection systems, raising questions about whether such tools constitute legitimate intermediaries or infringing actors.
- Courts typically require plaintiffs to demonstrate direct harm caused by circumvention; Google failed to show it suffered quantifiable financial loss attributable specifically to SerpApi’s actions beyond general competition concerns.
Industry Insight
AI startups building models trained on scraped data should prepare for increased legal scrutiny as tech giants seek new ways to monetize or restrict access to their platforms’ outputs. Companies relying on public datasets must evaluate whether current interpretations of fair use will hold up under evolving case law, especially when those datasets originate from sites actively fighting back against automation. Meanwhile, infrastructure providers offering APIs or scraping services need to carefully assess liability exposure while balancing innovation with compliance risks in an uncertain regulatory landscape.
Disclaimer: The above content is generated by AI and is for reference only.