1/3 web pages published since ChatGPT's launch show signs of AI authorship
Over one-third (35%) of web pages published after ChatGPT's November 2022 release show signs of AI authorship or substantial AI editing, according to a Pew Research study A random sample of 10,000 pages from July 2026 showed ~10% AI authorship overall, but this rose to 35% when pre-ChatGPT pages were filtered out .com domains exhibited AI authorship at roughly 10x the rate of .edu or .gov domains (~1% each), while .org domains sat at 4.6% The study used Common Crawl's web archive (nearly 500,000
Analysis
TL;DR
- Over one-third (35%) of web pages published after ChatGPT's November 2022 release show signs of AI authorship or substantial AI editing, according to a Pew Research study
- A random sample of 10,000 pages from July 2026 showed ~10% AI authorship overall, but this rose to 35% when pre-ChatGPT pages were filtered out
- .com domains exhibited AI authorship at roughly 10x the rate of .edu or .gov domains (~1% each), while .org domains sat at 4.6%
- The study used Common Crawl's web archive (nearly 500,000 English-language pages) and Open Pangram's AI-detection technology
- Secondary AI-writing indicators—em dashes, Oxford commas, and phrases like "it's not X, it's Y"—have also increased in prevalence over recent years
Why It Matters
This study provides one of the largest-scale empirical snapshots of AI-generated content on the open web, confirming that AI authorship has moved from a niche phenomenon to a dominant force in web publishing. For AI practitioners and researchers, it underscores the urgent need for robust detection, attribution, and content-authentication mechanisms as the internet becomes increasingly populated by machine-generated text.
Technical Details
- Data source: Nearly 500,000 English-language web pages collected from the Common Crawl web archive, spanning roughly five years beginning a couple of years before ChatGPT's November 2022 release
- Detection method: Open Pangram's AI-detection technology was used to classify pages as likely AI-written or substantially AI-edited; the study acknowledges that AI-detection tools can misclassify human-written content but argues the directional signal at scale is reliable
- Sampling methodology: A random sample of 10,000 pages from July 2026 yielded ~10% AI authorship; after filtering out pages published before ChatGPT, the rate jumped to 35%, demonstrating the importance of temporal normalization
- Domain-level analysis: .com domains showed ~10x the AI authorship rate of .edu and .gov domains (~1% each); .org domains registered 4.6%, suggesting commercial incentives drive the highest volume of AI-generated web content
- Linguistic markers: The study also tracked stylistic features associated with AI writing—em dashes, Oxford commas, and formulaic phrasing patterns—which have increased in frequency over the same period
Industry Insight
- The 35% AI authorship rate on post-ChatGPT web pages signals a fundamental shift in content ecosystems; platforms, search engines, and publishers should invest in content provenance standards (e.g., C2PA, watermarking) to maintain trust and differentiate human-created from machine-generated material
- The stark domain-level disparity (.com vs. .edu/.gov) suggests regulatory and academic spaces may remain relatively AI-clean for now, but commercial content farms and SEO-driven sites are prime candidates for AI saturation—organizations should audit their content supply chains accordingly
- As bot traffic has already overtaken human traffic (per Cloudflare), and a growing share of that content is AI-generated, the internet risks entering a feedback loop of bots reading bots; proactive investment in authentication, detection, and human-verified content channels will be critical to preserving information integrity
Disclaimer: The above content is generated by AI and is for reference only.