AI News AI资讯 2h ago Updated 57m ago 更新于 57分钟前 49

Pew study confirms sharp rise of AI-written text on the web since ChatGPT's launch 皮尤研究证实ChatGPT发布以来网络上AI生成文本显著增加

Pew Research Center analyzed nearly 500,000 English-language web pages and found over a third of pages published after ChatGPT's launch show signs of AI-generated text Commercial .com domains contain AI-generated text roughly ten times more often than .edu or .gov sites (~10% vs ~1%) Characteristic AI language patterns have spiked: words like "delve," "tapestry," and "pivotal" more than doubled in frequency; em dash usage doubled and Oxford comma usage jumped 63% since 2023 Current AI detection Pew研究分析近50万英文网页,发现ChatGPT发布后超三分之一新网页显示AI生成文本迹象 商业.com域名AI文本出现频率是.edu/.gov网站的约10倍,揭示不同领域AI使用差异 AI偏好词汇(delve、interplay等)和标点(em dashes、Oxford commas)使用频率显著上升 当前AI检测工具难以区分完全自动化文本与部分AI辅助文本,定义模糊影响结论精确性 研究警示公众对AI文本负面影响的感知可能超出数据实际支持范围

68
Hot 热度
72
Quality 质量
70
Impact 影响力

Analysis 深度分析

TL;DR

  • Pew Research Center analyzed nearly 500,000 English-language web pages and found over a third of pages published after ChatGPT's launch show signs of AI-generated text
  • Commercial .com domains contain AI-generated text roughly ten times more often than .edu or .gov sites (~10% vs ~1%)
  • Characteristic AI language patterns have spiked: words like "delve," "tapestry," and "pivotal" more than doubled in frequency; em dash usage doubled and Oxford comma usage jumped 63% since 2023
  • Current AI detection tools like Open Pangram cannot reliably distinguish between fully automated text, AI-assisted writing, or partial AI involvement, limiting the precision of such studies
  • A separate April 2026 study by Imperial College London, Internet Archive, and Stanford found ~35% of newly published websites were fully or partly AI-generated, with 33% higher semantic similarity and more positive tone overall

Why It Matters

This research provides large-scale empirical evidence that AI-generated content has become a dominant force on the open web, fundamentally altering the landscape of online information. For AI practitioners and researchers, it highlights the urgent need for better detection methodologies and clearer definitions of what constitutes "AI text," as the current binary framework fails to capture the nuanced reality of human-AI collaboration in writing.

Technical Details

  • Dataset: Nearly 500,000 English-language web pages from the Common Crawl web archive, analyzed in a July 2026 sample
  • Detection tool: Open Pangram AI detection model used to identify signs of machine authorship across web pages
  • Domain breakdown: .com sites at ~10% AI-generated, .org at 4.6%, and .edu/.gov at ~1% each, revealing a tenfold commercial vs. institutional disparity
  • Linguistic markers tracked: Frequency analysis of AI-favored vocabulary ("delve," "interplay," "testament," "pivotal," "landscape," "tapestry," "bolstered," "crucial," "meticulous," "vibrant"), em dash usage (2x increase), Oxford comma usage (+63%), and "it's not just X, it's Y" negative parallelism patterns (nearly 3x increase)
  • Corroborating study: Imperial College London/Internet Archive/Stanford April 2026 research found 35% of new websites AI-generated, with 33% higher semantic similarity and more positive tone in AI texts

Industry Insight

  • Content platforms and search engines must develop more sophisticated detection and labeling systems that account for the spectrum of AI assistance rather than a binary human/machine classification, as partial AI use is likely the norm going forward
  • Commercial website operators face increasing reputational and trust risks as AI-generated content saturates .com domains; brands should establish clear disclosure policies to maintain audience credibility
  • The polarization around AI authorship stigma—documented in workplace studies and debates over watermarks like Anthropic's planned Claude watermark—suggests the industry needs nuanced frameworks that distinguish between unethical spam generation and legitimate AI-assisted productivity tools

TL;DR

  • Pew研究分析近50万英文网页,发现ChatGPT发布后超三分之一新网页显示AI生成文本迹象
  • 商业.com域名AI文本出现频率是.edu/.gov网站的约10倍,揭示不同领域AI使用差异
  • AI偏好词汇(delve、interplay等)和标点(em dashes、Oxford commas)使用频率显著上升
  • 当前AI检测工具难以区分完全自动化文本与部分AI辅助文本,定义模糊影响结论精确性
  • 研究警示公众对AI文本负面影响的感知可能超出数据实际支持范围

为什么值得看

该研究首次大规模量化了AI生成文本在公共网络中的渗透程度,为内容生态治理和检测技术发展提供了关键基准。同时揭示了商业网站与教育机构在AI使用上的显著差异,对制定差异化内容策略具有参考价值。

技术解析

  • 数据集:基于Common Crawl网络档案,分析近50万英文网页,时间跨度覆盖ChatGPT发布前后(2022年11月至2026年7月)
  • 检测工具:使用Open Pangram AI检测工具识别机器作者痕迹,但该工具仅能粗略判断人机归属,无法量化AI参与程度或阶段
  • 语言模式分析:统计了AI偏好词汇(delve、interplay、testament等10个高频词)和标点符号(em dashes使用量翻倍、Oxford commas增长63%)的使用频率变化
  • 域名分类对比:按.com/.org/.edu/.gov域名分类统计,.com域名AI文本比例约10%,.edu/.gov仅约1%,商业网站AI使用频率是教育政府网站的10倍
  • 研究局限性:承认检测工具无法区分完全自动化生成与部分AI辅助写作,定义模糊影响结论精确性,且公众感知可能夸大负面影响

行业启示

  • 内容平台需建立分层审核机制,针对商业网站与权威机构采取不同AI内容治理策略,避免"一刀切"
  • AI检测技术应朝着细粒度辅助程度评估方向发展,而非简单二元分类,以适配人机协作的现实工作流
  • 企业应关注AI写作工具带来的语言风格同质化风险,制定品牌内容差异化指南,保持独特语调与表达风格

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

Research 科学研究 LLM 大模型 Policy 政策 Ethics 伦理