AI News AI资讯 3d ago Updated 3d ago 更新于 3天前 46

We still don't know how people are really using AI 我们仍然不知道人们究竟如何使用AI

The AI Observatory is a new independent research platform aggregating real AI conversations from seven datasets to provide unbiased insights into how people actually use generative AI models Major AI companies' usage reports (e.g., Anthropic Economic Index) filter out nearly half of conversations by focusing only on work-related uses, missing significant personal, sensitive, and harmful interactions The AI Observatory found that 48% of conversations would be filtered out by Anthropic's methodolo AI Observatory项目由斯坦福和MIT研究者联合创建,聚合了2023-2025年间5000名用户与52个AI模型的24,521次真实对话,提供独立于厂商的AI使用数据 主要AI公司(Anthropic/OpenAI)的报告存在选择性偏差,Anthropic Economic Index过滤了48%的非工作相关对话,而这些被过滤内容中敏感话题(健康/关系44.2%、成人内容7.9%、仇恨骚扰27.5%、性内容16.7%)比例远高于其官方数据 不同模型呈现差异化使用模式:Grok/Gemini用于信息检索(Grok集中了更多虚假信息),Claude用于编程,Gemini用于社交/角色扮演

65
Hot 热度
70
Quality 质量
60
Impact 影响力

Analysis 深度分析

TL;DR

  • The AI Observatory is a new independent research platform aggregating real AI conversations from seven datasets to provide unbiased insights into how people actually use generative AI models
  • Major AI companies' usage reports (e.g., Anthropic Economic Index) filter out nearly half of conversations by focusing only on work-related uses, missing significant personal, sensitive, and harmful interactions
  • The AI Observatory found that 48% of conversations would be filtered out by Anthropic's methodology, with non-work conversations showing higher rates of health/relationship topics (44.2%), adult content (7.9%), harassment/hate (27.5%), and sexual content (16.7%)
  • Usage patterns vary significantly across models: Grok and Gemini are used more for information retrieval, Anthropic for coding, Gemini for social/roleplay, and ChatGPT for homework assistance
  • Conversations became longer and more elaborate over time (2023-2025), with increasing small talk suggesting growing AI companionship, while sensitive exchanges decreased, possibly indicating improved platform safeguards

Why It Matters

This research exposes critical blind spots in how AI usage data is collected and reported by major companies, revealing that highly consequential policy and safety decisions are being made on incomplete and selectively filtered datasets. The AI Observatory provides the first independent, comprehensive view of real-world AI interactions across multiple models, enabling researchers and policymakers to better understand both the benefits and risks of generative AI adoption.

Technical Details

  • The AI Observatory aggregated 85,633 conversational turns across 24,521 conversations from seven real-world datasets, involving 5,000 users interacting with 52 different models including ChatGPT, Gemini, Claude, and Grok between 2023 and 2025
  • Researchers applied Anthropic's Economic Index filtering methodology to their dataset and found that 48% of conversations would have been excluded, revealing significant differences in sensitive content categories compared to company reports
  • The study analyzed conversation length (prompt tokens, response tokens, conversation turns), interaction styles, topic distributions, and sensitive content classification across models and over time
  • Datasets were collected with user consent through existing research projects, though researchers acknowledge this voluntary sourcing likely underrepresents sensitive uses due to privacy concerns
  • The research was co-led by Anka Reuel (Stanford STAIR Lab) and Shayne Longpre (MIT Media Lab), with collaborators from multiple institutions

Industry Insight

  • AI companies should consider publishing more comprehensive, less filtered usage data to enable independent verification and build public trust, as selective reporting creates information asymmetries that hinder effective oversight
  • Policymakers and researchers should not rely solely on company-published reports for AI safety assessments; independent datasets like the AI Observatory are essential for understanding real-world usage patterns and risks
  • The finding that sensitive exchanges decreased over time while AI companionship increased suggests companies should invest in both improved safety guardrails and better understanding of emotional attachment dynamics in human-AI interactions

TL;DR

  • AI Observatory项目由斯坦福和MIT研究者联合创建,聚合了2023-2025年间5000名用户与52个AI模型的24,521次真实对话,提供独立于厂商的AI使用数据
  • 主要AI公司(Anthropic/OpenAI)的报告存在选择性偏差,Anthropic Economic Index过滤了48%的非工作相关对话,而这些被过滤内容中敏感话题(健康/关系44.2%、成人内容7.9%、仇恨骚扰27.5%、性内容16.7%)比例远高于其官方数据
  • 不同模型呈现差异化使用模式:Grok/Gemini用于信息检索(Grok集中了更多虚假信息),Claude用于编程,Gemini用于社交/角色扮演,ChatGPT用于作业辅助
  • 对话趋势显示:用户与AI的互动越来越长、闲聊增多、AI自我披露减少、敏感内容下降,反映AI陪伴需求上升和安全措施改进
  • 研究局限性在于数据来自自愿提供,敏感使用可能被低估,且样本量(约2.5万对话)远小于厂商数据(Anthropic 100万、OpenAI 150万)

为什么值得看

本文为AI治理和研究领域提供了首个大规模独立AI使用数据源,揭示了厂商报告的"盲点",对政策制定者和研究者评估AI真实风险与收益具有重要参考价值。同时,研究揭示了不同模型的用户行为差异和演化趋势,为模型改进和安全策略制定提供了实证依据。

技术解析

  • 数据来源与规模:聚合7个现有研究数据集,包含85,633个对话轮次(prompt+response),覆盖24,521次独立对话、5,000名用户、52个模型(ChatGPT/Gemini/Claude/Grok等),时间跨度2023-2025年
  • 方法论创新:将Anthropic的筛选方法应用于独立数据集,发现48%对话因非工作相关被过滤,揭示了厂商报告的选择性偏差;通过对比分析识别不同模型的敏感内容比例差异
  • 趋势分析维度:从对话长度(token数/轮次)、话题类型(健康/关系/成人/仇恨/性)、互动模式(闲聊/自我披露)、模型响应特征等多维度追踪AI使用演化
  • 模型差异发现:GPT-3.5对话较短,GPT-4o对话更长且更具迭代性(与"情感依赖"现象吻合);Grok在新闻政治信息检索中虚假信息集中;不同模型版本间存在显著行为差异

行业启示

  • 数据透明度危机:厂商主导的AI使用报告存在系统性偏差,政策制定和风险评估依赖不完整数据,亟需建立独立、可验证的第三方监测机制
  • 模型差异化定位:不同模型已形成稳定的用户心智定位(Claude=编程、Gemini=社交、ChatGPT=作业、Grok=信息检索),产品策略应强化优势场景而非盲目扩展
  • 安全与体验的平衡:敏感内容使用频率下降反映安全措施见效,但自愿数据源可能低估真实风险;未来需在保护隐私前提下拓展数据采集渠道,建立更全面的AI使用图谱

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

Claude Claude Gemini Gemini LLM 大模型 Dataset 数据集 Research 科学研究