We still don't know how people are really using AI
The AI Observatory is a new independent research platform aggregating real AI conversations from seven datasets to provide unbiased insights into how people actually use generative AI models Major AI companies' usage reports (e.g., Anthropic Economic Index) filter out nearly half of conversations by focusing only on work-related uses, missing significant personal, sensitive, and harmful interactions The AI Observatory found that 48% of conversations would be filtered out by Anthropic's methodolo
Analysis
TL;DR
- The AI Observatory is a new independent research platform aggregating real AI conversations from seven datasets to provide unbiased insights into how people actually use generative AI models
- Major AI companies' usage reports (e.g., Anthropic Economic Index) filter out nearly half of conversations by focusing only on work-related uses, missing significant personal, sensitive, and harmful interactions
- The AI Observatory found that 48% of conversations would be filtered out by Anthropic's methodology, with non-work conversations showing higher rates of health/relationship topics (44.2%), adult content (7.9%), harassment/hate (27.5%), and sexual content (16.7%)
- Usage patterns vary significantly across models: Grok and Gemini are used more for information retrieval, Anthropic for coding, Gemini for social/roleplay, and ChatGPT for homework assistance
- Conversations became longer and more elaborate over time (2023-2025), with increasing small talk suggesting growing AI companionship, while sensitive exchanges decreased, possibly indicating improved platform safeguards
Why It Matters
This research exposes critical blind spots in how AI usage data is collected and reported by major companies, revealing that highly consequential policy and safety decisions are being made on incomplete and selectively filtered datasets. The AI Observatory provides the first independent, comprehensive view of real-world AI interactions across multiple models, enabling researchers and policymakers to better understand both the benefits and risks of generative AI adoption.
Technical Details
- The AI Observatory aggregated 85,633 conversational turns across 24,521 conversations from seven real-world datasets, involving 5,000 users interacting with 52 different models including ChatGPT, Gemini, Claude, and Grok between 2023 and 2025
- Researchers applied Anthropic's Economic Index filtering methodology to their dataset and found that 48% of conversations would have been excluded, revealing significant differences in sensitive content categories compared to company reports
- The study analyzed conversation length (prompt tokens, response tokens, conversation turns), interaction styles, topic distributions, and sensitive content classification across models and over time
- Datasets were collected with user consent through existing research projects, though researchers acknowledge this voluntary sourcing likely underrepresents sensitive uses due to privacy concerns
- The research was co-led by Anka Reuel (Stanford STAIR Lab) and Shayne Longpre (MIT Media Lab), with collaborators from multiple institutions
Industry Insight
- AI companies should consider publishing more comprehensive, less filtered usage data to enable independent verification and build public trust, as selective reporting creates information asymmetries that hinder effective oversight
- Policymakers and researchers should not rely solely on company-published reports for AI safety assessments; independent datasets like the AI Observatory are essential for understanding real-world usage patterns and risks
- The finding that sensitive exchanges decreased over time while AI companionship increased suggests companies should invest in both improved safety guardrails and better understanding of emotional attachment dynamics in human-AI interactions
Disclaimer: The above content is generated by AI and is for reference only.