Thomson Reuters bets $40M on owning its AI instead of renting from OpenAI or Anthropic
Thomson Reuters launched "Thomson," an in-house legal AI model built on Alibaba's open-source Qwen3.5-397B, investing approximately $40 million over two years in staff and compute The model underwent a three-stage process: safety/ethics retraining with Imperial College ("Snowdon"), pre-training on proprietary legal content, and agentic reinforcement learning within the company's own tool environments Thomson only edges out GPT-5.4 (0.83 vs 0.82) when granted access to exclusive Thomson Reuters c
Analysis
TL;DR
- Thomson Reuters launched "Thomson," an in-house legal AI model built on Alibaba's open-source Qwen3.5-397B, investing approximately $40 million over two years in staff and compute
- The model underwent a three-stage process: safety/ethics retraining with Imperial College ("Snowdon"), pre-training on proprietary legal content, and agentic reinforcement learning within the company's own tool environments
- Thomson only edges out GPT-5.4 (0.83 vs 0.82) when granted access to exclusive Thomson Reuters content; without it, it trails significantly (0.53 vs 0.65 on factual accuracy)
- The company built a "model factory" rather than focusing solely on the individual model, having cycled through approximately six different open-source foundation models during development
- Less than 10% of available proprietary content has been used in training so far, leaving substantial room for future performance gains
Why It Matters
Thomson Reuters' approach demonstrates that enterprises with exclusive data and domain expertise can build competitive AI systems without relying on frontier model providers, challenging the assumption that only well-funded labs can produce top-tier models. The case study is particularly relevant for professional services firms considering whether to build or buy AI capabilities, as it quantifies the trade-offs between ownership, cost, and performance.
Technical Details
- Foundation model: Built on Alibaba's Qwen3.5-397B, with the company cycling through roughly six different open-source starting points during development
- Training pipeline: Three-phase process — (1) safety, ethics, and political neutrality retraining with Imperial College ("Snowdon" intermediate model), (2) pre-training on proprietary content from Westlaw, Practical Law, Checkpoint, and Reuters, (3) post-training with domain experts and agentic reinforcement learning inside company tool environments
- Benchmark performance: On Stanford LegalBench, Thomson scored 0.823, trailing Gemini 3.1 Pro and GPT-5.5; on Harvey Legal Agent Benchmark, it sits just behind Opus 4.8; it leads on instruction following and PrBench Legal but falls sharply on reasoning and coding tasks
- Data scale: Less than 10% of available proprietary content used in training; the $40M figure covers staff and compute but excludes the value of decades of content and hundreds of domain expert hours
- Deployment: Initially deployed in CoCounsel Legal's Tabular Analysis feature for high-volume document review; a smaller open-weight version coming to Hugging Face under a non-commercial license
Industry Insight
- The "renting vs. buying" framework for AI adoption will increasingly define enterprise strategy: companies with proprietary data, domain experts, and measurable workflows can justify in-house models, while others should remain customers of frontier providers
- Data access matters as much as model quality — Thomson Reuters' razor-thin lead over GPT-5.4 came primarily from exclusive content access, suggesting that future competitive advantages will come from data moats and tool integration rather than raw model architecture
- The open-source community is closing the gap with frontier labs within months rather than years, as demonstrated by Qwen-based models achieving competitive results; enterprises should monitor open-source developments closely before committing to proprietary solutions
Disclaimer: The above content is generated by AI and is for reference only.