AIToday
Large Language ModelsOpen-Source AIAI Business & IndustryTHE DECODERPublished: Aug 24, 2026, 22:01 JST3 min read

Thomson Reuters builds own AI for $40M, skips OpenAI

Thomson Reuters builds own AI for $40M, skips OpenAI

Key takeaway

  • Thomson Reuters spent about $40 million to build its own AI model on Alibaba's Qwen.

  • The model trails frontier rivals on many benchmarks but leads on some legal tasks.

  • It only beats GPT 5.4 when using the company's own content.

3 Key Points

  1. What happened

    Thomson Reuters launched its first in-house language model, built on Alibaba's Qwen, after spending about $40 million on staff and computing over two years. The model, called Thomson, is now used in the Tabular Analysis feature of CoCounsel Legal.

  2. Why it matters

    The company chose to build its own model instead of fine-tuning a frontier model from OpenAI or Anthropic, citing cost, data control, and a compounding effect. An in-house model pays off on high-volume work like document review, and every expert review during updates becomes training data.

  3. What to watch

    On the company's own Deep Research benchmark, Thomson edges past GPT 5.4 only with access to company content (0.83 to 0.82). A small version is coming to Hugging Face as an open-weight model under a non-commercial license, and early talks with law firms about direct licensing are underway.

Ask the AI about this article →

Context & Analysis

Thomson Reuters' decision to build its own model rests on a specific economic logic: the real capital is not the $40 million or even the $450,000 final training run, but decades of content from Westlaw, Practical Law, Checkpoint, and Reuters, plus the working hours of hundreds of domain experts. The company's own benchmarks show why this matters. Thomson trails on reasoning and coding, but leads on instruction following and the PrBench Legal benchmark, and only surpasses GPT 5.4 when it can tap into its own content. This suggests the data is doing as much work as the specialized training.

The company's framing is careful about what it claims. Less than 10 percent of available content has gone into training so far, which implies room for improvement. CTO Joel Hron notes they have changed the open-source starting point close to half a dozen times already, and research chief Jonathan Schwartz emphasizes the "model factory" over any individual model. The comparison with frontier models is also lopsided: Thomson competes with test-time scaling, while GPT-5.5 runs without reasoning mode, and the company did not test newer models.

For companies with a similar profile, the case suggests the open-source community trails frontier labs by only months, and $40 million can be enough for a competitive specialized model. The deeper point, as Hron puts it, is that the next competitive advantage in AI will come from knowing how to orchestrate it and knowing which intelligence is important enough to own. Whether Thomson's razor-thin lead holds will depend on whether it moves to a stronger variant like Qwen3.8 and pushes past the 10 percent content mark.

FAQ

Why did Thomson Reuters build its own AI model instead of using OpenAI or Anthropic?
The company cites three reasons: standard fine-tuning degrades general capability, they avoid being locked into a provider for costs and roadmap, and the compounding effect of owning the model long-term, which builds equity over time.
Where is the Thomson model being used first?
It takes over the Tabular Analysis feature in CoCounsel Legal, where a smaller, cheaper model makes economic sense. The product stays multi-model, and administrators can switch.
How does Thomson compare to GPT 5.4 on accuracy?
On the company's in-house Deep Research benchmark with web access alone, Thomson scores 0.53 on factual accuracy while GPT 5.4 hits 0.65. With access to the company's content, Thomson edges past GPT 5.4, 0.83 to 0.82.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleEtron chairman sees memory boom lasting to 2028-2030