Written by Xiao Bing
On August 24, Thomson Reuters announced the launch of its proprietary large language model, "Thomson." The information giant, with annual revenues exceeding $7 billion, stated that the model was trained on an open-source foundation, with a total investment of approximately $40 million (covering talent and computing resources). The training data was sourced from the Westlaw legal database, Practical Law practice guides, Checkpoint tax research platform, and Reuters news assets. Early evaluations show that Thomson's performance on multiple tasks is "comparable to the latest state-of-the-art models."
$400 million—compare that to OpenAI’s latest funding round of $40 billion, Anthropic’s total funding exceeding $13 billion, and xAI’s single round raising $6 billion. Frontier labs are spending billions of dollars to train general-purpose models with massive compute, while Thomson Reuters achieved what it claims is frontier-level capability in a vertical domain with less than a fraction of that budget.
CTO Joel Hron said: "For years, the AI industry has treated scale as the answer—bigger models, more computing power, more money. Thomson proves there’s another path: starting from a strong foundation and deeply specializing for truly important tasks enables the creation of efficient, fully autonomous intelligence."
The era of enterprises building their own AI for vertical industries is beginning.
What did Thomson do?
Thomson, starting from an open-source foundation (the specific model is not specified in the announcement), incorporates Thomson Reuters's proprietary data through mid-training and post-training.
The announcement reveals that less than 10% of proprietary content has been used for training so far, and further exploration of new specialized directions will continue. Hundreds of subject matter experts participated throughout the entire process, from designing training objectives to final evaluation.
The first deployment scenario for the model is the Tabular Analysis feature in CoCounsel, targeted at law firms and corporate legal departments. CoCounsel maintains a multi-model architecture, using Thomson in scenarios where it has a clear advantage, and continuing to use external state-of-the-art models in other scenarios.
Thomson Reuters also released a "small" open-source version on Hugging Face for academic and non-commercial use, inviting legal and AI scholars to conduct independent evaluations.
Professor Jonathan Choi of the University of Washington School of Law tested Thomson, ChatGPT, and Claude with a challenging question from his corporate tax course and found that all three models answered correctly; however, he preferred Thomson’s response, particularly because its links to legal treatises made the answer more transparent and practical.
$40 million in economics
This number is the key to understanding the whole matter.
The cost of training general-purpose frontier models is rising exponentially. Meta used over 16,000 H100 GPUs to train Llama 3, with estimated compute costs in the hundreds of millions of dollars. OpenAI and Google DeepMind have even higher training budgets. These models aim to "do everything"—from writing poetry and solving differential equations to programming and role-playing.
What Thomson Reuters does is logically entirely different. It doesn’t need a general-purpose model; it needs a model that matches or even surpasses state-of-the-art general models in highly specialized domains such as legal research, tax compliance, regulatory interpretation, and professional document analysis. This goal has a much narrower scope, so the training cost is significantly lower.
But what truly made the $40 million possible was not the narrowed scope, but the data assets.
Westlaw, owned by Thomson Reuters, encompasses over 40,000 case databases and more than a century of legal precedent. Practical Law provides operational guides and templates continuously maintained by over 650 attorney editors. Checkpoint is the standard tool for U.S. tax professionals. The Reuters news corpus covers global events spanning more than 175 years.
This data is not scraped from the internet.
They are proprietary assets that have been professionally edited, annotated, structured, and continuously updated. Vertical models trained on this data achieve accuracy and citation quality in their specialized domains that general-purpose models struggle to match through simple RAG (Retrieval-Augmented Generation) or fine-tuning alone. General-purpose models can "access" this data, but accessing it is fundamentally different from "growing within" it.
Thomson Reuters themselves highlighted this distinction: the gains from domain-specific training cannot be replaced by mere content ingestion.
Who will take this path?
Thomson Reuters won't be the only one. Looking along the chain—proprietary data → vertical models → lower inference costs → greater data control → higher corporate gross margins—there are at least several types of companies capable of replicating this model.
Financial data company. Bloomberg trained BloombergGPT in 2023 using its proprietary financial terminal data and news corpus. Although there have been few subsequent developments, the data moat of Bloomberg Terminal is on par with Thomson Reuters’ Westlaw—these are proprietary datasets that no general-purpose model can legally access.
Professional service firms—including the Big Four accounting firms (PwC, Deloitte, EY, KPMG), major law firm alliances (such as Baker McKenzie and Kirkland & Ellis), and healthcare information companies (like Epic Systems’ electronic health record data)—possess vast amounts of highly structured, highly confidential professional data. These institutions are unwilling to entrust client data to general-purpose models, yet they need AI to enhance efficiency. Building vertical-specific models is not merely an option for them—it is a compliance imperative.
Industry data monopolists. MSCI owns global ESG ratings and index data, Verisk owns insurance pricing and risk modeling data, Wolters Kluwer owns European legal and compliance data, and RELX owns Elsevier academic journals and LexisNexis legal data. The common trait among these companies is that data is the core asset; the exclusivity of the data is the moat, and feeding data into others’ models dilutes their own value.
What does this mean for OpenAI and Anthropic?
A subtle detail in Thomson Reuters' announcement is that CoCounsel maintains a multi-model architecture: using Thomson’s models where they have an advantage, and continuing to use external models in other scenarios.
This indicates that even companies that have built their own models will not completely abandon general-purpose models.
General models still hold advantages in general reasoning, code generation, and multilingual processing. Vertical models replace the layer in specific domains where general models are "good enough but not excellent."
However, in the long term, if an increasing number of high-value enterprise clients begin building their own vertical models, general-purpose model companies risk being compressed to the infrastructure layer: providing foundational models for enterprises to fine-tune and offering inference APIs for general tasks, while losing pricing power in the most profitable vertical applications.
This is similar to the evolution of the cloud computing industry.
AWS, Azure, and GCP provide infrastructure, but the most profitable SaaS application layer is dominated by vertical companies like Salesforce, ServiceNow, and Workday. If the AI industry follows the same path, OpenAI and Anthropic may play roles more akin to "AWS for AI" rather than "Salesforce for AI."
What Thomson Reuters proved with $40 million is that when you have sufficiently unique data, you can match the performance of billion-dollar general models at a fraction of the cost in your domain.
Once this logic is validated, every company sitting on a proprietary data mine will recalculate its costs: continue paying annual API fees to OpenAI or Anthropic, or spend a significantly smaller amount to train a vertically specialized model under its own complete control?
For a market accustomed to the narrative of an "AI arms race," Thomson Reuters' $40 million represents a counter-signal. The next phase of AI competition may not be won by the company with the highest training costs, but by the one with the most unique and irreplaceable data. Models are tools; data is the barrier.
