Written by: Xiaobing
On August 24, Thomson Reuters announced the launch of its self-developed large language model "Thomson." The information giant, with over $7 billion in annual revenue, claimed that the model was trained on an open-source base, with a total investment of approximately $40 million (covering talent and computation), and training data sourced from the Westlaw legal database, Practical Law practice guides, Checkpoint tax research platform, and Reuters news assets. Early assessments showed that Thomson's performance on multiple tasks is "comparable to the latest cutting-edge models."
Comparing the $40 million: OpenAI's latest round of funding was $40 billion, Anthropic has raised over $13 billion in total, and xAI secured $6 billion in a single round. Cutting-edge laboratories pile billions of dollars to train general models, while Thomson Reuters achieved capabilities claimed to match cutting-edge without even spending a fraction of that in a vertical field.
CTO Joel Hron stated: The AI industry has treated scale as the answer for years, with bigger models, more computing power, and more money. Thomson proves there is another path: starting from a strong foundation and deeply specializing in truly important tasks can build efficient and fully self-controlled intelligence.
The era of building AI in vertical industries is beginning.
What did Thomson do?
Thomson started with an open-source base (the announcement did not specify which model) and injected Thomson Reuters' proprietary data through mid-training and post-training processes.
The announcement revealed that less than 10% of its own content has been used for training so far, and it will continue to explore new specialization directions. Hundreds of subject matter experts participated throughout the process, from designing training objectives to final evaluations.
The first deployment scenario for the model is the Tabular Analysis function in CoCounsel Legal, aimed at law firms and corporate legal departments. CoCounsel maintains a multi-model architecture, using Thomson in scenarios where it has a clear advantage, while continuing to use external cutting-edge models in other scenarios.
Thomson Reuters also released a "small" open-source version on Hugging Face for academic and non-commercial use and invited legal and AI scholars to conduct independent evaluations.
Professor Jonathan Choi from the University of Washington School of Law tested Thomson, ChatGPT, and Claude with high-difficulty questions from corporate tax courses. He noted that all three models answered correctly, but he preferred Thomson's responses, especially because the links to legal literature made the answers more transparent and practical.
The Economics of $40 Million
This figure is the key to understanding the whole matter.
The cost of training a general cutting-edge model is rising exponentially. Meta trained Llama 3 using over 16,000 H100 GPUs, with computational expenses estimated in the hundreds of millions of dollars. OpenAI and Google DeepMind's training budgets are even higher. The goal of these models is to "do everything," from writing poetry to solving differential equations, from programming to role-playing.
What Thomson Reuters is doing is completely different logically. It does not need a model that can do everything; it needs a model that performs just as well, or even better, in the extremely vertical fields of legal research, tax compliance, regulatory interpretation, and specialized document analysis. The scope of this goal is much narrower, and thus the training costs are much lower.
But what truly made $40 million possible is not the narrowed scope, but the data assets.
Thomson Reuters' Westlaw legal database covers over 40,000 case databases and more than 100 years of case law accumulation. Practical Law includes operational guidelines and templates maintained by over 650 attorney editors. Checkpoint is the standard tool for tax professionals in the U.S. The Reuters news corpus covers the globe, with a history spanning over 175 years.
This data is not scraped from the internet.
They are proprietary assets that have been specially edited, annotated, structured, and continuously updated. Vertical models trained on such data achieve accuracy and citation quality in their specialized fields that general models find difficult to match through simple RAG (retrieval-augmented generation) or fine-tuning. General models can "access" this data, but accessing it and "growing within it" are two different things.
Thomson Reuters itself pointed out this distinction: The gains from domain-specialized training cannot be replaced by mere content access.
Who will take this path?
Thomson Reuters will not be the only one. Along the chain of "proprietary data → vertical models → lower inference costs → stronger data control → higher corporate gross margins," at least several types of companies are qualified to replicate this model.
Financial data companies. Bloomberg has already trained BloombergGPT in 2023, based on its proprietary financial terminal data and news corpus. Although there has not been much follow-up action, the data barriers of Bloomberg Terminal and Thomson Reuters' Westlaw are of the same magnitude—these are proprietary datasets that no general model can legally obtain.
Professional service firms. The Big Four accounting firms (PwC, Deloitte, EY, KPMG), large law firm alliances (such as Baker McKenzie, Kirkland & Ellis), and healthcare information companies (such as Epic Systems' electronic medical records data) possess vast amounts of highly structured and confidential professional data. These institutions are reluctant to hand over client data to general models while still needing AI to enhance efficiency. Building vertical models may not just be a choice for them; it is a necessity from a compliance perspective.
Industry data monopolists. MSCI owns global ESG rating and index data, Verisk has insurance pricing and risk model data, Wolters Kluwer possesses European legal and compliance data, and RELX has Elsevier academic journals and LexisNexis legal data. These companies share a common characteristic: data is the core asset, the exclusivity of data is the moat, and feeding data to others' models dilutes its own value.
What does this mean for OpenAI and Anthropic?
There is a detail in Thomson Reuters' announcement that is easily overlooked: CoCounsel maintains a multi-model architecture. Use Thomson where it has advantages, but continue using external models in other scenarios.
This indicates that even enterprises that have built their own models will not completely abandon general models.
General models still have advantages in general reasoning, code generation, multilingual processing, and other areas. Vertical models replace the level at which general models are "adequate but not good enough" in specific fields.
However, in the long term, if more and more high-value enterprise clients begin to build vertical models, the risk faced by general model companies is being compressed to the infrastructure layer: providing base models for secondary training by enterprises and offering inference APIs for general tasks, but losing pricing power in the most profitable vertical application scenarios.
This is similar to the evolution of the cloud computing industry.
AWS, Azure, and GCP provide infrastructure, but the most profitable SaaS application layer is occupied by vertical companies like Salesforce, ServiceNow, and Workday. If the AI industry follows the same path, OpenAI and Anthropic's roles could become more like "the AWS of AI," rather than "the Salesforce of AI."
What Thomson Reuters demonstrated with $40 million is: when you have sufficiently unique data, you can achieve parity with general models trained with billions in costs at a fraction of the cost within your field.
Once this logic is validated, every company sitting on a proprietary data mine will recalculate: Should we continue paying OpenAI or Anthropic API fees each year, or spend a relatively much smaller amount to train a vertical model that we fully control?
For a market accustomed to the narrative of the "AI arms race," Thomson Reuters' $40 million is a counter-signal. In the next stage of AI competition, it may not be the company with the highest training costs that wins, but rather the one with the most unique and irreplaceable data. The model is a tool; data is the barrier.
免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。