AI has three major inputs: computing power, algorithms, and data. Over the past two years, public attention has mostly gone to the first two 鈥?the chip race and model iteration. But without data, even the strongest models cannot be trained; and at the application stage, how much model capability can be realized depends directly on how well the data has been prepared. To understand why large models hit the "data wall" and why AI often performs poorly once it enters specific industries, one must first return to a basic question: what exactly is data?
The Two Characteristics of Data
The most basic constituent elements of the universe are only three: matter, energy, and information. Matter and energy are "entities," information is a mapping of entities, and data is a record of information. Understanding the economic logic of data begins with two characteristics that distinguish it from matter and energy.
The first characteristic is non-rivalry. Matter and energy are both "rivalrous": a steamed bun can only be eaten once, and a ton of iron ore can only enter one blast furnace. Data has no such rivalry. The same medical record, the same batch of driving records, the same set of industrial parameters can be used simultaneously by multiple institutions and multiple models, and the cost of copying is extremely low. This means that the more widely data is used, the greater the total value it creates.
The second characteristic is exclusivity, which comes from property rights rather than physics. The fact that data can be infinitely copied does not mean that anyone has the right to obtain and use it. The scarcity of data is precisely defined by property rights. The ownership of data property rights varies widely: the property rights of data such as medical records, genes, payment records, and travel trajectories belong to individuals and are personal privacy; the property rights of data such as factory equipment parameters, yield data, and failure logs belong to enterprises and are trade secrets; data involving national security, such as energy, transportation, finance, and geospatial information, may belong to the government and constitute state secrets.
Together, these two characteristics form the core contradiction of the data economy: non-rivalry means that sharing can create greater total value, while exclusivity gives owners ample reason to restrict sharing. The high-value data that can truly and substantially improve intelligence is precisely the part that is hardest to obtain. When high-value data is held by enterprises, out of competitive concerns, the rational choice is often to hoard rather than share, so data exhibits artificial scarcity at the system level.
The Data Wall: The Era of Free Data Is Ending
What large model training consumes is, in a sense, humanity's accumulated "inventory." The internet text that makes up the bulk of pre-training corpora is information produced and sedimented by humanity over the past decades; the books and classic literature within it represent thousands of years of accumulation. But the growth of internet text has long since fallen from the doubling expansion of earlier years to single digits. The average text volume of web pages has grown only 2% to 4% annually over the past decade, while the appetite of frontier models continues to expand exponentially with each generation. A trickle of supply meets a flood of demand 鈥?this is the "data wall." But a key distinction must be noted: it is a wall of effective supply, not a wall of total data volume. What is truly tightening is only the portion that is publicly available and legally usable. The easier data is to obtain, the less valuable it is; the more valuable data is, the harder it is to obtain.
The other side of the data wall is a fundamental transformation in the mode of supply. Over the past decade, training data was basically "picked up for free." Now AI has raised the value of data. As a result, we see: unauthorized book corpora taken down due to copyright complaints, social platforms no longer allowing free scraping of user content, and news organizations shifting from lawsuits to paid licensing. Reddit opened its community content to Google for about $60 million per year; News Corp's five-year agreement with OpenAI totaled more than $250 million. From now on, data must either be bought with money, created with money, or unlocked through institutions 鈥?each of the three paths has its costs.
Buying with money means data becomes an asset with a clear price, and the two deals above are the earliest quotes in this new market. Creating with money means using machines to produce synthetic data, decoupling "high quality" from "expensive human annotation"; but synthetic data does not automatically expand the boundaries of knowledge, and its effectiveness depends heavily on external verification 鈥?in domains with verification methods (such as mathematics and code), it is a lever that amplifies capability; in domains without verification methods, what it amplifies is often just noise. Unlocking through institutions means using rules to change incentives so that previously locked data is willing to flow. Whichever path is taken, the era of picking "low-hanging fruit" is over; the next round of competition depends on deeper extraction capability: whoever can transform scattered data, tacit experience, and the traces of expert thinking into high-quality data that is learnable, governable, and sustainably supplied will be more likely to drive the next leap in AI capability.
The Industry Bottleneck Is Also in Data
The data wall describes the macro constraint at the frontier model level. When AI moves into specific industries, the data problem breaks down into two layers of bottlenecks closer to the ground.
The first layer is stuck at the training end: general large models learn from public text, but the most core intelligence in each industry is precisely not in public text 鈥?the differential thinking of doctors during consultations, the weighing of considerations by fund managers before placing orders, the reasoning paths of lawyers facing case files, and the intuition of maintenance engineers when troubleshooting. These "process records" of professional judgment barely exist on the internet.
The second layer is stuck at the application end: enterprises' own data is not ready. A 2024 survey by data company Precisely and Drexel University showed that only about 12% of surveyed organizations believed their data quality and accessibility were sufficient to support effective AI applications; in Informatica's 2025 survey of chief data officers, about 43% listed data quality and readiness as the primary obstacle to moving AI from pilot to production. In specific industries, manufacturing's difficulty is "available but unusable": factories are not short of data; the difficulty lies in the fact that various systems come from different vendors, different eras, and different protocols, with misaligned definitions; an even greater gap lies outside the systems 鈥?the judgment of senior engineers, what to check first when a failure occurs, and whether to dare to continue production when parameters are abnormal. Most of these have never been recorded; even if the data is assembled, there is still the verification hurdle. Production lines have near-zero tolerance for errors, and AI cannot learn by trial and error the way it does in a chat box.
Healthcare's difficulty is "abundant but disconnected": privacy and compliance lock data inside hospitals, making cross-hospital flow difficult; annotation consumes the scarcest resource 鈥?the time of excellent doctors; feedback on treatment effects is measured in years, making the verification loop extremely slow; and what is stored in medical records are conclusions, not reasoning. Doctors' true differential diagnosis thinking is not written into medical records.
There is also a positive example: programming. The application area where AI has penetrated most deeply and achieved the most remarkable results is programming. Programming satisfies almost every condition of "effective data." In terms of availability, GitHub alone hosts more than 500 million code repositories and has more than 180 million registered developers. Engineers around the world publicly host code together with modification records, review comments, and issue discussions, and open-source licenses open the door to reuse. In terms of quality, code is a formal language with clear grammatical constraints, far easier for machines to parse than natural language. In terms of verification, code comes with its own "grading" mechanism 鈥?whether it compiles and whether tests pass can be determined by the machine itself, and training signals are naturally verifiable. In terms of process, version history preserves in large quantities "how this problem was solved step by step." In terms of updates, new submissions from developers around the world pour in every day. Available, high-quality, verifiable, process-rich, and continuously sustainable 鈥?programming misses none of the five conditions. The open-source movement was an institutional experiment completed decades in advance. With an incentive design of reputation, collaboration, and licensing agreements, it pre-prepared for AI a verifiable, continuously updated data mine. In other words, programming first solved the economic problem of data 鈥?rights confirmation and incentives 鈥?before welcoming AI prosperity. For many industries, the key to unlocking data is not in algorithms, but in institutions.
Data Is the Only Home Ground Where Ordinary Enterprises Can Exert Strength
For enterprises, if AI is to be implemented, data preparation is not optional; it is the ticket to entry: collection, cleaning, alignment, annotation, governance, and rights confirmation 鈥?each is a real investment of money. Governance roughly has three levels: the bottom level is routine cleaning, deduplication, and noise filtering; one level up is establishing traceability and version control for important knowledge; in high-value professional fields, fine-grained annotation engineering and deep participation by domain experts are needed to build "small but refined" high-confidence datasets. At the same time, a pipeline is needed to continuously feed actions, results, and failures from business back into the system to satisfy the condition of "continuous updating."
Among AI's three major inputs, ordinary enterprises have only one card to play. Computing power points to chips and data centers and is a game for giants and nations; algorithms point to frontier model research and are the territory of top laboratories; only data is highly dispersed across vertical industries and scenarios. When the dividend from public data is exhausted, marginal improvements in models increasingly depend on high-quality data. And as model methodologies become increasingly homogeneous, the focus of competition is shifting from "whose model is stronger" to "whose data is cleaner, scarcer, and harder to replicate." In June 2025, Meta invested $14.3 billion in data service provider Scale AI for about a 49% stake 鈥?the fact that a giant is willing to pay such a price for data capability is itself evidence of the rise of data value.
Some of the data enterprises possess is intertwined with their own business: customer preferences, process parameters, failure experiences, and expert judgment. This type of data can constitute a true competitive barrier because it has three characteristics: it increases in value with use, has high migration costs, and is difficult to bypass. But having unique data does not equal having profit. Healthcare is a counterexample: the AI medical imaging industry is generally unprofitable. From data accumulation to profit, an additional round of testing is required: Is there substitute data? Will model generalization dilute uniqueness? Is the value capture link in one's own hands?
A more fundamental boundary lies in the nature of the feedback loop. In fields where feedback is difficult to establish at low cost, such as healthcare, industry, and credit, the barrier of proprietary data is structural; in fields where feedback can be automatically verified, such as mathematics and programming, models can continuously evolve through self-generated and self-filtered data, and the advantage of proprietary data is more likely to be only阶段性. Therefore, the first step in evaluating the prospects of a data asset is to determine which category its domain belongs to.
Beyond the Enterprise Boundary: Technology, Markets, and Institutions
When the social value of data is clearly greater than the private return, or when the risk exceeds what a single entity can bear, the problem must be solved beyond the enterprise boundary. What works here is mainly one technological force and two institutional paths. Technology answers "how to participate in training while reducing the risk of leakage," but it cannot answer "who has the incentive to participate and how returns are distributed." Liquidity and sustainable incentives still need to be solved through mechanism design.
The first institutional path is market-based mechanisms, giving data owners incentives to open data and share value-added returns under compliance, in forms including data trading markets, data trusts, and data asset capitalization. But this path is naturally constrained by the two characteristics of data: exclusivity determines that the tradable stock is limited, and scenario dependence makes pricing difficult 鈥?the value of the same piece of data varies by task, making it hard to form a unified market price. Therefore, marketization is still in an early exploratory stage.
The second institutional path is top-down public arrangement, suitable for areas of market failure and capable of being handled in layers according to the strength of externalities. Data involving national security and public interest, such as meteorological, geospatial, and basic demographic information, should be led by the government or industry organizations to build public data platforms and high-quality public datasets, turning fragmented resources into infrastructure; livelihood data with strong externalities such as healthcare and transportation requires stronger public arrangements under unified standards and strict privacy boundaries; while data oriented toward commercial competition should discover value through market-based methods, with government regulating rather than replacing the market.
A longer-term institutional imagination can return to the contradiction that appears repeatedly in this article: should incentives be given to owners, or should society as a whole benefit? The patent system actually long ago provided a compromise 鈥?using a limited period of protection to strike a balance between the two that has been broadly effective for centuries. Data may also need a similar arrangement: as a public policy, certain categories of data must be opened after a certain number of years. During the protection period, the owner's investment is rewarded; after it expires, the sealed intelligence returns to the public domain.
In summary, data is not only a technical issue, but also a management and economic issue. For entrepreneurs, data preparation is the ticket to making good use of AI, and data quality will become an important source of differentiation in competitiveness in the AI era, deserving strategic-level attention. And because of the externalities of data, we must also step outside the enterprise and look at it from the national perspective. China has a population scale, market depth, and a complete industrial system, and has accumulated extremely rich consumer data, industrial data, and urban operation data. Half of the wall depends on enterprise governance, and the other half depends on institutional innovation; let data owners receive returns, and let the value of data belong to society 鈥?the data economy of the AI era must ultimately find a balance between these two ends.