The article argues that following the challenge of lacking advanced AI chips, China is facing another serious challenge: a shortage of high-quality Chinese-language data to train next-generation AI models.
According to the US research institute Epoch AI, high-quality, publicly available human-created text globally could be completely exhausted within the next six years.
OpenAI co-founder Andrej Karpathy also warned of a “data wall” by the end of this decade, as the capabilities of AI models could plateau without additional new and reliable data.
US AI companies have accelerated offline data collection. According to court filings, Anthropic reportedly spent tens of millions of dollars to buy millions of physical books, unbind them, scan and digitize the entire content, and then discard the print copies, causing widespread controversy regarding ethics and copyright.
For China, the problem is more severe because Chinese accounts for only about 1.3% of content on the global internet, much lower than English (nearly 50%), Spanish (6%), German (5.9%), and Japanese (5%).
In June, China’s National Data Bureau announced a plan to significantly increase the supply, circulation, and commercialization of high-quality AI training data.
By 2028, China aims to build a verified data ecosystem for sectors such as manufacturing, energy, healthcare, finance, agriculture, embodied AI, autonomous vehicles, and low-altitude aviation.
Chinese experts are calling for the construction of large-scale Chinese data repositories by digitizing historical archives, local chronicles, ancient manuscripts, scientific documents, dictionaries, audio-visual content, and even dialects.
Beijing is also encouraging enterprises to develop synthetic data to supplement human-generated data. However, the digitization of documents is facing backlash from publishers. Some newly printed books clearly state clauses prohibiting the use of content for AI training and warn of legal action against violations.
📌 AI competition is shifting from chips and models to data. As high-quality public data sources increasingly dry up, China faces a greater challenge since the volume of Chinese content on the internet accounts for only about 1.3% globally. To maintain the pace of AI development, the country is accelerating the digitization of offline knowledge repositories, building a national data ecosystem, and developing synthetic data. However, the issues of copyright and data ownership could become major barriers in the next stage of the AI race.
