- China considers data a decisive factor in the AI race, alongside models and computing power.
- In 2023, researchers at the Beijing Institute of Technology evaluated that ChatGPT answered many questions about Chinese culture incorrectly and expressed biased viewpoints.
- Beijing fears that AI models trained primarily on English data will reflect Western values and perspectives.
- The government aims to turn China into a data powerhouse by 2028, according to the National Data Bureau’s plan.
- The plan proposes building high-quality datasets for over 20 strategic sectors such as science, manufacturing, and autonomous vehicles.
- China commits to sharing data with developing countries to help them build their own AI systems.
- State laboratories and media have released many free datasets on GitHub and Hugging Face.
- The goal is to attract global developers and narrow the gap with the US in high-quality training data.
- Experts note that China still faces difficulties as domestic data is fragmented between agencies and enterprises.
- The plan calls for breaking down “data islands” and developing a force of expert data labelers to replace cheap labor.
- Universities are encouraged to open training programs for AI data labeling.
- Western analysts warn that Chinese data could expand the influence of state propaganda.
- A study in Nature showed that ChatGPT and Claude respond more positively about China when using Chinese compared to English.
- The WanJuan dataset developed by the Shanghai AI Laboratory is introduced as being consistent with “Chinese mainstream values,” supporting multiple languages including Arabic, Russian, Thai, and Vietnamese.
- Observers say combining low-cost AI models with open datasets will help China expand its technological influence in developing countries.
📌 The AI competition between China and the US is expanding from chips, models, and computing power to training data. Beijing aims to become a global data supplier while building its own standards and AI ecosystem. However, this strategy also sparks debate about the risk of global AI models absorbing more content reflecting China’s official views, especially when this data is used to train or fine-tune AI systems in many countries.
