China expands high-quality datasets to support AI development

High-quality datasets expand to support research, industry and healthcare.

High-quality datasets in China reached around 120,000 by the end of June, with total data volume increasing by more than 60% since the first quarter.

China has built approximately 120,000 high-quality datasets as part of a broader effort to expand the data resources supporting AI development and the digital economy, according to the National Data Administration.

By the end of June, these datasets, covering fields such as scientific research, industrial manufacturing and healthcare, totalled more than 1,565 petabytes, an increase of over 60% compared with the end of the first quarter.

According to the administration, the total volume is equivalent to roughly 547 times the digital holdings of the National Library of China, illustrating the scale of the country’s expanding data resources.

China has also expanded data annotation through pilot programmes in seven cities, including Chengdu, Shenyang, Hefei and Changsha. Officials said these initiatives have processed more than 119 petabytes of annotated data while supporting around 140,000 data annotation jobs.

Why does it matter?

High-quality datasets are becoming a strategic resource for AI development, as they underpin the training, testing and evaluation of increasingly capable AI models. Expanding data availability and annotation capacity can strengthen research, industrial applications and the development of domestic AI ecosystems.

The announcement also reflects China’s broader strategy of treating data as national digital infrastructure. Alongside investment in computing power and AI models, expanding high-quality datasets is intended to support innovation while reinforcing the country’s long-term competitiveness in artificial intelligence.

Would you like to learn more about AI, tech and digital diplomacyIf so, ask our Diplo chatbot!