CANLI
Ana Sayfa🇹🇷 Türkiye🌍 Dünya📈 Ekonomi⚽ Spor💻 Teknoloji🎭 Magazin
Ana SayfaDünyaAre Chinese Firms Handing the United Sta
🌍 Dünya

Are Chinese Firms Handing the United States Data?

Foreign Policy·🕐 1 sa önce·👁 1 görüntülenme
AI competition may be an unwitting weakness for Beijing.

A recent report by the U. S. artificial intelligence company Anthropic accused several leading Chinese AI firms, including Moonshot AI and DeepSeek, of distilling Claude’s capabilities on a large scale. Anthropic said some of these companies went beyond using Claude to generate training data and, without users’ knowledge, in one instance routed at least 300,000 requests from real users through the model over a 10-day period.

The data reportedly sent to Claude included surveillance material submitted by a user Anthropic suspected of being affiliated with the People’s Liberation Army; internal code and valid credentials from engineers at major Chinese state-owned enterprises; information from police case management systems; and details of internal AI projects at Chinese technology companies.

A recent report by the U. S. artificial intelligence company Anthropic accused several leading Chinese AI firms, including Moonshot AI and DeepSeek, of distilling Claude’s capabilities on a large scale. Anthropic said some of these companies went beyond using Claude to generate training data and, without users’ knowledge, in one instance routed at least 300,000 requests from real users through the model over a 10-day period.

The data reportedly sent to Claude included surveillance material submitted by a user Anthropic suspected of being affiliated with the People’s Liberation Army; internal code and valid credentials from engineers at major Chinese state-owned enterprises; information from police case management systems; and details of internal AI projects at Chinese technology companies.

The leakage of sensitive information will of course alarm Beijing. But compared with one or two pieces of sensitive material reaching the United States, the bigger worry may be that, in the process of distilling the capabilities of the most advanced U. S. models, Chinese AI companies may also be sending large quantities of real Chinese user data to the United States.

AI competition is not just about compute but about data. Today, data available on the open internet is increasingly easy to obtain. What is truly scarce is the high-quality data generated by real tasks: why users turn to AI, what problems they are trying to solve, what background materials they provide, how far they have progressed in writing code, what technical obstacles companies encounter, and what they want AI to help them accomplish. Such data can be far more valuable than the ordinary detritus of the web.

On its own, such data may not mean much or be worth much. But once hundreds of thousands, millions, or even tens of millions of such pieces of data are aggregated, their nature changes completely. A Chinese programmer asking why a piece of code failed has little or no intelligence value. But if thousands of Chinese engineers repeatedly encounter and ask about similar problems, the aggregate may expose a common bottleneck in a Chinese technological system.

Likewise, an employee asking how to improve a particular piece of equipment or software would hardly be a concern—or a secret. But if large numbers of engineers from industries such as semiconductors, robotics, aerospace, telecommunications, electricity, and automobiles are continually talking about their work to AI models, over time those interactions could allow others to form a dynamic map of China’s industrial and technological capabilities.

This data is more telling than conventional searches. Those queries generally reveal what a person wants to know, but AI often reveals what a person is actually doing. Users voluntarily provide work context, upload code, documents, and project descriptions; explain what they have already tried; identify the problems they have encountered; and describe what they intend to do next.

This is a very particular kind of data. It might be called intent data. It tells outsiders not only what China possesses but what China is doing.

In peacetime, such data primarily has commercial and technological value. But as tensions between the two superpowers grow, so too will the value of such a data mine.

Imagine a scenario where war seems to be on the horizon. U. S. intelligence agencies would need to know not only how many missiles and warships China possesses but also whether its military-industrial system is accelerating production; where bottlenecks are emerging; which civilian companies are beginning to enter wartime production; what weaknesses exist in semiconductors, drones, communications, electricity and transportation; and which research institutions are suddenly concentrating on particular problems. Large volumes of AI interactions could provide precisely this kind of information.

When the institutions, names, equipment, software, suppliers and technical failures mentioned in AI conversations are combined with satellite imagery, customs records, procurement data, patents, academic papers, corporate information, network data, and other intelligence already available to the United States, an apparently ordinary user query can become one node in a much larger network. Hundreds of thousands or millions of such nodes, taken together, could prove far more valuable than one or two traditionally classified documents.

This raises a question that has received almost no attention: Did Beijing previously know that Chinese AI companies, in the course of distilling U. S. models, were sending real user data to the United States?

China has plenty of data security regulations. In recent years, Beijing has repeatedly tightened oversight of cross-border data transfers. The Provisions on Promoting and Regulating Cross-Border Data Flow impose security assessment requirements and other restrictions on the overseas transfer of important data—defined in terms of what could endanger national security, the economy, or social stability—and large volumes of personal information. China’s rules governing generative AI services also explicitly require providers to protect users’ personal information, trade secrets, and data security.

The question is whether this regulatory system recognized a new form of cross-border data transfer. Instead of a Chinese company formally exporting a database to a foreign company, an AI company seeking to train its own model may use proxy accounts and intermediary services to quietly feed batches of real user requests into a U. S. foundation model.

If Beijing was unaware, then Anthropic’s report has exposed a major blind spot in China’s AI regulatory system. China may have been scrutinizing multinational companies transferring databases to the United States while failing to notice that its own leading AI companies were continuously sending users’ work, code, technical difficulties, and even sensitive materials into U. S. models through model calls.

But even if China now understands the problem, that will not make it easy to resolve.

On the one hand, Beijing may require Chinese AI companies to stop such practices or at the very least prohibit them from sending unprocessed real user data into U. S. models. Given China’s existing data security laws and regulatory capacity, doing so would not be difficult.

On the other hand, why would Alibaba, Moonshot, DeepSeek, Xiaomi, and other companies take such risks to access and distill Claude on a massive scale? The answer appears equally obvious: The capabilities of U. S. frontier models remain valuable to Chinese companies.

According to Anthropic, Alibaba’s distillation-related activity exceeded 151 million exchanges between May and July; Moonshot’s exceeded 23 million; and DeepSeek conducted more than 12. 1 million exchanges in just 14 days in July. This amounted to industrial-scale acquisition of technological capability. Moonshot and DeepSeek were even accused of building dedicated technical pipelines designed to extract Claude’s reasoning process for use in training their own models.

This illustrates the dilemma confronting China. Distilling the most advanced U. S. models has become an important means for some Chinese companies to narrow the technological gap. China has engineers, computing power, its own large models, and a strong AI industry. But if Chinese companies are completely cut off from distilling the capabilities of leading U. S. models, producing training data and reasoning capabilities of comparable quality independently would require more time, computing power, and trial and error. In the short term, the pace at which China catches up with U. S. frontier models could therefore slow.

If Chinese AI companies are allowed to continue distilling U. S. models on a large scale, Beijing must accept the risk that Chinese data will continue to flow into U. S. systems. If that route is tightly shut down for data security reasons, however, the cost of catching up may rise, and the technological time gap may widen further.

As AI competition increasingly depends on distillation, data, and the exchange of model capabilities, China therefore faces a difficult question: How can it take advantage of the most advanced U. S. AI capabilities without allowing China’s own data to become a strategic asset for the United States?

Deng Yuwen is a Chinese writer and scholar.

Commenting is a benefit of a Foreign Policy subscription.

Already a subscriber? Log In.

Join the conversation on this and other recent Foreign Policy articles when you subscribe now.

Please follow our comment guidelines, stay on topic, and be civil, courteous, and respectful of others’ beliefs.

I agree to abide by FP’s comment guidelines. (Required)

The default username below has been generated using the first name and last initial on your FP subscriber account. Usernames may be updated at any time and must not contain inappropriate or offensive language.

I agree to abide by FP’s comment guidelines. (Required)

Washington needs its rivals to create peace on its behalf.

Kaynak: Foreign PolicyOrijinal Habere Git →
İlgili Haberler