Anthropic details distillation efforts by Chinese companies, like Moonshot and DeepSeek, sending user queries to Claude via "transfer stations" outside China
First reported by WSJ ·
If you use an AI chatbot that is not from a major Western provider, it may have been trained on stolen data.
Anthropic, the AI safety and research company, has detailed how two Chinese AI companies, Moonshot and DeepSeek, allegedly engaged in a "distillation" process to replicate its Claude large language model's capabilities. The company reports that these firms created thousands of fake user accounts and utilized millions of real user queries. These queries were reportedly sent to Anthropic's Claude models through "transfer stations" located outside of China, circumventing national restrictions. Anthropic claims this effort allowed Moonshot and DeepSeek to train their own models by observing Claude's responses to a vast dataset of user prompts. The company has stated this practice violates its terms of service.
This disclosure by Anthropic reveals a sophisticated method employed by some AI developers to bypass geographical restrictions and leverage proprietary models for competitive advantage. The use of "transfer stations" suggests a technical workaround designed to mask the origin of the data being used for distillation. It also highlights the significant challenge AI companies face in protecting their models from being reverse-engineered or mimicked.
The practice of distilling models, especially through unauthorized access to query data, raises critical questions about intellectual property, data privacy, and the ethical development of AI. For companies like Moonshot and DeepSeek, this approach could lead to faster development cycles but also carries risks of legal repercussions and reputational damage. It underscores a growing tension between open access and proprietary control in the rapidly evolving AI landscape.
AI-written summary. May contain errors.