OpenAI has accused individuals associated with China's Moonshot AI of conducting a "distillation attack" that began July 1. The company warns that extracting its models' reasoning at scale could enable rivals to train capable models without the same safety guardrails, with further coverage provided by The Register.
Model distillation is a machine learning technique where one model's outputs are used to train another. In this adversarial case, bulk queries were sent to reproduce the larger model's reasoning. OpenAI reported disrupting a coordinated campaign that ran throughout July, involving over 15,000 users and a significant volume of requests on July 24 and 25. While OpenAI noted it's unclear if all operators were linked to a single rival, the core cluster of the theft allegedly came from Moonshot AI, developer of the Kimi model. This follows similar accusations from Google, Anthropic, and US officials against Chinese rivals.
OpenAI stated that the attackers did not breach encryption or databases but manipulated model interactions. The extracted reasoning could be used to train new models without the original safeguards, posing safety and national security risks. OpenAI has since banned accounts, tightened controls, and shared details with other AI firms and government programs.
Source: The Register
