AI/ML

OpenAI disrupts AI model distillation attack linked to China’s Moonshot AI

(Adobe Stock)

OpenAI has accused individuals associated with China's Moonshot AI of conducting a "distillation attack" that began July 1. The company warns that extracting its models' reasoning at scale could enable rivals to train capable models without the same safety guardrails, with further coverage provided by The Register.

Model distillation is a machine learning technique where one model's outputs are used to train another. In this adversarial case, bulk queries were sent to reproduce the larger model's reasoning. OpenAI reported disrupting a coordinated campaign that ran throughout July, involving over 15,000 users and a significant volume of requests on July 24 and 25. While OpenAI noted it's unclear if all operators were linked to a single rival, the core cluster of the theft allegedly came from Moonshot AI, developer of the Kimi model. This follows similar accusations from Google, Anthropic, and US officials against Chinese rivals.

OpenAI stated that the attackers did not breach encryption or databases but manipulated model interactions. The extracted reasoning could be used to train new models without the original safeguards, posing safety and national security risks. OpenAI has since banned accounts, tightened controls, and shared details with other AI firms and government programs.

Source: The Register

An In-Depth Guide to AI

Get essential knowledge and practical strategies to use AI to better your security program.

Get daily email updates

SC Media's daily must-read of the most current and pressing daily news

By clicking the Subscribe button below, you agree to SC Media Terms of Use and Privacy Policy.

You can skip this ad in 5 seconds