The risk of “distillation attacks” – where malicious actors extract the knowlege from proprietary large language models (LLMs) – is escalating, and is no longer a theoretical concern. As security researcher Ben hultquist warned, the danger will spread as more companies develop and deploy their own models trained on sensitive internal data.

“We’re on the frontier when it comes to this, but as more organizations have models that they provide access to, it’s unavoidable,” he saeid. “as this technology is adopted and developed by businesses like financial institutions,their intellectual property could also be targeted in this way.”

In February 2024,OpenAI blamed DeepSeek and other Chinese LLM providers and universities for copying ChatGPT and other US firms’ frontier models in a memo [PDF] to the House Select Committee on China. The memo also noted some activity from Russia and warned that illicit model distillation poses a risk to “American-led,democratic AI.”

China’s distillation methods have become increasingly complex over the past year, moving beyond simple chain-of-thought (CoT) extraction to multi-stage operations. These include synthetic-data generation, large-scale data cleaning, and other stealthy techniques. As OpenAI wrote:

openai also notes that it has invested in stronger detections to prevent unauthorized distillation. It bans accounts that violate its terms of service and proactively removes users who appear to be attempting to distill its models. However, the company admits that it alone can’t solve the