Uncategorized

What is model distillation? The AI technique at the centre of the US-China technology battle

ET logo

Model distillation, a long-established artificial intelligence (AI) training technique, has emerged as the latest flashpoint in the growing technology rivalry between the United States and China. While the method helps developers build smaller, cheaper AI models, US companies have accused some Chinese firms of using it to extract capabilities from proprietary AI systems without permission.

Here is what model distillation is and why it has become controversial.

What is model distillation?

Training the world’s most advanced AI models, often called frontier models, requires massive computing power, data and investment.

Model distillation is a process in which a powerful AI system, known as the “teacher” model, helps train a smaller “student” model. Instead of copying the original model, the teacher generates outputs such as answers, explanations or computer code. These outputs then become training material for the smaller model.

The student model does not inherit the teacher’s architecture, parameters or full capabilities. Instead, it learns selected behaviours that allow it to perform specific tasks more efficiently.

Why is it important?

The biggest advantage of model distillation is lower cost.

Large AI models often require expensive chips and large data centres to operate. Distilled models can run on less powerful hardware, making them easier to deploy across smartphones, factories, vehicles and private enterprise networks.This allows businesses and governments to use AI more widely without bearing the cost of running frontier models.

What are reasoning traces?

The latest generation of AI systems has increased interest in transferring not just the final answer but also the reasoning process used to reach it.

These “reasoning traces” show how an AI system solves a problem step by step, helping a smaller model learn the approach instead of simply memorising answers.

Florian Tramèr, assistant professor at ETH Zurich, compared it to learning mathematics.

“If I give you a book of complicated math problems with final solutions, you will have a much harder time learning how to solve problems than if I gave you detailed solutions that describe all steps to take,” Tramèr told Reuters.

As reasoning traces become more valuable, companies have become more protective of AI outputs because they may reveal how advanced systems solve complex tasks.

Who uses model distillation?

Distillation is widely used across the AI industry and is not considered improper by itself.

Researchers and companies in the US have used it for years, including Stanford University’s Alpaca project and Microsoft’s Orca research, which relied on outputs from more advanced models to improve smaller AI systems.

Chinese researchers have also used outputs from US models in public research, including projects aimed at developing Chinese-language instruction models.

A key distinction is between open-weight and closed AI models. Open-weight models allow researchers to inspect and modify their parameters. Closed models such as OpenAI‘s ChatGPT and Anthropic’s Claude remain under company control and are generally accessed through proprietary application programming interfaces (APIs).

Why has it become a US-China issue?

The dispute is not over model distillation itself but over whether companies can use outputs from proprietary AI systems without permission.

US AI companies argue there is a difference between legitimate research and systematically collecting outputs from closed models to reproduce commercially valuable capabilities.

Anthropic has accused Chinese companies, including DeepSeek, Moonshot and MiniMax, of running large-scale campaigns to obtain capabilities from its Claude models, including software engineering and advanced reasoning functions.

OpenAI has also said it detected attempts by Chinese actors to use its models for distillation-related purposes.

According to Reuters, no Chinese company has publicly accused US AI firms of distilling capabilities from closed-source models.

Source link

Visited 1 times, 1 visit(s) today

Leave a Reply

Your email address will not be published. Required fields are marked *