White House technology advisor Michael Kratsios accused Chinese AI company Moonshot of copying Anthropic's Fable model and using high-end chips that were not approved for export to China when training Kimi K3. However, several researchers interviewed by TechCrunch expressed skepticism about this claim, arguing that distillation alone is insufficient to explain Kimi K3's performance.
Kimi K3 is considered one of the largest LLMs currently available with open parameter weights. Kratsios posted on social media that large-scale, covert industrial distillation is stealing US proprietary technology and undermining US research advantages. However, he did not disclose further sources of evidence, and Moonshot did not respond to questions regarding its training process.
Experts say the time window is too short.
Braden Hancock, a Laude Institute researcher and co-founder of Snorkel AI, said that relying solely on distillation would make it difficult to complete sufficient data collection, model training, and release within two weeks of Fable's public launch on July 1.
Nathan Lambert, a researcher at the Allen Institute for AI, also believes that as Chinese models gradually approach cutting-edge levels, the benefits of relying solely on supervised fine-tuning are declining. In his opinion, if distillation could truly replicate Kimi K3's capabilities quickly, other teams in the industry should be able to catch up more easily, but this is not the case.
Distillation is more like a replication style
Distillation typically involves continuously calling the target model to collect question-and-answer data, which is then used for post-training. Sometimes the model is required to demonstrate the reasoning process, while other times the prompts and responses are directly used for supervised fine-tuning.
Researchers point out that these methods are more likely to replicate the model's expression and response habits than to fully reproduce high-level capabilities in a short period of time. To approach the performance of models like Fable, more complex methods such as reinforcement learning are often required, which demand significantly more computing power, time, and funding.
The controversy also points to restricted chips
The report points out that Anthropic publicly accused Moonshot, DeepSeek, and MiniMax earlier this year of systematically distilling their models. Anthropic claims it identified millions of interactions between these companies and its models through IP addresses and other metadata, and believes these requests are clearly different from normal usage patterns.
However, distillation is not unique to Chinese AI companies. The report mentions that Musk also stated earlier this year that distillation is a common practice in the industry, and the line between distillation and synthetic dataset development is not always clear.
Kratsios's other allegation is that Moonshot obtained Nvidia Grace Blackwell 300 chips, which are subject to U.S. export restrictions, and used servers equipped with GB300 chips located in Thailand. Georgetown University researcher Sam Bresnick stated that a black market for advanced chips does exist in China, and he advocates for stricter customer identification and reporting mechanisms for global data centers.
Additional information:The report noted that the Biden administration proposed a federal customer identification rule for data centers in 2024, but no further progress has been seen during Trump's term; meanwhile, U.S. exporters still need to ensure that advanced chips are used only for approved purposes.












