AI & TechArtificial IntelligenceBigTech CompaniesBusinessNewswire

Why Kimi K3’s success isn’t about exploiting Anthropic’s Fable

▼ Summary

– White House science advisor Michael Kratsios accused Chinese company Moonshot of copying Anthropic’s Fable LLM via distillation using chips not cleared for export to China.
– Experts doubt distillation alone explains Kimi K3’s advanced capabilities, given Fable’s short public availability since July 1st.
– Distillation involves querying a model to generate training data, but researchers say its benefits are diminishing as models grow more complex.
– Anthropic earlier accused Moonshot, DeepSeek, and MiniMax of systematic distillation, citing millions of suspicious queries, though the practice is common in AI.
– Kratsios also alleged Moonshot accessed banned Nvidia chips, possibly via a black market, amid calls for stronger data center and chip export oversight.

The recent accusation by White House science advisor Michael Kratsios that Chinese AI company Moonshot built its Kimi K3, the largest publicly available open-weight LLM, by copying Anthropic’s Fable model using restricted chips has sparked intense debate. Kratsios condemned the practice as a “large-scale, covert industrial distillation aimed at stealing proprietary U.S. technology,” a charge that comes amid discussions about banning Chinese open-weight models. Moonshot declined to comment on its training methods, and Kratsios provided no further evidence for his claims.

His remarks echoed Treasury Secretary Scott Bessent’s earlier assertion that “watermarks of our U. S. large language models” appear on many Chinese models, though neither official specified what those watermarks entail. The Treasury Department did not respond to requests for clarification.

Yet many AI experts are pushing back against the narrative that distillation alone explains Kimi K3’s rapid advancement. Braden Hancock, a researcher at the Laude Institute and co-founder of Snorkel AI, told TechCrunch that replicating Fable’s capabilities through strict distillation in such a short timeframe is implausible. “Fable’s only been publicly available since July 1st. You can’t distill that much data, train a model, and release it in two weeks,” he said.

Nathan Lambert, an AI researcher at the Allen Institute for AI, expressed similar skepticism during a recent podcast. He argued that distillation is becoming less impactful as Chinese models approach the frontier and training shifts toward reinforcement learning. “If it were the case, everyone would be easily able to catch up to a GLM or to a K3 by using its data for distillation. But we have not, or we won’t see this, from supervised fine-tuning alone,” Lambert noted.

Distillation typically involves systematically querying a target model to generate data for post-training. This can include asking for chain-of-thought reasoning or using prompts and responses for supervised fine-tuning (SFT). SFT is where Lambert says a “model picks up its manners,” but he believes its benefits are waning. Replicating Fable-level capabilities would likely demand reinforcement learning techniques, which require an agent from the larger model to grade the smaller model’s responses and adjust accordingly.

Such advanced techniques demand far more infrastructure. Large reinforcement learning runs can involve tens of millions of agents. Using a frontier lab’s API for this purpose would be “insanely expensive and potentially a time bottleneck,” Lambert added, and might not even yield a performance improvement.

Anthropic has previously accused Moonshot, DeepSeek, and MiniMax of systematically distilling its models, citing millions of suspicious queries detected through IP addresses and metadata. Those queries were “distinct from normal usage patterns, reflecting deliberate capability extraction rather than legitimate use,” Anthropic claimed. The company did not respond to queries about Fable distillation. However, distillation is widely practiced across the AI industry, not just in China. Elon Musk testified earlier this year that SpaceXAI distilled OpenAI models to develop Grok, calling the practice common. The line between distillation and creating synthetic datasets is often blurry.

Hancock emphasized that American observers often underestimate Chinese technical expertise. “One of the founders of Moonshot was a CMU PhD student. These are legitimate researchers and engineers doing solid work,” he said. “If American models ground to a halt, I think China’s progress would slow, but would still continue. They’re not just riding coattails here.”

The second part of Kratsios’ accusation involves hardware. He claimed Moonshot obtained advanced Nvidia Grace Blackwell 300 chips and accessed GB300-equipped servers in Thailand, despite export bans. Sam Bresnick, a research fellow at Georgetown’s Center for Security and Emerging Technology, confirmed that a black market for such chips exists. In May, the founder of Supermicro, a U. S. server builder, was indicted for smuggling advanced chips into China. Bresnick advocated for know-your-customer laws for data centers globally, arguing that companies conducting large training runs on state-of-the-art hardware should face reporting requirements.

The Biden administration proposed such rules in 2024, but no further progress has been made under the Trump administration. Exporters shipping advanced chips abroad are still supposed to ensure they are used only for approved purposes, though enforcement remains unclear.

(Source: TechCrunch)

Topics

model distillation 95% us-china tech tensions 93% export controls on chips 90% open-weight llms 87% industrial espionage 85% reinforcement learning 82% supervised fine-tuning 78% watermarking ai models 76% government regulation 74% chinese ai capabilities 72%