Experts Push Back on White House Claims That Moonshot's Kimi K3 Was Distilled From Anthropic's Fable

AI researchers are publicly questioning White House science advisor Michael Kratsios's assertion that Moonshot built Kimi K3 by copying Anthropic's Fable LLM using chips not cleared for export to China. On July 23, 2026, TechCrunch reported that multiple experts consider the distillation claim technically implausible given the timelines involved (TechCrunch).
Kratsios characterized the alleged copying as "Large-scale, covert industrial distillation aimed at stealing proprietary U.S. technology and undermining American research." He did not share details about the sources of his allegations. Moonshot did not respond to TechCrunch's questions about its training process.
The claims come amid reported discussions within the U.S. government about banning Chinese open-weight models. Treasury Secretary Scott Bessent said "we are finding watermarks of our U.S. large language models on many of the Chinese models, and that that's unacceptable." The Treasury Department did not respond to TechCrunch's query about what those watermarks consist of. Treasury had separately threatened sanctions after the White House claims regarding Moonshot's distillation of Fable (TechCrunch).
The technical pushback from researchers centers on timing and feasibility. Braden Hancock, a researcher at the Laude Institute and co-founder of Snorkel AI, told TechCrunch he does not think Kimi K3's strength came from strictly distilling Anthropic's Fable. Hancock noted that Fable was only publicly available since July 1, 2026, leaving insufficient time to distill data, train a model, and release it within two weeks. Nathan Lambert, an AI researcher at the Allen Institute for AI, said distillation has become less impactful over time as Chinese models get closer to the frontier and training shifts to reinforcement learning. Lambert argued that if distillation were the key factor, competitors could catch up to GLM or K3 using its data for distillation, but that has not happened from supervised fine-tuning alone. He also said that distilling Fable-like capabilities would likely require reinforcement learning techniques with tens of millions of agents, which would be prohibitively expensive and a time bottleneck when using a frontier lab's API.
The broader dispute builds on a longer-running conflict. Anthropic publicly accused Moonshot, DeepSeek, and MiniMax of systematically distilling its models earlier in 2026 (Anthropic). In a February 2026 post, Anthropic warned that if distilled models are open-sourced, risk multiplies as capabilities spread freely beyond any single government's control. A subsequent June 2026 Anthropic post on Claude Fable 5 and Claude Mythos 5 stated that distillation of Fable 5's abilities could indirectly lead to the proliferation of near-frontier AI capabilities (Anthropic). Anthropic's research publication "2028: Two scenarios for global AI leadership" discussed how Chinese labs have remained close to the frontier by exploiting U.S. export control loopholes and carrying out large-scale distillation attacks (Anthropic).
Kimi K3 was released on July 16, 2026. The 2.8-trillion-parameter model is natively multimodal, features a 1-million-token context window, and is designed for long-horizon coding and end-to-end knowledge work (Moonshot AI). The Kimi K2 series was officially discontinued on May 25, 2026. Moonshot has stated that Kimi K3 "still trails the most powerful proprietary models, Claude Fable 5 and GPT 5.6 Sol" (TechCrunch).
Moonshot stated that full model weights for K3 will be released by July 27, 2026, and that the company will publish more details on the model's architecture, training methodology, and evaluation (Moonshot AI).
There is a policy dimension to this technical dispute that bears watching. The U.S. government is weighing restrictions on Chinese open-weight models, and the Kratsios allegations surfaced in the context of those discussions. The technical merits of the distillation claim matter because they underpin a proposed regulatory action that could restrict access to open-weight models broadly, not just those from Moonshot. If the distillation pathway described by Kratsios is not feasible within the available timeframe, as Hancock and Lambert argue, the evidentiary basis for a ban weakens considerably.
What remains unclear is what Bessent meant by "watermarks" found on Chinese models. Without a technical definition from Treasury, the term could refer to anything from residual artifacts in model outputs to provenance markers deliberately embedded in training data. The distinction matters. A deliberate watermark embedded by Anthropic and subsequently detected would be materially different from indirect behavioral signatures that could arise from convergent training approaches on overlapping datasets.
The open-weight release of Kimi K3 on July 27 may shed additional light. If Moonshot follows through on its commitment to publish training methodology details, independent researchers will have an opportunity to evaluate the distillation claims directly. Until then, the allegations rest on government assertions that have not been substantiated with technical evidence, set against expert skepticism from researchers with direct experience in distillation techniques.


