Ben Thompson, on the hype regarding Kimi K3, at Stratechery:
This is a point that bears repeating: because U.S. open weight
model makers must follow the frontier labs’ terms of service, they
(1) are worse than Chinese alternatives and (2) end up distilling
the distillation, just with a detour through Chinese labs.
Wouldn’t it be better if western open weight model makers could go
to the source?To that end, here’s an even more interesting question around
distillation: why exactly is it bad? After all, what are large
language models but the distillation of all of the knowledge on
the open Internet, scraped by the frontier labs and distilled into
the models that are themselves being distilled? Who is exactly
being wronged here?In fact, this paradox is the solution. I believe that open weight
models are good for innovation (and, per the above, I think that
labs on the frontier will be fine), but it’s a problem to be
dependent on China. The U.S. should pass a law that (1) makes
explicit that collecting data for training models is fair use, and
(2) bars terms of service that forbid distillation, for U.S.
companies at a minimum. Stopping distillation — which is
literally just querying the API — is nearly impossible; the U.S.
should go the other way and lean into a new copyright policy that
both indemnifies the labs and also guarantees that what they
learned fuels further innovation for everyone else.
To be clear — and Thompson emphasizes this point too — the leading Chinese models like K3 aren’t good only because of distillation. They’re not entirely rip-offs. But distillation is clearly an essential part of the formula they’re using to keep releasing models that are 6–9 months behind the U.S. frontier state of the art. So the Chinese treat all models as “open”, regardless of the terms of service. But anyone in the U.S. or other western countries that respect copyright who wants to distill a good model has to wait for the Chinese to release their models, like K3.
And that second paragraph I quote above distills (sorry) exactly why I want to bring out the world’s smallest violin to play a sad song for OpenAI and Anthropic regarding their objections to their frontier models being distilled against their terms of service.
