Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • One of the replies:

    > @MehdiKarech

    > I don't remember letting Anthropic or Open Ai scrapping my GitHub, my research gate and all my online writings L O L

    https://xcancel.com/MehdiKarech/status/2080000779859939678#m

  • At least with Github you did. What did you think would happen handing your data to a company owned by Microsoft?
  • Here's a site that asks the same questions to 22 models and compares how similar their responses are.

    https://typebulb.com/u/lab/you-re-relatively-right/full

    According to these results GLM 5.2 is very similar to Google Gemini and Kimi K3 is very similar to Fable 5.

    The American frontier labs are not similar to each other.

  • I think it’s fairly obvious the Chinese labs are doing mass distillation.

    I also think the Fable accusation is wrong and it was most likely Opus 4.8 which itself is likely a distillation of Fable.

  • Interesting. This dooes lend credence to the distillation idea. Good for them!
  • Super interesting. So Fable was really made available... a couple weeks ago? And K3 a few days ago? That's a really impressive feat to distill enough data AND train AND review to get a release that works really well in that time period. Mad props to the Moonshot team :flame:.
  • Fable was launched on June 9 (for 72 hours), then K3 launched on July 16.

    Timeline wise, Moonshot had over a month to post-train K3 on Fable distilled data, which is more than enough time.

  • the distillation everyone talks about in respect to LLM's isn't nearly as easy as most think.

    none of the frontier labs provide probability distributions over the tokens which is the actual method of distillation you use to train a smaller model based on a larger one. they don't even provide all the tokens.

    therefore this so-called distillation the frontier labs whine about is just a set of clever methods to work the existing LLM into the training process for a new model. methods like having the existing model grade the output of the new model and work those grades into the RL method. give the new models structured tasks and use the existing model as a source of truth for those tasks and a myriad of other hacks.

    efficiency scales with the gap between the models and generally allows an efficient bootstrap process. the implication that distillation wouldn't allow further advancement is false however, you can then start doing the same thing the frontier labs have been doing: dumping cash on humans to provide the signals or burning tokens on exploratory paths and grading the results.

    what openai and anthropic don't like is that fact that all the cash they burned can be used to benefit everyone and not just them. and that no matter how much more cash they burn to build up the gap it will closed at a small fraction of the price.

  • Nice! I hope to see continued liberation of these locked up SOTA models. Cloud is a virtual prison, since other people's policies on what they think a user should and should not do cannot be [easily] bypassed, if enforced remotely on a cloud. All digital natives should be skeptical of cloud-hosted services or software. Think like an intelligence agency or sovereign: how are you going to get screwed by the cloud? Your data is fully accessible by the provider, and they can surveil your activities. You probably can't pirate it, so you are a slave in their rentier model. You could be prevented from doing something you want to do, because the provider disagrees philosophically or economically with your desire. You could be stripped of your information/data by a ban due to their policy enforcement system triggering.

    One should live by the maxim: you don't have the thing if you don't possess the file or its processing. That goes for streaming, software, machine learning models, file storage, etc. But I digress; I am happy to see these paternalistic rentiers getting bit by these liberation/copying efforts, and human interests are served every time the digital and infrastructure locks are broken. I will always stand by the distillers!

  • Do you think China is doing this for the reasons you mentioned: liberation..efforts, human interests and freedom of policy?
  • I can understand that the AI labs might care about other labs distilling their models as it can eat into their competitive advantage, but do consumers care at all? Aren't consumers benefiting from this practice by getting better cheaper models as a result?
  • I demand my right to pay 3x more for AI access.

    cf https://www.reddit.com/r/codex/comments/1uyj6pq/kimi_k3_is_1...

  • It’s the same as the patent argument. If everyone could freely copy everything then yes in the short term prices would drop and consumers would benefit, but over the long term it would discourage investment into new technology because a return would be impossible.
  • They're probably going for the national security/domestic manufacturing angle.

    > Aren't consumers benefiting from this practice by getting better cheaper models as a result?

    Aren't consumers benefiting from cheap chinese batteries, EVs, and drones?

  • Commenters are overlooking the significance of this information and posting emotional reactions based on perceptions of fairness or feelings of schadenfreude.

    The economic viability of Anthropic and OpenAI rely on their being able to charge more for model access than their R&D and inference costs. If the market price for SOTA model access drops below that level, then these businesses will have to decide whether to continue to lose money or to reduce spending on R&D.

    Moonshot's papers [1] claim that their training load was primarily from synthetic data and model self-teaching rather than RLHF and therefore keep their costs low. If Moonshot genuinely does not rely on human-led training, they will surpass US closed-source model providers. The United States government considers US supremacy in "AI" as a national security consideration.

    This announcement is noteworthy because it implies that Moonshot's success is in fact due to distillation. It's in the interest of US frontier labs to place barriers to this if they find themselves in the position of subsidizing rival labs' research.

    1. Kimi K2, https://arxiv.org/html/2507.20534v1

  • I doubt distillation had anything to do with it. They barely had enough time. Can you distill Fable (which involves training!) in literally one week? No!
  • It doesn’t matter. Distillation is impossible to stop. They could release an extension that intercepts requests and in return gives you a discount like Honey and get the same data.
  • > The United States government considers US supremacy in "AI" as a national security consideration.

    They thought the same about SSL in the 1990s and the world didn't stop moving elsewhere.

  • People keep forgetting that over the last 6+ months a lot of increased action has been taken by OpenAI and Anthropic to detect and combat distillation. Several are public known.

    Combined with how short of a time Fable was around before K3 got released. I do not see how the data Moonshot is supposed to extract in such a short notice, that will enhance the model to such a point.

    It sounds to me a lot of cope from the US, so they can give this as a reason to ban Kimi models from the market.

    OpenAI/Anthropic their advantages used to be:

    * Early growth advantage

    * Access to a lot of client data to train upon

    * Access to a lot of hardware to train upon

    Several of those advantages have been eroded over time. That barrier has been shrinking. The US is not the only spot with a bunch of smart people (ironical seeing how many Chinese work in US R&D).

    Thing is, even IF they distilled from Fable and got the model so trained up, it means that K3 is a base for future model development. The cat is already out of the bag with how good the model is. When the model gets released on the 27'th, any Chinese company will be able to train their models against K3 openly.

    We are not in the past anymore, where DeepSeek was a unexpected hit, but where the Frontier models their advantages (compute, data, growth) prevented more Chinese models from growing.

  • s/announcement/claim

    You don't get to call Moonshot's a "claim" and this political hack's an "announcement." They're the same thing. Treat them the same. Diction designed to favor one of two equal positions is some weak sauce.

  • I think this is a really sober comment. There are lots of knock-on effects of this claim, even if it's not true which are consequential. The fact that a spokesperson for the US government is going out of their way to comment is concerning.

    Strong bee-hive pinata vibes here.

  • AI companies do not get to play the "Making an LLM using our data is unethical because the resulting LLM will replace us and hurt our profits" card.
  • > The United States government considers US supremacy in "AI" as a national security consideration.

    And we foreigners consider US supremacy in AI to be an existential threat. Your "national security" is directly harmful to us. I never thought I'd say this but the chinese are starting to look like a beacon of hope for the rest of us.

  • "Samuel Slater (June 9, 1768 – April 21, 1835) was an early English-American industrialist known as the 'Father of the American Industrial Revolution', a phrase coined by Andrew Jackson, and the 'Father of the American Factory System'. In the United Kingdom, he was called 'Slater the Traitor' and 'Sam the Slate' because he brought British textile technology to the United States, modifying it for American use. He memorized the textile factory machinery designs as an apprentice to a pioneer in the British industry before migrating to the U.S. at the age of 21."

    https://en.wikipedia.org/wiki/Samuel_Slater

  • Reminds of the quote by Bill Gates.

    > "Well, Steve [Jobs]… I think it’s more like we both had this rich neighbour named Xerox and I broke into his house to steal the TV set and found out that you had already stolen it."

    Source: https://www.goodreads.com/quotes/824084-well-steve-jobs-i-th...

  • That's an amazing quote LMAO

    Wow.

  • So what is the issue here? Distilling is still fair, on the same level like Anthropic scraped copyright protected material for their training.

    So here robbers are blaming robbers?

    These claims are just pointless, everytime

  • ++
  • > So what is the issue here?

    The issue seems to be the US only likes competition when it is winning. Markets in Asia are meant for cheap labor and resources, they're not meant to actually compete. /s

    by Matl
  • I'm by no means taking the side of the AI companies, but it's possible that Anthropic "added value" to the data they harvested. Stealing that does seem kind of uncool.

    Regardless, it was always inevitable—will continue to happen.