Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • > So the Chinese labs have thrown a lifeline to the Western loss-making labs, and I just have no clue as to why.

    "Thus the expert in battle moves the enemy, and is not moved by him."

    They figured out a clever method for avoiding excessive training costs via distillation. That forces the hand of frontier labs to move faster, produce better models, etc. (to avoid embarrassment and 'falling behind'—all the while shouldering most of the cost), which they can just keep distilling—or applying other techniques against—much to the dismay of said frontier labs.

    Checkmate.

  • Performance optimizations don't just help the western labs, they also help with running more powerful/useful LLMs locally.

    If local LLMs get "good" enough, people will soon paying for subscriptions to ChatGPT and Claude, which hurts their revenue.

  • “So the Chinese labs have thrown a lifeline to the Western loss-making labs, and I just have no clue as to why.”

    I don’t think it is intentional but this is actually quite bad for the western labs.

    The entire booster narrative has been “look at how their revenue is growing! $10bn to $100bn ARR in under a year! This’ll be a multi-trillion IPO!” and the extrapolated future growth from $100bn to $500bn and $500bn to $1tn justified future investment… but that revenue was just because inference was expensive.

    The revenue growth story is all that matters pre-IPO. If revenue falls from $100bn to $50bn that’s very very bad optics for OpenAI and Anthropic even if they are now profitable, it completely destroys the growth narrative.

  • I'm beyond thankful that Chinese AI models are so good. I desperately want us to cure the myriad of maladies that humans suffer needlessly with on a daily basis. We're going to need more powerful models than we have now if we're gonna do that and the Chinese are providing the competition needed to push this thing as fast as we can.

    I realize "going as fast as we can" is not the most popular position atm. But I'm far more interested in what good we can do than 10% apocalypse scenarios. I volunteer with a charity for childhood brain cancer and I do not want to see another 4 year old die. I'm willing to risk anything to stop this.

  • I've been running Qwen 3.8 27b (an opus 4.6 tier model), locally on a 5090 for just over two weeks @ 170 tokens/s. That's a frontier model from 9 months ago running on consumer hardware. Who knows where distillation and pruning gets us in another year.
  • The best explanation is that it's a goal of the CCP to generally commodotize LLMs, because LLMs will ultimately be a compliment to manufacturing (which China dominates), and you always want to "commodotize your compliments".

    I think this explains why they are open sourcing broadly. It's not to be nice. It's a strategic play by the Chinese government to help ensure there are many players in this race and not too much power accumulates to American labs (even if American labs benefit in the process)

  • I'm also grateful to the Chinese labs for providing workarounds for the walled gardens that the US based AI companies are attempting to create.

    Does anyone know if there are any distillation datasets available? I'd love to see these distributed on BitTorrent. I think it's critical that AI be democratized and not isolated in the hands of a few private companies.

  • Why is it a problem that the Chinese labs are just distilling down Anthropic’s models? Aren’t Anthropic’s models not just distilling down other people’s work?

    Feels like Anthropic crying do as I say not as I do.

Explore Birbla archives