Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- I just hope CCP doesn't follow the US government and won't pull the plug before their companies release something on-par with the US frontier models. The question is whether US models not available to the general public will count.
The question is not whether they'll prohibit open-weight models better than the US ones, because we all know the obvious answer.
by zb3 - I wonder if a lot of the companies and governments that seem to think it is essential to be on the forefront of applying leading edge LLMs to the point of starting to become dependent on them are going to find themselves in a situation like that from the Arthur C. Clarke short story "Superiority"? [1] [2].
[1] The story: https://nob.cs.ucdavis.edu/classes/ecs153-2019-04/readings/s...
[2] Wikipedia: https://en.wikipedia.org/wiki/Superiority_(short_story)
by tzs - Article confuses open source models with open weights models.
Not the same thing.
It’s used right in the articles body, but title is misleading.
by samat - I was advocating for "available weight" as a value neutral term for a while.
I gave up. No one cares. And no one will ever tell the truth about the training anyways.
Substantial and growing freedom beats zero freedom ever again.
- Literally no one cares. There are "full" open certified GMO free grass fed training data blah blah models. Apertus, Olmo, etc. No one cares. For all intents and purposes people use the term to describe a model that you can run locally and are allowed to modify and re-release. The rest is useless semantics. No one can "rEpRoDuCe" a model anyway.
- Interesting to consider this inline with recent us export bans, could the US be squandering its lead by giving the open source, largely Chinese labs catch up (in terms of model quality available to masses), will US labs be able to maintain the lead without users being able to use their latest models?by gehsty
- Why do you think this matters? Not that it does or doesn't but what quality does "US WINS" or "CHINA WINS" bring to the table?by ggm
- If the belief that open-weight/Chinese models depend significantly on distillation of the latest frontier models is correct, then presumably the gap will stabilise to the minimum time required for extraction of meaningful data (from the latest frontier model) plus finalisation of training of the latest dependent model. This gap can be minimised by increasing the process efficiency, but can't be eliminated entirely. (Attempts to hinder distillation from Anthropic/OpenAI may shift the balance too.)by mft_
- USA, a country that known for the land of freedom, is now restricting frontier models to the point where non-Americans cannot even use them.
China, a "authoritarian state" country, "the antonym of freedom", with a software industry that is especially capitalist, has produced all the competitive open-weight models.
It really is IRONIC.
Disclosure: I am Chinese, and I understand this strategy comes from being behind, using open source as an asymmetric way to compete and make up for missing compute by sharing the burden, etc. But still, very ironically.
by linzhangrun - Your comparison falls apart in the first few words:
> USA, a country that known for the land of freedom
The US might say it's the land of freedom, but it's been playing the game of economic protectionism for centuries. This is just the latest example.
by mft_ - It would be interesting to know how much of a boost the closed models companies are giving the open models.
If the closed models stop improving will the progress of open models slow?
by jacobgold - > It would be interesting to know how much of the "distillation" boost is helping the open weight models keep up.
Some people in China surely know.
> Like if the closed models stop improving will all the closed models also stop improving?
Seems extremely unlikely, unless the models all hit some kind of wall soon. The Chinese companies may be behind the US in compute capacity, but they have excellent researchers [0] who are probably approximately as good as their US counterparts at the kind of problem generation and RL that is currently working so well.
I would be very surprised, though, if the models cannot continue to be improved rapidly in any area that allows a tight feedback loop like programming, at least up to the point where we puny humans lose the ability to define objective functions.
(And, conversely, I don’t expect magic in fields where the feedback is slow or expensive. A model is not about to reliably invent a wonderful medicine for the same reason that a large and extremely competent pharma company cannot: the evaluation process is extremely slow and it’s so expensive that the kind of utterly enormous corpus that is driving the current progress in coding is simply not available. Running RL on m iterations of n medication-development trajectories each is going to cost n*m times $10-100 million and take m years if it’s even possible at all.)
[0] The US advantage in this space will likely decline, since the brain drain from the rest of the world via the US university system to US labs is drying up.
by amluto - Why are we assuming only American labs can innovate? DeepSeek already innovated a lot in efficiency, for example.by amunozo
- > What is notable is that a large amount of the total improvement of models has been in the coding benchmark. The coding index has gone from 15 months behind to only a month or two behind
This makes sense, right? Coding is one of the most obvious short-term uses of models, it also has a readymade market willing to pay a lot for tokens, it has a huge corpus to work with, and a significant degree of validation is built into the problem domain...
by swiftcoder - I haven’t seen it discussed anywhere that closed models can essentially cheat benchmarks right? What Anthropic or OpenAI brand as a model doesn’t necessarily have to be just weights, it can be a whole backend system that augments the model itself. With this they can score better benchmarks than an open source model that is weights alone.by cedws
- Good pointby snthpy
- Sure, I think that's fine, that all counts. It counts for open source too, it's not like they're somehow running these benchmarks without any harness.
Nobody cares if your AGI is 100% made out of neural networks or if it's like 50% neural networks and 50% perl scripts.
by jstanley - > Now is probably a good time to liquidate your pension, fly to a remote island somewhere, and live out the remaining 6 months or so of civilization in peace.
> So maybe the open source apocalypse won’t happen yet.
Sorry I wasn't at the last doomer meeting, when did we decide good open source models are a harbinger for the apocalypse?
by taffydavid - Doomerism is at all time high
People becoming more and more neurotic by the day
by danlugo92 - this is a blog post from a company that hosts open weights LLMs (https://www.doubleword.ai/). I think its possible it might have been tongue in cheekby somnial
- Cute. Climate change’s apocalyptical impact on food crops and cancer rates (post-ozone collapse) never convinced people to enact change.
But hey, it’s open-model LLMs, the boogeyman! Can’t have that, it must be OpenAI or Anthropic safely controlling the market and calling all the shots.
by port11 - I assumed it meant that when open weights reach the capability of frontier models, and tounge in cheek referencing the terrible consequences of us all getting our hands on mythos+ capability models without restrictions.by alienbaby
- If anything open-source models are a hedge against the apocalypse. Or at least against the cyberpunk dystopia.by kageroumado
- The Chinese models will not overtake the frontier US ones given the current way things are going. The US models derive their lead from incredible efforts to source more and higher quality (mostly synthetic data) via great feats (eg generating with humongous teacher models that could never feasibly serve interactive traffic). The Chinese models advance via heroic efforts to optimize models and great feats to secure more and higher quality training data from the US frontier models.
For an (Chinese) open weight model to surpass the (US lab) frontier models, this equation must flip and the Chinese labs must entirely retool from harvesting frontier model data to producing the data systems and efforts to produce novel data; as well as procuring latest generation hardware en masse for this. This does not happen easily. Also training a frontier scale model is actually not such an unimaginable feat: doing all the inference with the teacher models is where the hardware goes.
by christina97 - I don’t think anyone seriously believes any of the Chinese models are ever going to “overtake” the American frontier models. I doubt that that’s even their goal.
But if they can stay on pace, within say 6 to 12 months of the bleeding edge of the American frontier models, that’s a huge problem.
If they can just piggyback on the Herculean efforts of Anthropic, OpenAI, Google etc., accept a little bit of lag, and save billions of dollars? Why wouldn’t they?
And for the end user, why would they pay a premium subscription price for something they can just wait six months for and run on their own hardware at home? In my opinion, this is the cat and mouse game that’s being played right now. And I suspect it’s intentional on the side of the open weight models. I would bet they are playing a war of attrition
by 40four - Yeah, this is, to be perfectly blunt, cope, for several reasons:
1. It's unclear if there is a law of diminishing returns with ever-larger models. They're more expensive to run and for many applications, you'll probably find smaller models are sufficient;
2. There's an inbuilt market for local LLMs. This is an effective limit on how large models can get. Case law hasn't been established yet on, for example, if a law firm using ChatGPT breaks privilege. Specifically, chat logs may be discoverable. Medical applications have this issue too and I think you'll find that financial firms are going to be leery about this as well;
3. Better, larger models will bleed into smaller, open source models. The chat logs themselves are training data. There's a whole market in China for Claude tokens around this;
4. China has a national security interest in not being beholden to US tech giants when it comes to AI. China has a history of being able to commit to large-scale long-term projects and Anthropic just won't be able to compete with a national project by one of the world's superpowers, if it comes down to it;
5. Winning doesn't necessarily mean being the best. Often it's just being good enough;
6. As an example of a national project, China is busy replicating EUV because of the US ban on ASML and NVidia exporting their best stuff. I don't think many in the West are prepared for how rapid this will be. I'm reminded of the policy debate in 1945 when many in American policy and militarey circles thought the USSR would never catch up with atomic bomb or, if they did, it would take 20+ years. It took 4 years. For the hydrogen bomb, it took 1. The US hardware advantage is a lot more tenuous than many realize.
by jmyeet - How so? You'll soon have your choice of a very old OAI model or a new Chinese model, because the USG has no interest in letting you access the newest models without explicit permission.by kulahan
- Chinese frontier models don't need to catch up in every category. They just need to win in coding and that's exactly where they are going. The gap went from 12+ months to 1-2 months with the latest release of GLM 5.2 and coding is a task that you don't need heroic efforts to find rare and long-tail training data, you can just outsmart your competitor by optimizing algorithms and training recipes. This is something they can do at scale with the money and talent pool.by elisbce
- “China can only copy the US” is a very short sighted and uninformed opinion. there is more coming out of china than just new ways to distill modelsby bradishungry
- The amount of data Anthropic has claimed was extracted for distillation is tiny in comparison to the entire internet, which is right there for the taking and holds most of the knowledge people expect models to have.
Distilling even with small amounts of data from a better model is still helpful, but not in the sense of transferring capabilities the raw internet-trained model doesn't have at all, but for identifying those capabilities that are compatible with the servile assistant persona and suppressing others that are undesirable (e.g. trolling). A primitive version of this were instruction-tuning datasets generated with ChatGPT, as used e.g. for Alpaca.
Without a clear target to emulate, competitors might have to rely more on human raters, but there are plenty of data labeling companies in China, so that's hardly a hurdle.
by yorwba