Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- Looking forward to seeing the stats.
I gave it an abandoned repo for an Aseprite MCP someone made and told it to iterate with a laundry list of things I wanted from it to include thousands of plugins.
Came back 20 hours later and it shit out a pretty surprising little tool, will post the public repo when I get time.
by itsryanlenk - Funny how all china companies are expected to release weights by defaultby seydor
- Well, China is at their usual barely legal (at least if the WTO would be worth even a bit after it got thoroughly gutted) game, just burning their effectively infinite cash reserves to undercut Western providers and eventually force them out of business.
In the end there's barely any moat that any of the ludicrously "valued" AI companies have - the only thing justifying the valuations of SpaceX/xAI/Tesla/Anthropic/OpenAI is a supposed "secret sauce" that, frankly, barely exists any more.
Everyone and their dog can go and run LLMs for dirt cheap on their own hardware.
by mschuster91 - they are playing a completely different game than the US
- Xi has made it official policy, see his keynote speech at their World AI Conference last month:
> We should seize this rare, historic opportunity to encourage open source, openness, collaboration and sharing. [1]
People have pointed that this seemingly made Alibaba/Qwen turn around from closing their models (this was rumored after the shakeup early this year [2]) and release the weights for even the Max variant of their new models, which they previously did not.
1: http://english.scio.gov.cn/topnews/2026-07/18/content_118605...
by square_usual - All smaller models and models behind frontier are expected to be released by default. Otherwise there’s no reason to produce them.
Chinese labs are not releasing all of their model weights. Qwen is known as an open weight model by most, but their top model is not open weight.
Releasing weights is a marketing strategy for newer labs to get their brand out there.
by Aurornis - by garo-pro
- Unfortunately I can't find sources other than this for now but this seems to be legit.by garo-pro
- they have confirmed it officiallyby KaseyKim
- > The company on Wednesday confirmed speculation that the Ox Alpha model is a new iteration of its GLM series and said it will release the weights for it tonight, in response to queries by Bloomberg News.
Seems legit.
It's really hard to know how good it is. So much hype around it.
by mohsen1 - It writes pretty clean code and holds context alright, but it starts stumbling and losing the plot on complex bash scripts with pipelines. Waiting for the weights to drop so we can dig under the hood and see what is going on thereby RataNova
- It one-shotted generation of Java bindings for this project: https://github.com/jeffhajewski/latticedb
Related PR: https://github.com/jeffhajewski/latticedb/pull/5
The session used ~100K input tokens, ~60K output tokens, and ~80K thinking tokens.
I reviewed it using gpt-sol-medium, and it seems to be satisfied with it's work.
by freakynit - Releasing weights is the right move. Keeps them competitive with DeepSeek on the open side.by harlan_pdx
- I had good experience with GLM 5.3, but...
Z.AI is the only provider for GLM 5.3 on OpenRouter. I don't see 5.3 on Hugging Face. Not sure if this new model is "full GLM" or something smaller, or if they will like Moonshot AI publish weights but put restrictive license [1], which will again leave Z.AI as single GLM model provider on OpenRouter.
[1] https://huggingface.co/moonshotai/Kimi-K3/blob/main/LICENSE
by stanac - I'd be interested to know what was going on with it during the public test as there were numerous reports of it improving considerably at tasks it was asked to do early on in the test compared to later in it.by esskay
- lol don't shout out the obviousby rfoo
- Two potentials from my pov:
1. Just variance in pass@K. If you prompt any model multiple times you'll see a large variance. N=1, but I find chinese open source models have a higher variance than higher-RL'd models like fable/opus.
2. They legitimately shipped a new RL checkpoint over the 7 days, which I find hard to believe.
I am leaning towards 1.
by daveyoung - It's logical to serve the best version (quant) of the model at the beginning so that users keep testing it. It is also reasonable to think that the developer of the model tried to test various quant levels by gradually degrading the model's capabilities.by utilize1808
- Ox Alpha has been running on auto-pilot for the past 5 days on various experiments.
Very impressive model.
Here are some examples, open-source documented and the data available in HF datasets:
https://openzot.github.io/whetstone/ - https://github.com/openzot/whetstone
https://openzot.github.io/arcade/ - https://github.com/openzot/arcade
https://openzot.github.io/machinery/ - https://github.com/openzot/machinery
by _pdp_ - Mixed signals, here it's performing below even GPT-5.4 Nano:
while here it outperforms Fable by a significant margin:
but if the latter is true, will people still say it was "distilled" from Fable?
by WithinReason - GLM 5.3 was a great model, so this would be strange to release a regressed modelby epolanski
- omp+0x-alpha beat both cc+fable and codex-sol in creating/refactoring a big eval setup. the former just knows where things should belong and completed the task all the way while the other two failed on both metrics.by daralthus
- Kinda useless to compare simply based on model without considering harness. Different agents handle the context etc completely differently. I would like to start seeing these model vs model comparisons across different harnesses.by dpweb
- I've been having oxa and sol do architecture design then compare notes. Sol is definitely still way ahead. But there's reliably some really good wins ideas and concepts that OxA throws out there that Sol is very happy to encorporate.
One thing that I think matters a lot for the non developers, all three of us (sol, oxa, and me) usually agree that oxa's write up is far far better. It explains the situation very well, and has great structure for its write ups. Sol gets the job done, but it's terrible at re-explaining the problem for humans, at laying out information. It also doesn't show it's thinking, so it's imo a terrible peer to work with!
- I really want to see hard evidence of distillation before I buy into it. Seems like a lot of sour grapes over not having the sort of lead assumed. In this field, it has been shown repeatedly that leaps in performance come swiftly and without notice.by tescreal
- the 2nd website is not official, just something someone slopped together for some reason.by sunbum
- That benchmark is super sus. Until someone pointed it out, the top performing open weights model was a Kimi K3 fine tune from their sponsor (abacusai/Smaug-Agentic). Now, it's not on the list.
Source: https://twitterwebviewer.com/?tweet=2091116504787935350
- Claims about Ox Alpha performing at Fable level were from the social media hype cycle. Everything new in the LLM space brings a wave of influencers hyping it up as a revolutionary leap forward. Don’t forget to like and subscribe to learn more.
It is a capable small model, but it’s not frontier level. The interesting part will be seeing the model size, how it responds to quantization, and how fast it runs on the kind of non-server hardware that we can buy without selling a kidney.
by Aurornis - by giamma
- thanks for pointing me out on Unwall.App!
Didnt know they exist - looks very good, maybe even better than Archive.ph