Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- I ran the same Mac SVG drawing prompt through GLM 5.2 and 5.3 across every reasoning effort level, and 5.3 showed improved performanceby sumedh
- Tied for #1 by agentic index (with Opus 5).by colingauvin
- Still yet, I cannot justify switching from dirt-cheap Luna model, which is pretty damn "intelligent" and works well for my flowby yipinwong
- Very impressive score for the size, though token use is higher than k3 and far higher than proprietary models, and its price to performance isn't all that far ahead of k3 as a result
- >token use is higher than k3 and far higher than proprietary models
GLM sets effort to max by default historically.
by Havoc - Does Artificial Analysis use OpenRouter for model access to do their benchmarks?by swingboy
- Is it worth using these models if I have a claude code subscription already? The appeal of lower cost is nice but I haven't gotten over the switching cost yet.by Zaheer
- no, at subscription prices claude is a better value than GLM.
They're only a better value if you're paying API rates
by notatoad - If anything, it's going to be more expensive. Price/performance ratio isn't there yet for frontier open weight models.
But regardless, you definitely should use a harness where switching models on the fly is easy. There's a reason why Anthropic uses their own proprietary formats/conventions anywhere they can - to lock you in when inference eventually commoditizes.
by glub - At least by API usage, they aren't yet lower cost than subscriptions. Not sure about GLM's subscription plans though.by colingauvin
- Yes. Please seriously try other models. See relevant thread here: https://news.ycombinator.com/item?id=49296740by karimf
- I use the $200 plan w/ Anthropic and run out of tokens half way through the week and supposedly they are progressively reducing the limits on all their subs even further.
At some point I will switch, $200 buys a lot of tokens on OpenRouter.
by oceanplexian - Use a unified proxy that lets you switch between models seamlessly. We are far from an equilibrium in this market and you will continue to have FOMO no matter who you pick if you go all in on one companyby culi
- FYI, you can use your Claude subscription pricing with OpenCode via Meridian[0], which also makes it easier to try out other models when they come out. You can also use your other subscriptions in OpenCode with CLIProxyAPI[1]. The switching cost was relatively high, mostly from claude code plugins but completely worth it. I'm now mostly using GLM-5.3 and Codex models via OpenCode and barely using Claude which seemed unfathomable less than two months ago.
[0] https://github.com/rynfar/meridian
[1] https://github.com/router-for-me/CLIProxyAPI
edit: reworded for clarity
by robertn702 - I understand that running these benchmarks can get expensive, but it would be really nice to see AA include more benchmarks of models at reasoning settings other than the maximum, at least for the biggest releases. They have that nice graph of cost vs. composite benchmark score with the Pareto frontier line, but who knows if those are actually the optimal choices? There are already a few non-max-reasoning models on the Pareto line, among the few that were tested.by AnodicElegy
- You can turn on various levels of some of many of the models in the UIby apitman
- Beware of the benchmarks listed. SciCode and EnterpriseOps for instance: https://shukla.io/blog/2026-08/gym.htmlby BinRoo
- The Chinese models also like to cut corners on stuff like science. Their scores on stuff like biotech and scientific knowledge is far from ChatGPT unfortunately. (Claude is pretty good but it just refuses all prompts).by Onavo
- Sol is an underappreciated model. Dropped Claude today and went to codex. None of that god awful prose Claude used for me any longer.by Escapade5160
- I've tested GLM 5.3 on the release day and Artificial Analysis is spot on. It's a really good model.
But my main takeaway was something else. I've used closed weight models for long enough that I've forgotten how good it feels to see reasoning tokens.
With GPT/Claude, you kind of hope that intent was captured well, that agent had all the information, all the tools it needed, because you won't see "hmmm it seems like nix flake isn't available here and I shouldn't install something globally" until it slopped out millions of tokens and wasted hundreds of dollars for 8 hours. With GLM and the likes, you just stop the disease right where it begins.
by glub - Generally, are closed sourced models hiding their traces? I was making an agent to develop and deploy apps and fed the traces to dispel time-consuming detours and made it a few times faster.by aitchnyu
- With GPT/Claude, hiding those from users to waste their tokens is a feature, not a limitation.by tw1984
- Yes, not necessary often but being able to stop something that is going off the rails is super useful. Especially if the root cause is prompt ambiguity - inject a clarification & it recoversby Havoc
- I like to compare models with a similar score on cost per task and output tokens per task since those measure two things I'm interested in: cost efficiency and token efficiency. Here's how GLM-5.3 compares to other models in a similar score and against GLM-5.2 to save a few clicks for others who care about these metrics:
Edited for accuracy and more models.Model Score Cost / Task Output Tokens / Task ------------------------------------------------------------------------- GLM-5.3 (max) 59.5 $0.68 41,107 GLM-5.2 (max) 53.0 $0.56 32,200 Claude Opus 5 (high) 61.5 $1.52 21,353 GPT-5.6 Sol (max) 60.9 $1.23 16,879 Grok 4.6 (high) 60.9 $0.84 21,735 Kimi K3 (max) 59.7 $0.84 25,474 GPT-5.6 Sol (xhigh) 59.0 $0.87 11,098 Claude Opus 5 (medium) 58.6 $0.98 12,459 Qwen3.8 Max 58.1 $1.13 38,287 Qwen3.8 2.4T A95B 57.7 $0.95 32,472 Claude Opus 4.8 (max) 57.3 $1.65 33,557 GPT-5.6 Sol (high) 57.3 $0.52 7,545 Muse Spark 1.2 (xhigh) 56.8 $0.40 30,430 GPT-5.6 Terra (max) 56.6 $0.51 20,838 GPT-5.5 (xhigh) 56.3 $0.69 16,893 Gemini 3.7 Flash (high) 56.0 $0.40 36,847by scotttrinh - these $/task figures aren't very useful in my experience. it doesn't tell you how well it did the task.
generally I choose models by their intelligence and then personal preference from direct experience.
by teravor - this is not very useful.
for over 1 billion real world users living in China, they don't have the option of paying $1.52 per task to use Opus 5, they are banned doing that due to US politics.
by tw1984 - It would make reading and comparing a bit easier if the data was sorted by a dimension.by dudeinhawaii
- Muse Spark has a nice balance. not to mentions the Contribs version is old deepseek flash prices.