

Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- > if you run one model, run glm-5.3
That is a horrible take-away from this, with only 28 tasks and a high pass rate for most models, it says almost nothing.
Test a model for your use case and use the fastest, smallest, cheapest model that 100% satisfies your use case.
Or, if you truly do need a model with strong generalized performance, definitely do not take a benchmark like this serious with such a limited task set.
by CMay - For the last week, I've been heavily immersed reverse engineering a device with help of GLM-5.3 and it surpassed all my expectations - I actually managed to achieve very way more than I thought I would. I never worked on such low level stuff, it would have taken me months to learn ARM assembly and how to find for and write exploits. Initially, I attempted this with Claude, but it blocked me on the very first message, so I got a refund and decided to try z.ai. The only downsides are that it's maybe a bit slower than my day job Opus and I had to pay ~200 EUR for a monthly plan in order not to bump into weekly limits in a couple of days. If this level of capability cost maybe 50 EUR, I'd strongly consider getting a long time subscription.by Klaster_1
- Not sure I trust a benchmark where Haiku gets a nearly perfect score and Fable is tied for last placeby ac29
- IDK I use open models every day for personal projects, and closed models for work.
Open models are all decidedly far behind Fable and a good bit behind Opus as well. All of these posts read like motivated/wishful thinking to me.
I get that people badly want the open frontier to be where the closed frontier is, but it is just so obviously not the case if you actually use the models on a real project.
by solenoid0937 - There's no way GPT 5.5 is better than GPT 5.6 Sol, and gemini-3.6-flash is better than both of those. I wouldn't trust this benchmark at all.by Fe2O3
- One problem is that $30 per run is really noisy for many verifiable tasks. Ours usually run at least an order of magnitude more for a model in GLM 5.3's price class.
We just published GLM 5.3 results on our multi-agent coding evaluations and it's definitely an impressive model, coming in around #6. For the price, it's actually not Pareto optimal, falling slightly behind Grok 4.6 (which is faster and the same price) and Sol 5.6 (which uses far fewer thinking tokens for comparable results). As with most Chinese models, it excels at iterating in a harness while its first answer/base fluid intelligence is below the American frontier.
Data at https://gertlabs.com/rankings
by gertlabs - Almost 10 models are passing the benchmark >95% - isn't that... substantially overly saturated?
I'm also really skeptical of benchmarks that place any Haiku model very high. I've been thoroughly unimpressed with Haiku and my opinion has been that you basically shouldn't use it. Yet here, it ties Kimi K3. HMMMM.
"Rubric Quality" seems a bit more realistic than "Pass Rate", but it is apparently judged by Fable 5...
This entire write-up is also obviously clearly very heavily AI-assisted, which doesn't help matters any.
by jchw - This whole thing immediately reads as Claude generated, making it hard to take seriously.
Why do these results contradict existing serious attempts at benchmarking LLMs? Namely: https://artificialanalysis.ai/ https://arena.ai/leaderboard/agent
by hellohello2