Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • The GLM-5.3 weights are not yet open, and they've said that the delay is due to the need to nerf them for "safety."

    So I have a feeling a lot of these early claims are not going to pan out in the long run.

  • I'm not super worried about "safety" nerfs. Most refusals can be finetuned out, it's usually not a dealbreaker for open-weight models.
  • The stealth model "Ox Alpha" has been crushing benchmarks and appears to be the next release in the GLM family
  • ai;dr

    can't take any generated benchmark seriously. if you produce actual results, then produce actual copy to go with it.

  • my prediction is even if open source Chinese models are 90% as good (or even a bit better, which I don’t really believe because of benchmark hacking) enterprises will still pay for Claude / ChatGPT and the harness, integrations, and peace of mind versus using some Chinese cloud.
  • In the real enterprise world companies are running their processes writing Gemini "gems" or using copilot because they were already google/Microsoft customers.

    Am I the only one that knows people in industries like insurance, banking, consultancy, materials, etc? Cause none of them gives two damns about what the leading SOTA is, procurement and compliance matter.

  • > versus using some Chinese cloud.

    Just use Openrouter.

  • On a tech level I’d say that Kimi and GLM 5.3 on Max reasoning are good enough for non-trivial planning and exploration and on High are good enough for various implementation tasks. They can easily replace Opus 5 for me and mostly even Fable (webdev with some ML and DevOps work on the side, as well as local software).

    All of that pretty much means nothing for the orgs that just want to do the AI equivalent of picking IBM.

  • peace of mind versus using some Chinese cloud

    These are open weight models (GLM-5.3 soon too). You can run them on the Together AIs or Firework AIs of this world. Use OpenRouter or HF Inference Providers in between and you can effortlessly switch between models and providers.

    I have been using GLM and Kimi models the last few months mixed with the latest Anthropic models and for my daily work there is barely a difference anymore (except for pricing).

  • Enterprise's peace of mind is being able to use the model and not have the US government decide on a whim to block access.

    Additionally I may want to run attack simulations which requires the removal of safeguards. My only option is to use an open model I can run on my own hardware.

  • > if you run one model, run glm-5.3

    That is a horrible take-away from this, with only 28 tasks and a high pass rate for most models, it says almost nothing.

    Test a model for your use case and use the fastest, smallest, cheapest model that 100% satisfies your use case.

    Or, if you truly do need a model with strong generalized performance, definitely do not take a benchmark like this serious with such a limited task set.

    by CMay
  • The point I'm making is that most models are good enough for most tasks, so choose on speed/cost.

    Definitely benchmark on your own tasks. GLM-5.3 is the winner on mine, on yours maybe not. I am not trying to be a universal benchmark, as these serve no one but the person doing the benchmark.

    My previous post sank like a stone, but all the evidence + code to run this + what you have todo to adapt it for your own use cases is all here https://github.com/ed-is-ai/featherbench

    Encourage everyone to eval like the devil

  • For the last week, I've been heavily immersed reverse engineering a device with help of GLM-5.3 and it surpassed all my expectations - I actually managed to achieve very way more than I thought I would. I never worked on such low level stuff, it would have taken me months to learn ARM assembly and how to find for and write exploits. Initially, I attempted this with Claude, but it blocked me on the very first message, so I got a refund and decided to try z.ai. The only downsides are that it's maybe a bit slower than my day job Opus and I had to pay ~200 EUR for a monthly plan in order not to bump into weekly limits in a couple of days. If this level of capability cost maybe 50 EUR, I'd strongly consider getting a long time subscription.
  • > z.ai

    > The only downsides are

    The catch's in their revolting terms of service.

  • GLM 5.2 and 5.3 are exceptional at reverse engineering. you just point them at an IDA Pro MCP and off they go doing whatever you want from them.
  • Not sure I trust a benchmark where Haiku gets a nearly perfect score and Fable is tied for last place
    by ac29
  • I've been not just unimpressed by Fable, but actively find it to generate negative value.

    It hallucinates more, and in more destructive ways, than other models I've worked with and generates truly atrocious jargon and bizarre inhuman explanations that end up cluttering things. The code it writes is terrible too. Overly complex with a lot of technical debt.

  • Fable got heavily beaten down by its refusals, which is not too surprising; although a couple of problems got refused for reasons I can't even imagine and the page doesn't quote the refusal.

    Some of the other failures like the colicky baby one are also probably soft refusals, it's not clear what the grading criteria are but I'm guessing it got docked for not going anywhere near a possible diagnosis.

  • IDK I use open models every day for personal projects, and closed models for work.

    Open models are all decidedly far behind Fable and a good bit behind Opus as well. All of these posts read like motivated/wishful thinking to me.

    I get that people badly want the open frontier to be where the closed frontier is, but it is just so obviously not the case if you actually use the models on a real project.

  • There's no way GPT 5.5 is better than GPT 5.6 Sol, and gemini-3.6-flash is better than both of those. I wouldn't trust this benchmark at all.
  • benchmark is saturated but your claim that gemini-3.6-flash is better than 5.6-sol is also not trustworthy or accurate.

    maybe if you mentioned 3.7-flash it might have been slightly more believable (but still false).

  • One problem is that $30 per run is really noisy for many verifiable tasks. Ours usually run at least an order of magnitude more for a model in GLM 5.3's price class.

    We just published GLM 5.3 results on our multi-agent coding evaluations and it's definitely an impressive model, coming in around #6. For the price, it's actually not Pareto optimal, falling slightly behind Grok 4.6 (which is faster and the same price) and Sol 5.6 (which uses far fewer thinking tokens for comparable results). As with most Chinese models, it excels at iterating in a harness while its first answer/base fluid intelligence is below the American frontier.

    Data at https://gertlabs.com/rankings

  • Almost 10 models are passing the benchmark >95% - isn't that... substantially overly saturated?

    I'm also really skeptical of benchmarks that place any Haiku model very high. I've been thoroughly unimpressed with Haiku and my opinion has been that you basically shouldn't use it. Yet here, it ties Kimi K3. HMMMM.

    "Rubric Quality" seems a bit more realistic than "Pass Rate", but it is apparently judged by Fable 5...

    This entire write-up is also obviously clearly very heavily AI-assisted, which doesn't help matters any.

    by jchw
  • yeah that haiku really diminishes the claims behind the benchmarks. luna-max is significantly cheap and it is a strong performer but i dont see it on the benchmarks.

    I just get a feeling this site started with an intent to elevate Chinese models above the rest so wouldn't be surprised if the whole prompt sail was set to that tune

  • Interesting - when Kimi K2.6 came out I switched over from Anthropic models, with at that time comparable to better results for me. I was using Anthropic via API, heavier months were roughly $400 worth of Anthropic tokens - I can get the same thing done via a $100 ollama subscription.
  • > Almost 10 models are passing the benchmark >95% - isn't that... substantially overly saturated?

    The only true benchmark for any of these models that I've discovered isn't if they can pass precanned SWE tests, but rather can they create something novel? This isn't even too difficult to test, just give it a seemingly impossible task let it spin and see where it ends up.

  • Just read through some of the code "benchmarks", and I see why: https://reinvently.co.uk/tools/ed-o-meter/tests/

    Most are extremely trivial tasks. I would be surprised if a model from 2 years ago failed these...

  • This whole thing immediately reads as Claude generated, making it hard to take seriously.

    Why do these results contradict existing serious attempts at benchmarking LLMs? Namely: https://artificialanalysis.ai/ https://arena.ai/leaderboard/agent

  • Because it's notoriously hard to benchmark LLMs. Ultimately every benchmark is different and measures different things. This is why companies that make LLM models have private benchmarks - they find the areas where the model is weak and make that their goal.
  • Haha theres even an em-dash in the title