Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • ARC-AGI3 doesn't seem like a great benchmark to me in the first place. It assumes a lot of human like tendencies which an AI either shouldn't or wouldn't have. Particularly in the genre of "gameplay" where unspoken assumptions from prior games inform our understanding of rules.
  • It's actually crazy to see the difference between opus 5 and the next best model on ARC AGI 3 when you actually look at the ARC AGI problems
  • Why? 30% is passing the first two problems only, which are really very simple.
  • How believable is this benchmark? EG maybe opus was training on this? (You can try to identify the IP of wherever previous ARC questions came from)
  • Why is Fable not on here? I wish Fable hadn’t come out because it’s taking the wind out of every release because that feels like the cap above which the US government will not let LLMs improve anymore and everything they’re releasing from this point has to be worse than that.
  • I don't know why exactly, but Fable has felt the most human LLM to arrive.
  • > Why is Fable not on here?

    Because the data retention policies didn't guarantee that the ARC team could run the semi-private set of problems without fear of them being trained on later on. They only run the semi-private set when they get assurances like ZDR.

  • I have a suspicion that they are just trained on puzzles by now
  • It is RLVR, Not a puzzle, Not leetcode
  • It's like we've come full circle:

    First people practiced L33t3cod3 problems for interviews

    Then people built AIs to build software

    And now the AIs are studying L33t3cod3 problems

  • There are private datasets, and 3rd party providers of these models. Fable doesn’t have a datapoint here because of its particular data retention policy. Even if you don’t trust AWS, do you think Opus on AWS is also sending the data to Anthropic? Do you have any evidence?
  • Also top on the freshly released Frontier-Bench, by a large margin: https://www.frontierbench.ai/
  • ARC-AGI is a beauty contest for pigs where the pig's owners compete to see who can apply the lipstick most convincingly.
  • Great comparison! We only have to take into account that applying lipstick well bears no consequences, but applying it poorly (i. e. new model tanking the benchmark) could amount to potentially losses of billions of dollars for the pig-breeders (AI labs).
  • Why is there no Kimi 3, and GLM5.2 didn't run the third benchmark? I am more interested in knowing the abilities of open weight models.
  • > Only systems which required less than $10,000 to run are shown. (Notes[1])

    Am I lost or are their many models on this ranking (Opus 5 included) that clear this?

  • Many models are much cheaper through their subscriptions' included usage. That could be what's happening here.

    Claude gives you something like $5000 of tokens on a $200 plan.

  • ARC-AGI-3 launched a few months ago which would suggest that prior models likely had no knowledge of ARC-AGI-3 or training on similar challenges.

    I could be wrong, but given the large outsized jump solely in the ARC-AGI-3 score, it would suggest that the model didn't become significantly more intelligent overall, but significantly better at solving those specific problems.

    This could mean one of two things (I think):

    - Opus 5 was not benchmaxxed on ARC-AGI-3, but has benefited significantly from discussions about the various challenges and mechanisms deployed in ARC-AGI-3 such that it has far better heuristics to solve its challenges.

    - Anthropic looking for buzz around their latest model picked a well regarded benchmark with significant room for improvement and focused some of Opus 5's training compute on ARC-AGI-3-style problems.

    Or it could be some combination of both. Personally, given how much of an outlier the ARC-AGI-3 jump is I struggle to see it being the product of a significant improvement in general intelligence.

  • I also noticed that Opus 5 doesn't show a corresponding gap on the ARC-AGI-2 leaderboard. There is a significant increase in performance between Opus 5 and 4.8 on ARC-AGI-2 though.
  • > it would suggest that the model didn't become significantly more intelligent overall, but significantly better at solving those specific problems

    We would actually need a test that shows the ability of a model to export its skills to more problems ("interdisciplinarity" etc.).

  • Why Anthropic models are always leapfrogging these benchmarks, but in real life work I do feel like after 3 weeks I am back to Claude Opus 4.5? (regardless of the model I use, Fable was exception for 1 day when it was released)
  • I honestly just use GPT models nowadays, Claude models are too restrictive and more of a quitter and fable/whatever is just too expensive to be worth it.
  • Same, like I prefer 5.3 codex over the “stronger” models.
  • Going to call it user error if you find Opus 4.5 better than 5, sorry.
  • Well, what kinds of things do you see Opus 4.5 completely fail at? Maybe those are not the ones that newer models have improved on.
  • 5 seems incredibly smart to me in my conversations today about some pretty niche ideas in.longitudinal modeling. It.felt.like a big step up.from 4.8, to me
  • It's called frog boiling.

    We get used to the new level of intelligence so fast, any deviation feels like going back to the stone age.

    If you don't believe me, create something complex with Opus 5 and then with Opus 4.5, and notice the difference.

  • Some 20 years ago, the telecommunications sector in Germany was liberalized. Many telephone card providers entered what had previously been a barely competitive market. They advertised their products with aggressive claims like: “Buy our €10 top-up card and get 660 minutes to destination X.”

    For the first few weeks, they would actually provide those 660 minutes to establish trust in their cards. But after a while, they would quietly start reducing the number of minutes on subsequent top-ups—say, from 660 minutes down to only 300. They wouldn’t do this for every card, so it was difficult to prove. Instead, they relied on averages across their customer base to make the economics work.

    Lately, I’ve found myself wondering whether something similar may be happening with frontier AI models. Companies launch with an exceptionally strong model and generous compute limits to build adoption. Once the model is established as a market leader, the incentives change, and users may start perceiving the service as becoming more constrained or less capable over time.

    I don’t have evidence that this is what’s happening with Anthropic—or with any other AI company. It’s simply a pattern that the current situation reminds me of.

  • It could be that the set of your day-to-day workload which could feasibly be accelerated by AI just happens to be saturated around Opus4.5, but you can still see lots of “reasoning” which makes you think the model is more performant in the first days of use. That’d mean you couldn’t perceive any meaningful difference in more powerful models’ results, even though you can see a difference in the raw output due to the length of reasoning traces leading up to the result.

    So for example, if your workload was literally just addition of sets of numbers, you’d never have noticed progress in the result beyond GPT3.x level models. But you would perceive a difference in the now-Tolstoyan length reasoning text accompanying the result.

  • The last time I checked, for the arc agi 3 leaderboard, the models are given a simple prompt and the game input and asked to play the game, no harness/tools. If harnesses were allowed, I would expect the benchmark to be saturated. There were a few harness attempts, but they could only be evaluated on the public set, so it's not an apples to apples comparison.

    My guess is, the large score jump for Opus 5 is mainly because of getting the right RL envs for training.

    It's becoming harder and more expensive to build and run meaningful benchmarks, it would be interesting to see what they do with arc agi 4, maybe just give it gameboy/steam games and see how they compare vs a human baseline? The latency requirements and very long horizons in games could be an interesting challenge for llms.

    by dinp