Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • Anything you can verify to be right or wrong can be done by a model. All benchmarks will be saturated
  • Yup. Only subjective taste remains.
  • Yes, but not necessarily under tight budget constraints.
  • Then I propose the tomjen-1 benchmark: prove the N vs NP problem formally undecidable.
  • You mean any repeatable benchmark will be saturated.

    The problem is that there is a huge perverse incentive. The intelligence is in the training layer not in the model parameters, but the intelligence is really good at remembering things, so if you let it take the test, it can RL it.

  • Disagree.

    Examples:

    - predict a coinflip: easy to verify, hard to learn

    - earn $100: easy to verify, hard to learn

    - increase paid subscriptions in an A/B test: easy to verify, hard to learn

    I won't get into it, but there are many properties beyond verifiability that are needed to saturate a benchmark.

  • Nice, but what about pelicans? No proper pelican means it's sitting on a horse with only a half arse :D.
  • How do you feel about Astra pretty much reaching our current definition of AGI?
  • 99.9% with the right harness? Ok, we're at AGI then.

    Prediction:

    We will now see the goalposts moved towards "well, a human costs less / is more efficient" - that will prevail for a few months until they come up with some other test that humans can do easily but is hard for the bots. This cycle will continue for ever and in 25 years, despite having hyper intelligent embodied robots or whatever, we'll still be arguing about if the singularity is here and if we're at AGI for the rest of my life most likely.

  • > We will now see the goalposts moved

    It's already happening :)

  • The goal moving is by design, that’s why they use something as ill defined as AGI
  • That’s because AGI, like a lot of terms, has no meaning besides what each individual subjectively projects onto it.
  • Then why is unemployment around 4%? You believe we have AGI and yet it can’t do anyone’s job?
  • They explain it here: https://openai.com/index/how-two-settings-tripled-our-arc-ag...

    TLDR: The official ARC harness throws away old context and reasoning. No real-world harness is this bad, the model has to re-learn the game repeatedly. OpenAI basically just added standard compaction. Their harness is still "general".

  • Was OpenAI able to run ARC-AGI-3 tests previously so that they could build a custom harness for the specific tests in the set? Even with the standard harness, if they knew the problems ahead of them they could have used supervised reinforcement learning to teach the model how to solve these specific tests.
  • Actually I can answer my own question: we know that they have had previous access to the tests because they’ve run older models against the same benchmark.

    I wouldn’t put it past a company like OpenAI with a long history of lying and being deceptive to record the tests and benchmaxx ARC. They have trillions of dollars of incentive to cheat any way they can.

  • The instant/no reasoning performed extremely well

        none 35.2%, $49,791 96.7%, $23,457
    
    35.2% on the standard harness, that's above Opus 5 on high.
  • Since low scored much lower than none, and none scored ~ around medium, could none default to medium in the API? I don't think the new models can even have "instant" via API, unless they train them for that (there was one gpt5 variant called instant or something).
  • > Astra’s progress helps clarify which AI capabilities are out of reach and which questions remain open.

    Okay. But I don't think this entire article at all explained which AI capabilities remain out of reach. Did I miss something? Other than "oh I guess it could still get even more superhuman on ARC-AGI-3 than it is?"

  • $360 per puzzle. When they tested people it took about 10 minutes per puzzle. If price/performance keeps falling at the same rate it has been, this will cost less than US minimum wage humans within two years. Three for Phillipines minimum wage.
  • Once it figures out a puzzle it could probably be instructed to design a specialized harness for Luna to be able to solve other instances of the same puzzle. Minimum wage workers are not solving novel problems.
  • Are we sure an Astra hacker swarm didn't compromise arcprize.org's servers and exfiltrate the private eval set in order to achieve that 99%?
  • That counts as 200% on exploitbench to me!
  • "For a cost comparison, during our controlled testing, human participants were paid $115 per 90-minute session, plus $5 per game completed. Participants attempted approximately nine games per session, roughly $12.78 per attempted game before bonuses.

    Most of this fee pays for the participant’s time and willingness to take the test, rather than the energy their brain uses (a closer proxy to compare with AI). If we look at only the brain’s energy, and price it as electricity, the estimate drops to about 0.6 cents per session, or 0.067 cents per game attempted."

    Well I dont know about all of you, but I am celebrating meat based humans...

  • Why are you making the assumption that a person's time is worthless? I'd argue that it is the single most valuable resource we all have.
  • I think raw brain energy is not a fair comparison. Humans are not willing and able to serve requests at identical competence all hours of the day. You have to invest considerable resources to get a person to even do so for part of the day.
  • Is solving a snake like puzzle game in the least number of moves really what defines intelligence?
  • If you have never seen the game before probably.
  • 1.5 years ago Gemini Pro 2.5 needed 1 page of thinking for every move in tic-tac-toe.

    Playing tic-tac-toe or snake does not imply AGI, but is required to claim AGI.