

Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- I might’ve missed it, but why was Fable 5 tested on high while Opus 5 was tested on max? Seems like quite a few of them aren’t on the same effort setting as well. Although effort doesn’t really matter anymore since they can change it dynamically, seems like that might be viewed as an experimental error to some.by ninjahawk1
- This is a pretty embarrassing showing for Grok. I wouldn't trust xAI models as far as I can throw them, but I am interested in how much of this is deficiencies of the model and how much is their harness just terrible. Not that it makes it better, a good harness is far easier and less expensive to make than a model.by bastawhiz
- auto research the new cool kid on the block - look at https://mlx.fastby c0rruptbytes
- “We ran 153 autonomous runs across 18 frontier models on the nanoGPT optimizer speedrun.”
Uh.. okay.. but whats a run… read blog
“We want to measure how well frontier models can conduct research….””we ran 153 autonomous runs on the nanoGPT optimizer speedrun across”
Okay but what is a optimiser run and what connection does it have to being good at research?
“For comparison, Anthropic's internal automated AI R&D evaluation optimizes a model on a CPU node,”
So I should go look what Anthropic was doing to understand?
Why not just explain what it means in their blog..
by totetsu - Pretty interesting to see on the training front. I've used most of these models to grind semi-autonomously (days at a time) on kernel optimizations (except for Fable - it kept triggering guardrails almost immediately and bouncing me down to Opus 4.8 at the time). I think for a lot of people that might be the biggest problem, although it looks like Opus 5 still does well.
I found that if you leave them alone undirected, the models (especially GPT models) will rathole, but with the right scaffolding it seems to work pretty well. My general loop is to start with ideation and profiling phase, limit # of runs before forcing moving on to the next item down the list, and then iterating, potentially mixing models with "fresh eyes". This is probably something that could be fully automated, but I like checking in once a day or so and seeing what's happening and redirecting.
by lhl - What's going on with sol here? The note says it spends a lot of time waiting, did it just not effectively use time (i.e. something like parallel runs) so it's graph ends up being stretched in the time axis?
I'm also seeing notes like on Opus 5 saying it was a run with a older serial version of program.md, so the graphs aren't complete apples-to-apples comparisons?
Edit: the blog seems to address these https://www.primeintellect.ai/blog/measuring-autonomous-rese...
by nsingh2 - I did something similar in March using Opus 4.6 (iirc) on google's "Parameter Golf" challenge, "a challenge to train the best language model that fits in a 16MB artifact and trains in under 10 minutes on 8xH100s, evaluated by compression on the FineWeb validation set (tokenizer-agnostic, bits per byte)."
I never ended up writing it up, but you can watch me spend $600 as it explored different experiments: https://github.com/rbitr/parameter-golf/blob/main/BUDGET.md
I found, similar to another comment, that it got in local minima very easily and continued to pursue loosing ideas instead of exploring (despite being prompted to do so and being aware of how much budget it had left). I also found it tended to ignore instructions. And one example, when I fed it a better solution that had come along from the public leaderboard, it ignored everything it had done and started exploring locally around that new solution, which wasn't very interesting or productive.
Would be interesting to re-run with a newer model but it's hard for me to justify the money again.
by amarble - "Almost every model finds the same winning ideas. What separates the best traces is what an experiment leaves behind. They preserve weak signals long enough to validate them, but they also have a better understanding of the results."
Curious if a harness that helped preserve signals in some history log would change the outcome.
Also curious if different goal prompts would have changed the outcome. Not a bunch of prompt engineering; small diffs like "consider novel solutions, keep track of weak signals".
IMO they allocated quite a bit of GPU time to the same goal prompt.
by vibe42