

Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- While I like this index, calling in "Intelligence" might be confusing - it is a mix of coding and knowledge.
Compare and contrast with ARC-AGI, BabaIsBench (https://quesma.com/benchmarks/babaisbench/), or MazeBench (https://mazebench.com/blog?post=introducing-mazebench).
In particular, in one Baba Is Bench post (https://quesma.com/blog/baba-is-aug-2026/), while quoting a Pareto frontier chart from AA, I noted:
> Intelligence Index vs. Cost per Intelligence Index Task from Artificial Analysis. Note that it is based on score of benchmarks like Terminal-Bench v2.1, SciCode, Humanity’s Last Exam, GPQA Diamond - not necessarily fluid intelligence like in abstract puzzle games of ARC-AGI-3 or Baba is You.
by stared - This update really gives OpenAI a boost. Not saying there's anything inaccurate or untoward about that, but the timing is unfortunate. It would have looked better had it been done prior to the Fable 5.1 and GPT 6 releases. I guess AA would say that there's no perfect time to do these updates, given the rapid fire pace of releases!by AnodicElegy
- I wonder if Artificial Analysis could be influenced by certain model companies. Looking at the changelog [0], they updated a few times after new models appeared, so US models progression was much higher than that of other vendors (like when they updated the algorithm after Kimi K3 versus Opus 4.7, so Kimi dropped in the rankings). Or maybe thats just coincidence.by __natty__
- This is really a great achievement: "Astra dominates the output token frontier"
Many labs used increased thinking to boost benchmark scores and performance. Most of the Chinese models were doing that for a while. Google and Anthropic as well.
Not OpenAI. 5.6 already was much more token efficient than other models and Astra beats Sol in token efficiency by a wide margin.
Edit: Just to make the point: Astra (max) has the 2nd highest score and the third lowest output tokens (among the models shown by AA).
by __jl__ - IMO one of the biggest losses of the OpenAI/Cursor breakup will be the loss of OAI models on CursorBench [1]. Their bench has always been one that most-fit my mental model of how good each of these models are. I find AA’s Intelligence index to often be out of alignment with my own subjective evals.by jjcm
- I have no idea how artificial analysis got to be something anyone took seriously. This is their new benchmark set?
A glance at their new index shows that whatever they're measuring, it isn't useful.
Spend an hour with gemini 3.8 and tell me that model belongs in 2026. It feels like the model has Alzheimer's. It gets confused about whether what it reads is what it did. Just crazy bad.
I haven't tried muse spark 1.3. But it must have been a miracle since 1.2 to hit that rank.
Video game journalism vibes all over this.
- They realized Astra having the same score as Sol was silly so they rushed to update the index so it fits what people expect.
The old index was clearly bad (Astra is way better than Sol) but it's also unscientific to tweak it like this.
by redox99 - Imo the omniscience index they have has the highest correlation to actual usefulness of the models.
https://artificialanalysis.ai/evaluations/omniscience
> measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer.
This is so useful because it makes you actually trust a models output. A high score on benchmarks is not as useful because a model overtrained to always answer will give confidently wrong responses. But this index measures how often it is correct while penalizing wrong responses so that a high score means you can trust this model more and when it doesn't know it is more likely to tell you that it really doesn't know rather than making shit up.
Fable also performs a lot better than opus 5 here which correlates very strongly with perceived strength despite the models performing similarly on e.g. DeepSWE
Astra is a big jump from sol and performs the same or slightly better than fable here.
by jascha_eng