Comments
Hacker News
by entity002
At least two models (GPT-5.6, Kimi K3) match its score (~1-2% diff) for half the cost.
by chmod775
by andy99
by blfr
by aarondong
The top AI models by Intelligence Index are: 1. Claude Opus 5 (Adaptive Reasoning, Max Effort) (61), 2. Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) (60), 3. Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) (60), 4. GPT-5.6 Sol (max) (59), and 5. Claude Opus 5 (Adaptive Reasoning, High Effort) (59).
Which means Opus5 at Xhigh is still smarter than Sol at max, and Opus5 at High is equal to Sol at max.
That would make Opus5 High same as Sol max, and now I wonder what the price and speed difference between those is?
by didibus
AA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer.
This seems to be a good proxy for param size/density and the ranking breaks down as such: Claude Fable 5 (with fallback), Gemini 3.1 Pro Preview, Claude Opus 5 (Max), Grok 4.6 (high), Gemini 3.6 Flash, GPT 5.6 Sol (Max)
I've thought for a while that Gemini 3.x has 'big model smell'
by firasd
A single metric ranking is useless because each model has strengths and weaknesses for specific domains and tasks. There is no "one best model" anymore, and you might not even need the best model for the level of complexity for your task.
For e.g. you might use Fable for UI design, Sol for systems design backend work and Kimi K3 for exploit development.
The only purpose these metrics serve is bragging rights for the model companies.
Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- I like how Opus 5 doesn't re explain EVERYTHING to me like 4.8 did. GPT 5.6 SOL reasons WAY too hard over nothing, and Opus 5 is an amazing mode. Way to go anthropicby entity002
- The more interesting finding is that it's still the second most expensive model (after Fable 5) by a long shot.
At least two models (GPT-5.6, Kimi K3) match its score (~1-2% diff) for half the cost.
by chmod775 - #1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id checking - which I have not experienced personally), it’s not worth whatever slight benchmaxxing they did for the latest release.by andy99
- So if you're in Google leadership, you sleep in the office, right? Not merely because you have a ton of work but also because you're deeply ashamed to be seen in public.by blfr
- Before getting too excited, take a look at the intelligence vs cost matrix: https://artificialanalysis.ai/models?intelligence-index-toke...by aarondong
- What's interesting is this:
The top AI models by Intelligence Index are: 1. Claude Opus 5 (Adaptive Reasoning, Max Effort) (61), 2. Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) (60), 3. Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) (60), 4. GPT-5.6 Sol (max) (59), and 5. Claude Opus 5 (Adaptive Reasoning, High Effort) (59).
Which means Opus5 at Xhigh is still smarter than Sol at max, and Opus5 at High is equal to Sol at max.
That would make Opus5 High same as Sol max, and now I wonder what the price and speed difference between those is?
by didibus - Very interesting that one of the components is "AA-Omniscience Index"
AA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer.
This seems to be a good proxy for param size/density and the ranking breaks down as such: Claude Fable 5 (with fallback), Gemini 3.1 Pro Preview, Claude Opus 5 (Max), Grok 4.6 (high), Gemini 3.6 Flash, GPT 5.6 Sol (Max)
I've thought for a while that Gemini 3.x has 'big model smell'
by firasd - The funny thing is that these leaderboards have become completely meaningless for end-users to make decisions on when to use what model.
A single metric ranking is useless because each model has strengths and weaknesses for specific domains and tasks. There is no "one best model" anymore, and you might not even need the best model for the level of complexity for your task.
For e.g. you might use Fable for UI design, Sol for systems design backend work and Kimi K3 for exploit development.
The only purpose these metrics serve is bragging rights for the model companies.