Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • I'd be curious to see the results, especially with some models having 1.5m and 2m context sizes, if the first 75% of the context was filled with unrelated info.
  • i used for several hours now and my verdict is that its no better or worse than sol

    its surprisingly bad at UI which is unexpected

    its also lacking in depth vs sol 5.6 which goes above and beyond (which in itself is also an issue at times)

  • Anthropic has imo underrated marketing and positioning skills, mythos/fable hype/fear being the most obvious indicator but even the way they almost haphazardly position their models with no intentional cohesion, people see model names and numbers, it's easy to think of them as more intentionally accurate like how cars make S models or AMG, but then the performance and surprises surpass the prior expectation that was set by previous models, rather than having it be more obvious, suddenly the Anthropic Camry will outperform their Corvette without any fanfare.
  • I don't have a horse in this race, but to me this makes GPT-5.6 Sol Max look better. It is about half the cost for nearly the exact same performance. It just goes to show how expensive Fable really is when Opus 5 is still this expensive relative to GPT 5.6.
  • Make sure you’re comparing opus high to sol max. That’s where the comparison makes sense
  • I like how Opus 5 doesn't re explain EVERYTHING to me like 4.8 did. GPT 5.6 SOL reasons WAY too hard over nothing, and Opus 5 is an amazing mode. Way to go anthropic
  • Yes, I didn't appreciate this because I was giving Sol a brief to implement and it was doing very well.

    So then I just told it to do its thing without a brief and it went for 2.5 hours and used 30% of my week. I tried the same task with the brief and Sol went for 30 minutes and used 2% of my week. Compared the two and the 30 minute brief-based Sol output was much better factored, shorter, validated better, scoped better, and of course cheaper.

    Left to its own devices, Sol goes out of control.

    Now I ask Sol to write the brief and Terra to implement it, works pretty well and overall usage is down.

  • The more interesting finding is that it's still the second most expensive model (after Fable 5) by a long shot.

    At least two models (GPT-5.6, Kimi K3) match its score (~1-2% diff) for half the cost.

  • The chart shows max effort, used mostly by price-insensitive enterprise users. At medium effort it drops to almost half K3’s cost, and is probably sufficient for 95% of coding tasks.
  • #1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id checking - which I have not experienced personally), it’s not worth whatever slight benchmaxxing they did for the latest release.
  • Maybe you just aren’t doing anything meaningful. Terrence Tao doesn’t whine about woke AI models.
  • The utter meme-think direction this company takes with regards to sycophancy of its models is disgusting.
  • I told Claude Opus 5.0 to use a global api key for a PFAAS to deploy some web applications in a test environment and import some data into them. It balked at using a global api key because the security issues surrounding the permissiveness run afoul of it's sensibilities.

    I have done this task with Opus 4.5, Opus 4.6, Opus 4.7, Opus 4.8 and Fable, without issues.

    I have done this task with Codex 5.4, Codex 5.5, and Sol 5.6 without issues.

    Opus 5 is too cautious to be productive for me. It needs more tuning.

  • Yep. I cancelled my Claude Max subscription 2 weeks ago after feeling like Anthropic was doing everything it could to fuck with my day to day. Their lead would have to become significant for me to ever go back.
  • Every time an online chatter (e.g. "limits are better", "model is better") makes me to reevaluate my principle of never paying Anthropic, I go to the model card, which strengthens my belief in the principle.

    Why is Anthropic is so hell-bent on this auto/silent downgrade? Do they have a single user who prefers an auto-lobotomization instead of a refusal? Have they learned nothing from the backlash the first time?

    by gck1
  • It's definitely not benchmaxxing from my experience with it. I have a test I use on all the models to create a game and Opus 5 feels like a generational leap compared to the rest. Benchmarks don't paint an accurate picture, you have to try them for yourself.
  • What are you asking that you’re so regularly running into censorship?
  • So if you're in Google leadership, you sleep in the office, right? Not merely because you have a ton of work but also because you're deeply ashamed to be seen in public.
    by blfr
  • If you're Demis, at least, you sleep fine because you were personally an early investor in Anthropic.
  • No. Demis is busy creating another documentary about how great of a human being he is . and giving interviews to fawning journalists projecting profundity over his every word.
  • Is Google trying to compete with OpenAI and Anthropic re: maximally intelligent models? Google seems to be the only one of the three that doesn't pray and self flagellate at the altar of AGI.
  • They are tied for first using the Google-proof question and answer benchmark: https://artificialanalysis.ai/evaluations/gpqa-diamond

    Maybe that's their only goal?

  • Google still has several enormous advantages here:

    1. Google Books, Youtube and the Google Search index all provide vast amounts of legally acquired training data.

    2. They can easy people into AI using the info box. I think this strategy is working even if it does cannibalize their main revenue source. Better than just withering and leaving all of the money to OpenAI/Anthropic. I would not be surprised if Google has significant layoffs due to reduced ad revenue at some point, but I think they'll still be on top.

    3. They already have their hooks into people's lives through Gmail, Google Calendar, Android, etc. The only other companies that come close are Apple (but for a much smaller number of people), and Microsoft (but only for business).

    The fact that Google's models might be 20% worse, or a few months behind Anthropic's is completely insignificant in comparison to those things.

  • Gemini models are at or near the top in several categories, though, so I'm not sure the takeaway is that they're shamefully far behind.
  • I don’t think Google cares about being the most intelligence AI as much as it cares about monetizing it with all its products, which requires speed.

    Google has long said that this is what it cares about most, the fastest at giving the correct answer to questions.

  • They can sleep just fine being the only player in town actually not loosing subsidized money.
  • What's interesting is this:

    The top AI models by Intelligence Index are: 1. Claude Opus 5 (Adaptive Reasoning, Max Effort) (61), 2. Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) (60), 3. Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) (60), 4. GPT-5.6 Sol (max) (59), and 5. Claude Opus 5 (Adaptive Reasoning, High Effort) (59).

    Which means Opus5 at Xhigh is still smarter than Sol at max, and Opus5 at High is equal to Sol at max.

    That would make Opus5 High same as Sol max, and now I wonder what the price and speed difference between those is?

  • Opus5 is simply not as inteligent as Sol max. To me it looks worse than Opus 4.8 on some tasks. When I say worse I mean mainly superficial. I basically have to teach him how the whole app/framework works before he just jumps doing stupid stuff(I.e adding features already supported but in a different form)
  • > That would make Opus5 High same as Sol max, and now I wonder what the price and speed difference between those is?

    According to AA's "intelligence vs cost per task" and "intelligence vs time per task" graphs, Opus 5 High and Sol Max are roughly evenly matched on cost and time.

    On DeepSwe, Opus 5 beats Fable but not Sol.

    On FrontierCode, it destroys everyone, unless you set it higher than Medium effort, and then it tanks, falling to Sonnet level?

  • Before getting too excited, take a look at the intelligence vs cost matrix: https://artificialanalysis.ai/models?intelligence-index-toke...