Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- Of course I echo the universal: Opus 5 sucks. But what also sucks is still using 4.8. Are you all seeing this? It's like the older models got dumber just before Fable and Opus 5 were coming out? I've heard the theory that it's because 4.8 is getting put on older hardware? And I imagine that just ratchets down reasoning time then possibly?
Sometimes I'm just trying to sus out if I'm truly seeing things these days or going a little nuts :)
by nate - I'm not an expert by any means whatsoever but deploying models is not a straightforward task. There's a lot of levers to pull and I bet when models get "downgraded" to older hardware they do so WITHOUT the same stringent quality control of the output as they do when they release it.
I don't think it's something deliberately malicious like planned obsolescence but it's more like startup culture of "just make it fit in this sprint".
by gandreani - For me, Opus 5 mostly sucked because of its incomprehensible writing style. Having the system prompt focus on writing in terms that are easier to understand helped a bitby make3
- RIGHT!? The data suggests this isn't true but every fibre of my being is convinced that Opus 5 High was EXCELLENT at launch and has been lobotomised since then. My benchmark is Sol High. I've been using both consistently and either Sol High suddenly became MUCH more capable - and the data does not support that - or Opus 5 became much dumber. It's so bad I can't even use it anymore.by Gareth321
- I definitely agree. I honestly cannot do tasks that require even minor complexity. Opus 5 keeps forgetting things in context as well and coding conventions. Really cannot build with CC without Fable.by samrj12
- This is yet another reason why I think local models will win in the future. They're almost certainly A/B testing all sorts of opaque stuff that people have no clue about, hence the various 'How's Claude doing this session?' popups.
So what you are paying for may vary on a day by day basis, which is quite undesirable, even if their main goal is simply to make a better model. When it comes to a tool, I'd rather have consistent mediocrity than instability.
- I have noticed very serious degradation in performance from Anthropic. I've switched away from Opus 5, but 4.8 is still much worse than how it was before Fable came out. It's constantly making what seems like obvious mistakes... I point them out, it's constantly apologizing.
I don't know why...
Is it A/B testing?
Is it load shedding?
Is it because I'm in Canada?
Is it because I'm not on the Claude Max plan?
Is it because I'm not paying via API?
Is it because I'm not paying via Bedrock?
Is it because the U.S. is worried people are distilling?
Is it because the U.S. wants to keep the top capability to themselves?
I think open models are the future. Anthropic is killing their reputation so fast. If they don't come clean I think they're cooked.
by purpleidea - Anthropic killed it for me with all these silly security guardrails. As a programmer I use the models to check for vulnerabilities in my work. I also use the models for fun reverse engineering projects on the side. Neither of these activities are illegal. And I don’t appreciate being treated as a criminal. I asked Claude to do something the other day and it refused. I copied and pasted the exact same prompt into codex and it happily churned away at it for hours. What Anthropic has done to kill their own company is a bit sad.
- They're expecting to IPO bigger than SpaceX later this year.
HN's take that they "kill their own company" seems wildly out of touch.
Truth is, Sam and Dario and Elon are all terrible and so are their orgs. Leaving any one of those companies for any of the others is wild.
- Transformer models are quickly becoming a commodity, and I suspect in time we'll all be running them locally. Even now, you can run something pretty useful on a 16 GB graphics card, and I suspect a decade into the future, entry level hardware will be running better models than high-end graphics cards can run now, as entry-level hardware gets better and models get more efficient.
It doesn't mean hosted frontier models wont exist, they'll just be rare. It's no different than any other commodity market, for example most cars are cheap commodity models, with rare individuals buying expensive luxury cars and businesses buying expensive trucks and specialized equipment.
by dlcarrier - This won't happen so long as the bandwidth bottleneck is a thing. Local models are just several orders less efficient at scale, and this limitation is inherent to the architecture and won't go away unless both hardware and model structure change enormously. There are really only two usecases I can see for local models going forward:
- 10-1000 person groups such as corporations where you can amortize serious hardware with parallel use
- Porn.
- >and I suspect a decade into the future, entry level hardware will be running better models than high-end graphics cards can run now, as entry-level hardware gets better and models get more efficient.
I doubt it. The play seems to be: lock what was once commodity compute up into datacenters depriving us regular folk of it, then sell it back to us on subscription. Even if my #NeverSubscribe movement succeeds, all that misdirected hardware [into datacenters] won't likely be practical for home use.
by rustcleaner - Yep local models will be good enough for most things you need to do, in the same way as most people need a laptop not a supercomputer.by chr15m
- For many coders including myself, LLM based coding agents work well enough to be useful, and in some cases worth paying for.
What I don't see is vast areas of industry finding $10s to $100s of billions of value in LLMs. There's no lint or compiler that can check for correctly constructed contracts. So LLMs, which should be useful to law firms, incur a lot more manual checking of their work than coding agents.
Less formal document production in other industries is likely to have less structure. That might not matter in some settings but I'm having trouble thinking of an example off the top of my head.
by Zigurd - > What I don't see is vast areas of industry finding $10s to $100s of billions of value in LLMs.
Translation between languages.
That value dwarfs all programming value that can be had. Economically, culturally, scientifically, spiritually.
by carlosjobim - > There's no lint or compiler that can check for correctly constructed contracts.
There are definitely linters and this exists https://catala-lang.org/
by _joel - I have a family member that is an attorney in housing law. She claimed that LLMs are not particularly useful for her work. If she asked a simple question like, "Find all the <insert specific housing laws> for all 50 states," then she still has to go and check every single one of the laws. Since the legislature is modified so often, she cannot look at, say, Maryland's law and know if it the LLM output was the 1990, 2014, 2018, or 2026 version of the codified law. In order to fact check the law, she has to look it up, and by that point in time, she has the answer she did the work of the LLM.by hirvi74
- While there's no linter for writing contracts, my experience (as a commercial lawyer) is that frontier LLMs are far better and error checking and far quicker at writing than the average senior lawyer. The main thing holding back further deployment (in my jurisdiction) are concerns around data residency, privilege and how fundamentally it will break an industry that is so heavily reliant on time based billing.by ray_kay777
- My company still hasn’t been able to deploy wide access to Fable because it’s not available on a ZDR basis. This wasn’t mentioned in the article but I imagine this factor is not irrelevant.by semiquaver
- ZDR for fable is coming, very soon.by jimmydoe
- It was in one paragraph mentioned, without much detail. But yes, I agree it is one of the largest barriers.by usaar333
- That's the reason at my shop. Real shame, as there's really no limit on token spendby topbanana
- This is the biggest factor. FT (or the analysts it cites) fumbled the ball in this article.by kwisatzh
- The same for my organization. It's explicitly forbidden in my organization because it's not available with ZDR.by x3n0ph3n3
- Same for me, they’re not offering it with the same data residency features of Opus, many enterprise companies can’t accept compromises.by Lucasoato
- >ZDR
Zero Data Retention, for the uninitiated readers in this thread.
The no-ZDR is clearly to permit surveillance. I would be shocked if NSA wasn't all up in these SOTA model providers' systems.
Never subscribe!
by rustcleaner - Wow the sentiment here is so negative.
I'm on the $200 plan (work pays) and I also have the $20 OpenAI plan (I pay) and keep a balance on OpenRouter.
There is nothing as good as Fable, not even close.
I recently had it run a 18 hour autonomous rebuild of a project (moving from Spark to Pandas for performance/data size trade off issues).
It orchestrated Opus sub-agents flawlessly for 18 hours. It even did a great job of managing the number of agents to keep them within the 5 hour budgets (I think I had to restart it twice).
After 18 hours I ran a /simplify, /code-review, /simplify cycle which went for another 6 hours.
2 billion tokens (mix of Opus and Fable), 24 hours of continuous coding and a bug free outcome. It would have cost $2000 at API prices and worth every cent.
Fable's ability to keep other models on track while working on these long horizon goals is so much better than anything else.
Far from neutered, I've never had a cyber refusal, and Fable's English is actually readable (unlike Opus 5).
As an aside: while I hate reading Opus 5 English it still is a noticeably better model than Sol in my experience.
But I could handle losing Opus5 is I got Sol instead. But there is nothing even close to Fable.
by nl - the sentiment is negative and justified. anthropic nanny states what you can do.
in your instance, anthropic may decide, arbitrarily, to stop 'autonomous rebuilds / refactors and ports' because they could pose some alignment/rights/etc risk to whatever slop their philosophers dream up while they're out eating $200 avocado toasts. then you can't do the thing anymore.
fable is good, absolutely. agree it roasts Sol which is, comparatively, a little receipt-hunting jack**
but now imagine being an enterprise, and having another organization not only taking your workflows and baking it into your models, but then deciding they can arbitrarily cut you off.
when you can instead own your data, use an agnostic provider, and get better results (through model combinations), it will take 1-2 quarters to figure it out.
the main reason anthropic is killing it is because they really do understand the enterprise development experience and lifecycle and have built products and have a sales-team that can deliver.
business-model and vibes-wise they have lost all goodwill in the past 6 months, and that momentum will be quite hard to regain.
by epsteingpt - Like others have suggested, you should give more time to GPT models. I sometimes launch Fable with elaborate review personas, it might take 30 minutes or an hour, exceed limits, to produce a review of a PR. Then I ask the same thing GPT without any elaborate 'come up with personas, review the reviews, do rebuttals, etc', and it can find problems that hours of Fable couldn't.by elAhmo
- > 2 billion tokens (mix of Opus and Fable), 24 hours
That's 1.3M tokens per minute. I suppose you mean input tokens? This would not have cost $2000 at API prices because you would have cache hits.
by acchow