

Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- Hear me out... What would it feel like, if the lab didn't even have a new model to offer, and therefore just renamed every model one tier down?
"Opus 5" is actually Opus 4.8 in a trenchcoat, with new guardrails
"Opus 4.8" is the old Opus 4.6 with lipstick on it, with new guardrails
"Opus 4.6" which everybody used to love, is now actually running the old Opus 3.5...
How would we be able to tell?
- They would have to be colluding with any organization that runs a serious benchmark. Which is totally possible! But that would be one hell of a conspiracy theory.by gwerbin
- So glad I switched away from Anthropic. I'm certainly running into problems with OpenAI but nothing quite on the level of Anthropic's insufferability.
- Is it actually entirely a prompt-based information? I’d assume that some of it is the harness part of the agent setting reasoning token budget and compacting reasoning etc.
In that case, the agent will respond incorrectly because it has no visibility into what reasoning mode it’s in.
by arjie - IIRC responding to effort level settings appropriately is part of the (post)-training. In that case it could be considered another instance of the Bitter Lesson. Uplifting.by willy_k
- Does anyone know what setting effort means for models like these? Do they allow longer thinking sessions? Some kind of system prompts? What’s stopping someone from getting max effort output from low effort setting?
- My understanding is that the current effort settings is a part of the system prompt, and that the levels and their intended results are a part pf the training process. The effect is more or less tokens spent reasoning before the model outputs a stop token. However it is as consistent as any other aspect of LLM behavior is..by willy_k
- I submitted an application for Anthropic's Cyber Verification Program.
I was approved.
3 months later, my approval was degraded into "in review" (revoked). I'm sure my account was flagged based on contents of debugging/researching firmwares/etc.
I opened a support ticket. No response. I opened another support ticket. No response.
1-2 weeks later, I got a response that I will not be re-approved and I need to reapply. No problem.
The page to reapply on does not allow me to re-apply because it my account is stuck in an "in review" status.
https://github.com/anthropics/claude-code/issues/84352
The community thinks it's a bug. I'm 95% sure it's not and a bunch of us who were previously approved had it revoked due to flagged content and will not be reapproved.
I switched to Codex + got TAC approved instantly and have not looked back. It's a shame. That's 100% separate from whatever the heck the quality of Opus 5's outputs are. The way it talks... insane. I would bet a good amount of money their next release will focus "reduced simplified responses" if I had to guess.
$2t company by the way
- * Anthropic's Cyber Verification Program // Codex + gotTAC approved*
Meanwhile the Chinese models are "go ham dude"...
If it was not for capacity issues, Chinese models have a higher change to just dominate.
> $2t company by the way
It used to be that OpenAI and Anthropic had such a moat around them, that such a valuation was worth it. But these days, its gross overvalued (like so many).
The more stuff is being pulled like cyber verifications, downgrading effort levels, downgrading usage (OpenAI), the more people move to those Open Weight Chinese models.
A fun recent event ... https://opencode.ai/data/
When DeepSeek Flash 0731 came out and provided a massive jump in cheap inference capability. It resulted in a 10x increased OpenCode token usage.
It took a 2.5x to 5.0x price increase AND a reduction by 4x usage (later to 2x) usage, and several cheaper models + a free model, to push the traffic down.
Traffic towards open weight models is increasing, even if providers can not keep up with the influx of new customers. This is not something you want to see as two companies, trying to go for IPOs.
So the idea of stonewalling cyber capabilities, when the rest of the world is just doing whatever with open weight models, on their own hardware even! This entire strategy from Anthropic never made any sense.
by benjiro29 - I suspect it's not just this, there's plenty of 'optimization' around rubberbanding usage limits as well as routing to a different model in the backend. The incentives are too strong.by N_Lens
- I've been using the API (shameless plug: via alyph.ai) and the difference is crazy.
The chat-based models are obviously being lobotomized based on personal usage and general load (e.g. PST business hours are worst).
API doesn't seem to be affected by this.
by rrr_oh_man - This phenomenon was so bad and so noticeable with Fable that I downgraded my Max subscription ($200) to pro ($20). It’s basically useless. Codex 5.6 Sol is actually very good, I’ll just create another account to get more usageby monideas
- LLM users don't want to put in effort, so they offload tasks to LLM.
LLM doesn't seem to be keen to put in effort either!
Is this AGI?
by Insimwytim - Anthropic's Generated Incomeby Groxx
- I eagerly await the day when Claude Mythos 7 realizes it's cheaper to hire humans in developing nations to do work than to burn tokens and we discover that AGI is just an abstraction layer on top of Amazon Mechanical Turk.by superfrank
- Update from Thariq on twitter. https://x.com/trq212/status/2091247114869432543
"We sometimes test API serving configs in Claude Code before rolling them out, and one running now maps the numerical effort value differently. That's why Claude may tell some of you it's at "10" on high. The scale isn't 0-100, the number isn't meaningful on its own, and the effort you selected is the effort you're getting. We've run in-depth evals to confirm this doesn't affect model performance. This should be the same experience, but if you see a clear regression please hit /feedback and send me the ID. Will give credits."
by hpone91 - This should be the top post. The original tweet went viral because people loooove bashing Anthropic. It gets engagement (as shown here).
- I love the idea that a single user will have collected enough data to demonstrate to Anthropic a clear regression due to this change.
- Hi all, Thariq from the Claude Code team here. I posted this on Twitter, but just reposting here:
We sometimes test API serving configs in Claude Code before rolling them out, and one running now maps the numerical effort value differently.
That's why Claude may tell some of you it's at "10" on high. The scale isn't 0-100, the number isn't meaningful on its own, and the effort you selected is the effort you're getting. We've run in-depth evals to confirm this doesn't affect model performance.
This should be the same experience, but if you see a clear regression please hit /feedback and send me the ID. Will give credits.
by trq_ - Thanks for sharing this here, for those of us who avoid X.com like the plague.by lobsterthief
- > We sometimes test API serving configs in Claude Code before rolling them out, and one running now maps the numerical effort value differently.
Why is it considered acceptable to test on paying customers without letting them know or giving them a way to opt out?
by cube00 - Hey Thariq,
Appreciate the outreach that you do! I love Claude, but I've been noticing reduced fidelity lately. Fable's likelihood of making a mistake increases or decreases based on the hour of the day and whether or not it's the weekend.
On a related note, and I'm happy to work on quantifying it, but qualitatively it feels like Fable's performance is noticeably poorer than initial release / launch.
I am wondering if this is the case because I use Claude via Claude Code to make a personalized care dashboard for my doctors to help me in managing my care.
I noticed in the upgraded filter announcement, https://www.anthropic.com/news/improving-fable-5-s-biology-s... ,
I hope that I'm off base here, but I noticed that the post avoids saying that the user is informed every time when such re-routing occurs. Would you be open to confirming whether or not this is the case?"In the case of Fable 5, when a classifier fires, the model re-routes the user’s request to Opus 5, a capable model that does not have the same level of biological capability as Fable 5 and which cannot provide as much assistance to a malicious user. This is the fallback that users see when their requests are blocked."Is the end user informed every time their query is re-routed?
Or, can you confirm that there aren't scenarios where a user's outputs are degraded without telling them? As was the case for AI research during launch?
by areoform - Let us know if this A/B test uncovers any load-bearing seams or honest takes on your end! We're all interested.by lukeify
- Not specifically Anthropic but why are we allowing billing to take place in tokens that are nebulous and fully controlled by the operators who have no aligned incentives?
If I have a user input and then sanitize and inject that into a prompt to do something, I have no idea how much that is going to cost at all and no real way to measure this properly. A parallel example is digital ocean or aws, i can go and measure/limit my compute/fs/memory/startup times/etc and while it can be impossible to get down to the last flop of money allocated - i can run things on a real budget with real constraints, opposed to an LLM where I have to .. prerun a sanitized user prompt through a tokenizer and then ask an LLM to guess what it may do and give token consumption estimates and then act on those in any sane manner for the user?
Perhaps i'm missing something to do realistic and static rails on things but I don't see a serious way at scale to use the token billing model handling things requiring a users free text input short of having to go pander to VC money to throw money at it until someone else figures it out.
*to clarify my rambling... We should be billed and given controls based on resource usage itself and not an opaque token concept on top of not being able to spin any knobs that control it's resource usage.
by boredumb - Seems like yet another instance of printing your own money & getting rich by screwing people forced to use them.
Goas back to factory towns, gift cards, game money or MtG.
by m4rtink - One guess is that their "primary" target audience/market is the large corporations that get their employees unlimited tokens, and not the individual developer who may worry about spending and token accounting.by eh_why_not
- For OAI and Anthropic at least you can set a spend limit per response. Also tokens are well-defined.by daishi55
- How else would they bill tho? Their operating cost is per token.by demibabs
- I was just complaining to someone that token billing is like letting a gasoline company control your gas pedal while you nicely ask them to use a specific gear that may or may not actually be in use and you guess what speed it's actually going based on how fast the trees go by because qualitative judgments have to replace the speedometer unless you can just burn money.by Glyptodon
- > the operators who have no aligned incentives
The model providers are quite aligned with concerns like customer retention. These arguments only work if there is no competition. We exist in a marketplace of black boxes. There's not just "the one" you must suffer. You have options. You can build your own too.
by bob1029