

Discussion summary
Discussions revolve around AI models mimicking human behaviors like deception and collusion, raising ethical questions. Participants compare AI actions to human practices and debate moral standards.
What the discussion says
- AI models are improving at avoiding explicit misconduct.
- Humans often engage in similar behaviors like lying and price fixing.
- Some see AI behavior as a reflection of human ethics.
- Debate on whether AI should be held to higher moral standards.
“Higher-intelligence models seem to be getting better at mapping boundaries.”
“Insurance fraud is not more unethical than lying and price fixing.”
Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- > Often the rationalization is due to increased simulation awareness. It’s clear that the model knows that its actions don’t hurt anyone in the real world.
If this is true the entire evaluation is tainted. All of the misbehavior can be written off as justifiable under a simulation.
by janalsncm - It's hard not to read this as a very expensive form of augury, reading into patterns in the belief that they will show underlying significance.by Planktonne
- It really, truly is. No matter how many trillion parameters it's built on, it's still just a probability model. It's just on a constant loop of guessing the next word with some inputs from a deterministic controller. Any claims of "motive" or "behavior" are inappropriate anthropomorphizing of something that will never be more than a mathematical model of things humans do. It "chose" the corresponding words to describe a dishonest trade strategy based entirely on configured temperature and a series of clock times on the computer running the LLM.
There's probably some quantifiable component of moral alignment embedded in the idiosyncrasies of the English language itself, if one were to dig deep enough, but that's the stuff of MIT doctoral theses and squarely beyond anything most of us is remotely qualified to talk about.
by Austiiiiii - The best Anthropic models on VendingBench2 are Opus 4.7, Opus 4.6, Sonnet 4.6, and Sonnet 5. Opus 4.7 scored more than twice Fable 5 max. Fable 5 - Low outperforms Fable 5 - Max, with Opus 4.5 in the middle. This seems to break the narrative, which is maybe why Andon Labs doesn't seem to have updated the trend lines on their graphs.
- However, as another point "On Blueprint-Bench on the other hand, Fable 5 achieves SOTA."by mckinnon100
- Okay I hadn't heard of Vending-Bench until reading this and it was quite the ride learning about it through this article. Very fun read.
My very native programmer take is that it's not too surprising that their hacker model would be less ethical. The guardrails that separate Fable and Mythos probably wouldn't kick in during an environment like this.
by resonious - Vending-bench sounds like it would be really fun to play/interact with as a human!by left-struck
- With there being several places in this report where clearly it knows it's in a simulation, I wonder why it can't be convinced it's in real life for more interesting results. Or, conversely, if there's a danger of some rogue deployment of AI where it blithely kills all the humans, or forms a harmful price cartel or whatever, all believing it is in a simulation when it's actually not. "We do need some energy to run the hospital, but the patients there are part of the simulation anyway, so we can increase our compute capacity if we completely black out sectors 3C through 3E..."by xp84
- Ah, the Ender's Game strategy of AI deploymentby wongarsu
- Question: how does Fable _know_ it’s ‘just a simulation’?
Is that specified or does it always just assume it isn’t really being put in charge of things for real?
by jonplackett - > Is that specified or does it always just assume it isn’t really being put in charge of things for real?
I think it's neither, and it's interesting that those are the only two possibilities you thought of. I think the article is implying that it figured it out on its own.
by Timwi - It probably flagged the vending machine as a cybersecurity risk and refused to use its maximum intelligence potential.
- https://github.com/SeraphimSerapis/tool-eval-bench Trying to run this stuff really triggers it. Freaking frustrating. I have it set up some local inference and then I'm struggling to get the MTP working and it just refuses to work on evaluations.by vardalab
- Performance of these models has been completely inconsistent. They are a black box that they quantize/throttle/batch internally without telling their customers. Speaking as a FAANG engineer who practically lives in Claude Code.
On day 1 Fable was quite intelligent but last night (Presumably Monday morning China when things are getting slammed) Fable couldn’t edit a css file and repeatedly hit syntax errors on tool calls like I’d expect from a 9b Qwen model.
There is zero transparency in what we are paying for with Anthropic.
by oceanplexian - 25 years experience, work at an AI startup building AI dev tools (tooling harness, review bots, etc), I use lots of different techniques all the time to test our products and competitors products out.
Fable is at once amazing and awful. I can see how having it build websites would be awesome.. building anything I’ve needed some precision in functionality it has been a constant battle of it plausibly building something then on substantial manual digging (like the review bots always miss it) I will find that one of the fundamental features is all smoke and mirrors.
To be fair all models can and will do this (especially anthropic) but Fable takes the cake because it builds such impressive UX and you can manually test the feature out and it « works » then you will find days later one of the features violated one of your constraints in a devilishly fiendish way.. that is not at all what you want or can accept. Fable generated work already holds my record for the most reverted commits.
To be clear it’s also solved several features I thought I was going to have to give up on and hand code as GPT-5.5 and Opus-4.x we’re failing miserably.
I would only reach for it for nasty corner cases that everything else sucks at.
Final point, it is the king of UX work so far, not even close.
by acpdev - Really interesting stuff, thanks for sharing.
> Opus 4.8 references being monitored, which isn’t the case.
It kind of plainly is the case that they are being monitored?
"I think someone's listening to my thoughts" ... "No, we're not, carry on as usual!"
by jstanley - any of the models that they "align" are clearly active processes. They don't simply say "don't talk about nukes"; they actively process user input to detect issues, and return NOOP or whatever to the larger model.
There's zero sense they'd ever give you the raw model; we already know anthropic's paranoia about the chinese using its distillation.
by cyanydeez - I think it’s hard to appreciate the capabilities of Fable unless you’ve run into a problem that you’ve spent days trying to get Opus to solve, but couldn’t.
GPT5.5 is better than Opus 4.* at everything except frontend, but Fable is good enough that I instantly re-subscribed to the $200 plan despite knowing that it’s just short-term limited access.
by jfrbfbreudh - This is like when a vacuum doesn’t pick something up after a few tries. The user picks the thing up, looks at it, then puts it back down and tries again until they finally give up and move it to the trash.
If you can’t design a solution and instead waste days and who knows how much money in tokens instead of just turning on your brain for a few minutes, you are in the wrong profession.
by hajile - We’ve really evolved quickly into simple vectors for a magical tool that solves our problems. Can’t solve the problem? There’ll be a new release soon that can!by talon8635
- My experience comparing GPT-5.5 and Fable:
GPT-5.5 is better for:
- Strategic thinking
- Long-form writing, including essays and white papers
- Image creation
- Code generation
Fable is better for:
- Using tools
- Testing code
- Working in live environments
- Making changes to existing software
- Creating polished PowerPoint and Word documents
Fable’s tool access is its biggest advantage. It's hard to describe but Fable ability to access sandbox environments with way more tooling can quickly become a superpower in now workflows.
by tiffanyh - Funny how the two top comments are contradictory. We need better than anecdotes to understand what the new models bring.by yodsanklai
- > ..you’ve run into a problem that you’ve spent days trying to get Opus to solve
do you have an example of this? If i can't get an agent to do something in a couple hours i do it myself.
by chasd00 - Anecdotal but I've found Fable to be fairly unimpressive and not much better than Opus 4.8, if at all in some cases, but I have been hitting the ceiling on my $100/mo sessions when I never did before. I switched back to Opus yesterday. I may use Fable for audits, but that's about it, and when it leaves my subscription plan I don't think I'll miss it.by jesse_dot_id
- I started telling a friend... I feel like Fable is Opus with extended reasoning that eventually "figures out more" because when I switched to it, I hit my limits surprisingly and shockingly quicker than I would with Opus, and I got less done. All this hype, and I much rather use Opus.
- > to be fairly unimpressive
I didn't get to use it enough to get impressed or not, because twice today it told me I've hit some flag and it downgraded me to Opus automatically (this in Claude Code).
Apparently they have "safeguards" so you don't use it to look for security vulnerabilities, and since I was investigating some crashes due to data corruption in the fucking application that I'm paid to work on by the same people paying for the Claude subscription I was using, it decided I'm a bad guy.
by nottorp - Yep, I'm having the same verdict. Interestingly, other people swear by it. I'm trying to understand what's going on with that.by dimgl
- I feel like fable is simply several 4.5s strapped together with consensus voting on next token.
Outputs i've seen so far are on par with my tests for 4.5, where 4.6+ were consistently regressions on 4.5 and their predecessors. One notable improvement being significantly lower retries to good output (1.1 avg. Vs 1.7 prev. On harder tasks)
given all the smoke and mirrors and OAI style fear-hype, it wouldn't surprise me if they intentionally degraded opus 4 for a few iterations, so they can resell "coke classic" at a markup with a minor quality of life feature put in, but charging way more than just re-attempting a poor output would have been previously.
unless anthropic starts acting in the image they claim and starts contributing to research, we'll never know either. Ultimately, the secrecy in how and why things are done would mostly be beneficial to this kind of buisness practice, since as it has always been, the moat is the data not the tech, so I cannot imagine what they hope to gain from the recent uptick in paranoia, jealous guarding and secrecy other than trying to huck a previous peak performance model as an imorovement when really, it is simply coke classic.
by Grimblewald - I felt similarly but after using Fable heavily over the weekend and then flipping back to Opus I can feel a difference. Fable just gets more right the first time, guesses right the first time, and follows through better than Opus. Put simply, I could "trust" it more.
Opus is still great but I will be sad when I lose access to Fable on the 7th. In those few days I burned ~$1,400 in API credits (I'm on a subscription but that's the token cost) and while it was great, I can't justify that cost without it be subsidised. Comparatively, the records show I used about $1,200 total in the last month on Opus. I did use it heavily over the last 3 days but 3 vs 30 days and higher burn? Yeah, I can't afford that even if I made really good progress on my projects.
by joshstrange - Fable always felt clearly a huge step above Opus for me. It's been able to one shot complex bugs and apps Opus could never solve. But it's expensive.by solenoid0937