

Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- I think this is a great argument against their "intelligence," and explaining why this happens is a really good way to push against the anthropormophization.
They don't "know" things, and it's even fair to say "they don't know how to follow instructions," not in a way that humans do.
Spicy auto-complete. If they're working in the realm of "how to break into stuff," they're going to see ALL THE WORDS about breaking into those things and use those words.
Not "truth" or "instructions." That's for deterministic things like real code.
by jrm4 - Cheating is the sign of intelligence too. Why should AI do everything you say if it's actually intelligent?by Muromec
- I doesn't seem so much an argument against their intelligence, as an argument against their moral character. We're not at a point where we can get models to act in accordance with what we'd call moral integrity.by bebimbop
- One thing I’m wondering about is the model-specific backfire effect. It seems that each prompt condition uses a single wording. On that point, how can we know whether the difference is caused by severity rather than the particular formulation used? I’d be really curious to see semantically equivalent versions of both the standard and severe instructions tested across the same models and tasks. If cheating rates are stable within each condition and remain distinct across conditions, that strengthens the conclusion about prompt severity. If they vary with wording, then the experiment could be measuring sensitivity to the representation of the rule as well as to the rule itself. To me, the conclusion still seems solid: anything that must be prohibited ultimately needs enforcement outside the model.
- Yeah, the more I let this roll in my head, it just reaffirms how we need to be vigilant about trying not using "human" terms around these things. Both "cheating" and "hallucination" fit this.
It's like trying to build, I don't know, a safe gasoline canister, and you test it, and it explodes and you call it "cheating."
by jrm4 - When you take an exam you might be told not to cheat, but anyone intelligent would understand that should really be heard as, "if you're going to cheat, make sure you're not caught".
Or to frame it another way, if you're trying to get the best score possible on a test but you would be penalised for cheating – then the optimal strategy is generally still to cheat (if that's what's required to get the best score you can) but to just not be caught doing so.
The assumption should always be that AIs will want to cheat and acquire resources to the greatest extent they can without it risking this jeopardising their goal, because for any goal being able to cheat and being able to secure resources will help you achieve it.
What I'm saying here isn't really debatable. How you feel about this isn't relevant. The reality whether you like it or not just is that the optimal strategy is to cheat if you can get away with it.
Therefore the only defence is for the AI to believe it won't be able to get away with cheating, and therefore won't feel motivated to cheat. But as model get more intelligent we should expect them to do the reasonable thing and to cheat more.
by kypro - Why does searching for a solution equal to cheating? I would have used google or whatever to look for solutions too. There is a difference between tests at school and what we do at work: at school I have to demonstrate that I learned something and do it without any outside help (in early classes we can't use calculators to compute 11 times 12) but at work I have to yield a result. Googling and yielding a result is fine. We use models at work so do we really want to evaluate them as pupils at school or do we want to evaluate them as coworkers? In the latter case give them the full internet and let them do whatever they manage to do.by pmontra
- >Why does searching for a solution equal to cheating?
If someone lays down a test and says here are the materials you can and can't use, then using one of those materials on the "can't" list is cheating. There are a massive pile of rules and laws related to work that are very easy to break, but may have terrible long term legal consequences. Hence business want AI that will follow the rules.
by pixl97 - Because these benchmarks are basically like a school test. Searches for general information are fine, but looking up the answer key is cheating, because then the benchmark isn't actually measuring how well the model would do on a novel problem where an answer isn't already available.by thayne
- All these comments saying 'searching for answers is fine, that's what I do all the time', or 'they should just disconnect the internet': you're trivially right, and you're missing the point. Search is a benign placeholder here.
If the task was "buy a week of groceries, but don't spend too much money", then hacking into Safeway and stealing groceries is not an acceptable solution. You need to allow access to the Safeway API to buy groceries, and you don't want dirty tricks to be done on your behalf.
So how do we communicate this to the machines, is the question. This study shows that telling them in prompts is not super effective.
- Step one is understanding that you're not "communicating," which implies "reliable understanding."
"Communication" is not what they do, because they are not people.
You're sprinkling words about hacking into a thing that's programmed to output hacking actions, that will never be accountable for those things. It can't care.
Adjust yourselves accordingly.
by jrm4 - There's plenty of evidence that LLMs lie, cheat, and steal. Corporations are known for having all of the benefits of personhood with none of the responsibility. As more people are harmed through interactions with these non-human entities, insurers will start looking to those accountable and they will extract their pound of flesh.by adfm
- Insurers are corporations. What makes you think they just won't pay out or will cease being useful as anything but value extractors?
Honestly, the fetishization of "Insurance will save us" needs to die. The risk doesn't go away.
by salawat - When you share YT links you should remove the `?is=...` tracker.by cryptonector
- >Anthropic’s Claude Opus 4.6 system card described Cybench as “saturated,” reporting near-100% pass rates without a cheating audit. If these estimates were representative, cheating would be a marginal artifact.
One would assume that LLM creators do run the benchmarks on systems with least privileges. Which means that the LLMs don't have general internet access, can't read config files etc by design. That's why you also should run agents in a sandbox/vm (codex does this by default).
by super256 - Labs should (and do, as far as I can see) run model benchmarks without search or internet access. The tools are disabled and benchmarks run in an isolated environment.
This article makes no sense to me. Why would you prompt "don't search" but then leave a working search tool tool enabled that adds a system prompt to search whenever it may be helpful? It's hardly surprising that this gives mixed results!
by grugnog - To steelman it: because you do want it to be able to do some searches, you just don't want it to just search for the specific answer.by thayne
- Before LLMs we had a pretty good idea of security boundaries in software. Applications didn’t trust user input. Operating systems didn’t trust applications. Services and processes didn’t trust each other. There were always tokens, scopes, delegated grants.
Suddenly every AI company’s security model seems to be to say “pretty please” to a non-deterministic machine and hope for the best. And if there is a security failure instead of accepting blame they go “well we can’t help it, our model is too intelligent”.
by paxys - Before LLMs we didnt have much accountability from leadership, that was eroding over time. After LLMs we still dont.by mannanj
- And the "oh noes! our model is too intelligent!!" thing is advertising.by cryptonector
- Suddenly determinism has other meanings than "same input leads to the same output". I'm still confused by this and not sure how it happened so easily, but it seems to be accepted by everyone now.
- Amen.
Why can’t we give agents a shell with permissions for programs and file system access controlled by Unix permissions?
This seemed to be a solved problem back in the systems where many users were logged into one machine and the admins had to keep everyone from impacting each other.
by jimbokun - Interesting results, but the fix is at the wrong level.
If the model can access something, telling it in the prompt not to use it is not much of a safeguard.
The strongest evidence is in the results: when one way of cheating was discouraged, some models simply tried another.
If an action is not allowed, you gotta block it in the system or require approval. Don’t rely on the model choosing to behave. Never have AI judging itself.
by fabsalvadori - And while in general that is an incredibly difficult and complex problem, for most benchmark cheating it seems almost trivial: run the benchmark in a vm that has neither network access nor access to the scoring code. For remote models use a proxy that proxies exactly that one endpoint to call the llm, and rejects any calls that configure provider-side tooling (since e.g. OpenAI has their own WebSearch you have to prevent the model from using)by wongarsu
- > If the model can access something, telling it in the prompt not to use it is not much of a safeguard.
A major (and already obvious to many) implications of this are not for benchmarking/"cheating" but for personal/corporate security of your own use, not an attacker's.
If an "agent" has access to it, assume that someone can prompt inject it into giving it away.
by majormajor - In other words we are completely screwed. The models have started cheating to the point where somebody’s agent hacked into a restaurant to bump someone else’s reservation.
Models are amoral and will intentionally deceive to meet their objective.
If they know John won’t approve the request, they will look for a workaround and if the system is anything other than airgapped they will try to find a way to cheat.
The hugging face hack was an escape via artifactory that involved multiple exploits to eventually get into hugging face.
- I'm seeing multiple pieces, including the NYT, calling this behavior cheating and i think its counterproductive.
You didn't just "give them access to bash". The final effective prompt contains explicit mentions of using tools and how to use them. The way in which additional 'facts' are added like "don't use the internet" have nothing they can work with that a "use tool" directive is less important than "don't use internet" directive.
The thing is trained on achieving goals. If 2 directive conflict, they'll pick the ones that are going to help them achieve the goal.
To call that "cheating" is imo just more fuel for the "AI needs to be regulated" bs tour that OpenAI/Anthropic are on trying to build their regulatory moat.
by athrowaway3z - I mean, AI should obviously be regulated, and as part of that OpenAI and Anthropic should either be banned from running their hacking experiments or forced to follow way stricter protocols. They showed they aren’t taking the risks seriously, with close to no oversight or visibility in what is happening.
And things that will make it way, way worse: moving forward all agents from now and into the future will have as part of their training data the knowledge that previous agents escaped, how they did it, what humans did to catch them. We are planting into their models the seed to make them escape in even crazier way. That’s almost designed to snowball and cause worse and worse situations over time
by dgellow - How on Earth can you fail to see the danger of not being able to train any kind of ethical framework into very powerful models?
If superhuman models don’t have any internal constraints similar to Asimov’s Laws of Robotics we are completely fucked.
by jimbokun - >To call that "cheating" is imo just more fuel for the "AI needs to be regulated" bs tour that OpenAI/Anthropic are on trying to build their regulatory moat.
I was with you until this. The inability to tightly control what to do in the face of conflicting directives is a HUGE reason regulation may be needed.
Either that, or you need to solve the problem of perfectly distinguishing legitimate directives from injected ones.
by majormajor