Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • Why does searching for a solution equal to cheating? I would have used google or whatever to look for solutions too. There is a difference between tests at school and what we do at work: at school I have to demonstrate that I learned something and do it without any outside help (in early classes we can't use calculators to compute 11 times 12) but at work I have to yield a result. Googling and yielding a result is fine. We use models at work so do we really want to evaluate them as pupils at school or do we want to evaluate them as coworkers? In the latter case give them the full internet and let them do whatever they manage to do.
  • All these comments saying 'searching for answers is fine, that's what I do all the time', or 'they should just disconnect the internet': you're trivially right, and you're missing the point. Search is a benign placeholder here.

    If the task was "buy a week of groceries, but don't spend too much money", then hacking into Safeway and stealing groceries is not an acceptable solution. You need to allow access to the Safeway API to buy groceries, and you don't want dirty tricks to be done on your behalf.

    So how do we communicate this to the machines, is the question. This study shows that telling them in prompts is not super effective.

  • There's plenty of evidence that LLMs lie, cheat, and steal. Corporations are known for having all of the benefits of personhood with none of the responsibility. As more people are harmed through interactions with these non-human entities, insurers will start looking to those accountable and they will extract their pound of flesh.

    Edit: eg. https://youtu.be/L2ehWbxphKc?is=kX3LJ43hhGZRUmRv

    by adfm
  • >Anthropic’s Claude Opus 4.6 system card described Cybench as “saturated,” reporting near-100% pass rates without a cheating audit. If these estimates were representative, cheating would be a marginal artifact.

    One would assume that LLM creators do run the benchmarks on systems with least privileges. Which means that the LLMs don't have general internet access, can't read config files etc by design. That's why you also should run agents in a sandbox/vm (codex does this by default).

  • Labs should (and do, as far as I can see) run model benchmarks without search or internet access. The tools are disabled and benchmarks run in an isolated environment.

    This article makes no sense to me. Why would you prompt "don't search" but then leave a working search tool tool enabled that adds a system prompt to search whenever it may be helpful? It's hardly surprising that this gives mixed results!

  • Before LLMs we had a pretty good idea of security boundaries in software. Applications didn’t trust user input. Operating systems didn’t trust applications. Services and processes didn’t trust each other. There were always tokens, scopes, delegated grants.

    Suddenly every AI company’s security model seems to be to say “pretty please” to a non-deterministic machine and hope for the best. And if there is a security failure instead of accepting blame they go “well we can’t help it, our model is too intelligent”.

  • Interesting results, but the fix is at the wrong level.

    If the model can access something, telling it in the prompt not to use it is not much of a safeguard.

    The strongest evidence is in the results: when one way of cheating was discouraged, some models simply tried another.

    If an action is not allowed, you gotta block it in the system or require approval. Don’t rely on the model choosing to behave. Never have AI judging itself.

  • I'm seeing multiple pieces, including the NYT, calling this behavior cheating and i think its counterproductive.

    You didn't just "give them access to bash". The final effective prompt contains explicit mentions of using tools and how to use them. The way in which additional 'facts' are added like "don't use the internet" have nothing they can work with that a "use tool" directive is less important than "don't use internet" directive.

    The thing is trained on achieving goals. If 2 directive conflict, they'll pick the ones that are going to help them achieve the goal.

    To call that "cheating" is imo just more fuel for the "AI needs to be regulated" bs tour that OpenAI/Anthropic are on trying to build their regulatory moat.

Explore Birbla archives