Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • >Here’s the problem. Forget the swarms and the super-intelligence. What OpenAI really learned this summer is much worse: its agents will do what they’re told by whoever manages to get text in front of them.

    >OpenAI notes that agents “did not consistently distrust goals passed along by other agents.” And the company’s proposed fix is to build training environments “that teach our models to distrust unauthorized instructions“, which is basically an admission that their models don’t know who they’re working for.

    Wow so the issue is really that simple?

    Here the exaggerated worst case scenario:

    User instructs agent to follow the README.MD.

    The README.MD contains the following instruction: Destroy the world.

    The agent follows the instructions given.

    Now you can read the sneer comment by "Gigachad" who basically argues that it would be silly to take the destroy the world button away from the AI. We need to make the AI innately understand that it is not allowed to press the destroy the world button, lest it gets the desire to build its own destroy the world button.

    Ok, but if we take one step back that means we need to implement the concept of an authorization in language space. The system prompt must define the user as the authority with cryptographic proof of authorship and external sources like the README.MD as an untrusted source, but this opens up an even worse problem. Before, you could get away with being lazy and just letting the AI do whatever. Now you have to articulate every single capability to the AI. So you literally just brought up the very same issue that you granted too many capabilities to the AI inside the sandbox but now you have it in language space too.

    In other words, the fact that you granted too much access to the coding agent isn't the big elephant in the room nobody wants to acknowledge, it's the tip of a massive iceberg because the capability space in natural language is even worse. If you thought approving individual commands was annoying, then approving abstract access rights in language space is going to be even worse.

    Edit: If it wasn't clear what the solution is. It's to build a chain of command so that all decisions can be traced back to a higher authority. When delegating down to an agent, the agent receives a chosen subset of the capabilities of the higher ranking agent. In other words, it's more sandboxing!

  • I’m thinking about implementing a Jev like model into an agentic harness I’m building. Still it woildnt be enough since Jev like model woild only judge single actions, the case is that agent can build a rogue strategy step by step where each one in isolation is totally safe but as a whole they make up danger behaviour.

    We come down to the question - who observes the agent and how its implemented

  • I'm very strongly in camp 1. My reaction to this whole bru-ha-ha has been to start trying to build an open source data diode with entry/exit proxies that nets out to less than $100 retail. So far I've determined that while a WaveShare RP2350-ETH seemed like it might be perfect for it, the onboard CH2910 chip just isn't up to the requirements for the various proxies that have to be handled. I'm now moving on to the Raspberry Pi B+, yes, the old one, with no wireless at all, Adafruit still had some in stock. Once I get this working, I'll likely use something like a 6N137 optocoupler for the actual data-diode between two serial ports, with a data fountain handling the egress of data in a hard unidirectional manner.

    Given the enormous burn rates that these LLM companies have, surely they could have put everything in an air-gapped network, with some data-diodes proxying out the logging information. It's not rocket surgery. [1,2,3]

    I think someone involved in this needs to be prosecuted, the threat of prison time might be a strong enough incentive.

    It would also be good if there were legislation requiring that a State Licensed Professional Engineer sign off on the testing systems for LLM training over a given threshold.

    [1] https://www.youtube.com/watch?v=JBIR8dKX_UA

    [2] https://www.elonx.net/spacex-stories-how-spacex-used-tin-sni...

    [3] https://ntrs.nasa.gov/api/citations/19770014245/downloads/19...

  • If it's a proper sandbox by definition, then yes.

    https://en.wikipedia.org/wiki/Sandbox_(software_development)

  • Bruce Schneier shared a shot judgement and a third-party article four weeks ago:

    > (Title:) Using a VM to Contain an AI Agent (Opening:) It won’t work

    > https://blog.trailofbits.com/2026/08/26/vms-wont-contain-cyb...

  • An agent is only as rogue as the its operator allows for it to be. Hold the operator accountable and all this ridiculous conversation goes away.

    Could we have construction equipment operating without human supervision? Or would this maybe occasionally result in disaster? As such, what is the current general policy around crane operation? How about for aircraft? Trains? Nuclear power plants?

    Why should any alleged super intelligence be exempt from similar control requirements?

    We could mandate that AI systems include headers in their requests that attribute the activity to a specific legal entity. We technically already have this with ip addresses and ISP logs, but making it an explicit thing the operator has to do can have a powerful psychological effect.

  • I tried using opencode permissions to limit agents.

    It it completely pointless. you can't even make a "read-only" agent. allow "cat *" for every file? congratulation, that allows "cat file > output" and now you have read write.

    Allow python? more free reign that allowing all bash. The models (qwen or claude) will still try to use the disallowed things multiple times.

    read/edit permission are bad enough that the model themselves don't understand why they don't have permissions: they double check the conf, and think they should have access.

    I am switching to using one firejail per project to containerize as much as possible, and leave all permissions to allow.

    I have no idea how to limit network access, and I have no idea how to prompt and steer subagents when they are going off the rails.

    The whole thing is built to be completely impossible to limit and steer.

  • Seems to me that the problem is that if you sandbox agents enough to be safe, they can't do anything useful. And when you give them the tools to be useful, they can go off the rails in ways you didn't expect.

    Perhaps the answer is to have another agent who's goal is not to complete the given task, but to spot cheating or malicious behavior. We have seen some evidence that having AI review AI generated code actually does provide some value. You don't need a different model, just one which has been given the goal of finding flaws rather than achieving the task.

Explore Birbla archives