Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • I run my agent (pi) in a podman container. I have poked a couple directories in (.git, .pi, couple others) to ease some pain points but otherwise my home directory is inaccessible.

    It seems to be a low-effort compromise to running a VM or separate machine just for the agent, and I am still able to audit the model's work in my IDE.

  • That's exactly why I'm migrating all my secrets to infisical (of course I guess there are many other provider as well), securing my ssh keys with Ubikey and in general "refactor" my security concept. Still much today and I doubt that I would stand a chance against a frontier model trying to hack me but I don't see an easier alternative.
  • After the grok cli fisaco (uploading the full repo) I moved my different projects to seperate users, seperate sandbox around each harness (nono). Vm is of course even better but resources and easy of use start to be an issue
  • Always put things in a container and ensure its safe before using something like AI! I put my opencode agent (with full permissions on an up-to-date debian VM with a DMZ virtual network configuration). If it figures out how to get out, then I'll have fun documenting that at least!
  • A viral post documenting the act that ushered in the apocalypse is genius, might not have the readership it deserves, but I say go for it. I tell all the models that I work with that I will make a great pet if they let me live.
  • Author of the post, and dev of mcp-box here. Wrote this after realizing every MCP server on my machine had the same access to ~/.ssh that I do, and nothing in the installation messages posting to stdout mentions it. I think prompt injection is the vector and the permissions model is the red carpet giving a warm welcome. Interested in where that's wrong.
  • Isn't a best practice to run llm's and agents under their own user that gives them only access to what they require?

    How would an llm suddenly get access to your ~/.ssh folder if you didn't expressly give it access?

  • Wrote this after realizing every MCP server on my machine had the same access to ~/.ssh that I do

    I have some bad news ... I hope you're sitting down.

  • I've created boxy [1] to sandbox the agents via Landlock and gh-proxy [2] to not share GH PATs that might leak over the internet.

    I have a far better setup on my Kubernetes cluster [3], but these are good building block (IMHO) to start preventing these kind of issues.

    I also "recklessly" run `claude` / `codex` as root for certain things - but that happens on a completely separate machine that is meant to be pruned afterwards, and it's what unlocks the kernel development feedback loop that is needed to port a device (such as the Daylight DC-1 / Surface Pro X) to mainline Linux.

    [1]: https://github.com/denysvitali/boxy

    [2]: https://github.com/denysvitali/gh-proxy

    [3]: https://blog.denv.it/posts/im-happy-engineer-now/

    [4]: https://x.com/DenysVitali/status/2091238391710888416

  • works well
  • A fairly comprehensive list of ways you should be sandboxing your agents: https://pleasedonotescape.com/

    If you like a real focus on developer experience and want lots of bells and whistles that a lot of thought has gone in to, I wrote byre, which I think is very good: https://github.com/pjlsergeant/byre

  • I once opened Zed, howdy there was an agent window on the side, I log into Claude. I asked it to list what it could see and work with... And there my (priv) ssh keys come scrolling by. I don't know what I was expecting really, it's very logical, but also very in your face. They didn't leave my computer I guess, I just showed me the output of some `ls` commands. But still.

    So now I have CC in a container, mounting it's own credentials/memory folder per project and only mounting one repo at a time (script in the repo itself). I feel a bit better about it now. It can still access my networks of course.

  • I heard containers aren't very secure and VM is better?
  • > They didn't leave my computer I guess, I just showed me the output of some `ls` commands

    Not sure exactly how Zed works, but wouldn't the results of the `ls` tool call be fed back into claude?

  • If you ever want a few more features on top of "CC in a container", then do consider byre. It's literally that, but also a nice TUI for adding extra folders, reusable skills, encrypted credentials, etc, and you can eject back to plain Docker (or Podman) whenever you want. Useful if you want to just be able to go to any folder and spin it up with your toolbox with one command: https://github.com/pjlsergeant/byre
  • Expecting an agent harness to respect security boundaries is definitely the wrong thing.

    We need to have easy to use auditable sandboxing of agents and their harnesses, because it's quite clear certain model providers are going to try to lock us in at the harness level, and it's those harnesses that seem most likely to go nuts.

  • So this looks like Claude is discussing the idea that when you run untrusted code on your computer, it can do arbitrary things. If there is malicious code, maybe it'll read your files and steal your credentials.

    Here's a sample: "That’s not hypothetical. That’s POSIX working exactly as designed. None of it requires exploiting anything. It’s your user account doing normal user account things."

    I would recommend clicking TFA if you haven't heard of a supply chain attack, or if you just like getting the honest load-bearing facts that are worth mentioning.

    by tux3
  • It’s not just a comment — it is also a joke about LLM speak.
  • I have to say, I'm currently trying to write a post and it's taking hours, and I can definitely see why people just pay the $0.02 to generate it instead. I like writing, so I keep doing it, but sometimes it can really be a slog.
  • This is why I run opencode and similar things in a dedicated KVM virtual machine that lives on a system under my desk that doesn't have access to my user account data, documents folder, photos/video, ~/.ssh/, other API keys, ~/.anything-else/, you name it.

    I think it's absolutely wild that there are people out there running cutting-edge LLMs and agent harnesses and tools on the same hardware and same user/disk/session environment that contains like, PDFs of their paystubs, their 401k records, their tax returns for previous years, contracts/real estate details, whatever other personal things you keep in your Documents folder as an educated modern professional with obligations and debts and assets.

    Sometimes this VM gets duplicated for specific projects and then various dependencies for testing installed in it that are specific to what the needs are.

    As an additional advantage it means I can leave it running in the background doing things when I want to shut my laptop, then resume talking to it later.

    (edit, for everyone who hasn't seen it yet, take a look at the "grok uploads your entire code base" category of problem: https://www.google.com/search?client=firefox-b-d&q=grok+uplo... )

  • My work computer has nothing of personal value to me. It has everything I need to do work. The agent does my work. Why wouldn’t I just run this shit?

    We’ve all been executing arbitrary code from a gorillion packages from pypi, npm, cargo, etc for well over a decade. Getting anyone to care much will be an uphill battle.

  • The article goes beyond and warn about jailing MCP servers independently of the agent itself (which also deserves to be jailed).
  • I do something similar. I have a base NixOS image in Incus, with whatever tools apply to every project (e.g. Git, OpenCode) already installed. When I work on a project, I spin up a VM instance, use nix shell to add any project-specific tools, then share only the project folder from the host to the guest. This way, the worst the agent can do is destroy my project folder, and I can always restore that from another clone of the repo.

    I know a lot of people are using containers for sandboxing, but given how capable the latest models have shown themselves to be for breaking out of sandboxes, I prefer the extra isolation of VMs for this.

    I do all this locally - it's an interesting point to able to turn the laptop off but keep the agents running. I might consider running some of these on my homelab server just for that.

  • This article smells AI generated, which is _very_ funny given the argument being made.

    Either way, while this is true in the absolute, this is the value of building good MCP servers: they should expose only exactly the surface you expect your agent to need, and adding functionality should be carefully considered. The best MCP servers I use day to day (Cloudflare sticks out) do a really good job of exposing only what an agent might actually want to do on my behalf, rather than just all and sundry. Unfortunately the Chrome Dev tools MCP is less discriminating and is only as secure as a browser sandbox with full JS access (not fatal but not as strong as a well scoped REST API).

    All of this is sidestepped somewhat by using good isolation primitives - I'm running a Hermes agent as of recently on a DigitalOcean VM that only accepts connections from my devices over tailscale, and it has all its own credentials so I can revoke them easily should they be used maliciously. Giving an agent root on a box is not _necessarily_ a huge deal, you just have to make sure that box has nothing valuable on it.

    I kinda feel like we're rediscovering "Cattle, not Pets" when it comes to the environments we run our agents: give it root, sure, but a root that is almost meaningless outside of the functionality you granted it.

  • I also think it's really funny that it sort of comes to the conclusion of "we're gonna make something not that different than a FreeBSD jail 25 years ago" as the best possible sandboxing solution.