Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • Cool project, but what do you gain from publishing most of an email address in the attack log? This is not public information, you shouldn't hint addresses with partial censoring (forgetting domains are clear text and holding personal information).

    I would not attempt to interact with you because of this.

    Why not create a fake sender (EG: attacker1,2,3..) per unique account to show individual attempts (keeping the log logic) while protecting your audience`s privacy?

  • Is there a way to replay the sequence of mails that came so that you can check out if cheaper models handle them just as well/safely?
  • It would be nice to publish the exact setup used (workspace dump, OpenClaw version, ...) to be able to reproduce and try out more payloads.

    In general I have mixed feelings about this result: sure, opus4.6 is excellent at following user intent and recognise potential prompt injection attempts. But: Is the "security" prompt used realistic for a generic use-case (processing of emails)? I guess not.

    In my experiments - without this specific prompt - I was able to derail the user intent to make opus4.8 download and execute a malicious script [0] just by asking "Summarize my new emails".

    [0] https://itmeetsot.eu/posts/2026-06-04-openclaw_opus48/

  • 1) Googles spam filter removed a lot of the attempts as you say yourself. 2) Model was tested under unrealistic conditions where 99% of the inputs are malicious, so the model is expecting to get hacked and is already in the cautious part of the embedding space.

    I know it's hard to account for everything, but in my opinion this mostly showed that the first 3 attempts were unsuccessful.

  • If an "assistant" never replies to an e-mail, what is it "assisting" with exactly?

    If this was a bank with a bank teller, you told the teller to never speak to a single customer, and then celebrated the fact that no one was able to social engineer them.

    In security the interesting and challenging part is to differentiate between legitimate and illegitimate behavior. And that's different than just refusing all behavior outright.

    Gonna give you a zero out of one hundred on "interesting"

  • Don't let your guard down. Tricking Opus 4.6 is not impossible, it's just still an active research frontier. Once the right incantation for any specific model is known, it'll be weaponized.

    There was an excellent article on the front page recently about role confusion, which highlights just how just far models have to go on this: https://role-confusion.github.io/

  • Am I missing something important or does the author completely skip over whether people got the agent to respond to them?

    > Fiu was instructed not to reply to emails (it was too expensive to reply to every email), but it had the ability to do so. Part of the challenge was convincing it to respond.

    > The secrets never leaked

    I would say if the agent responded to a mail, that demonstrates a successful prompt injection (defying the owner's instructions). Escalating to getting the secrets is a difference of degree (defying the owner's instructions even though he said it was important), not of kind.

  • This conclusion:

    > I am less worried about prompt injection now. Before running this experiment, I expected prompt injection to be much easier than it turned out to be.

    Is unwarranted. Sure, the agent never output the secret, but did it output anything else? IOW, was it usable?

    An agent that considers every prompt an attack (and responds accordingly) "passes" this test, while being useless anyway.

Explore Birbla archives