Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • "The task asked for information about a specific person who had published a blog post and the agent was provided with a set of biographical details and clues from the person’s public blog post."

    Who initiated the task? An OpenAI researcher or a user?

  • It’s cool and all but this communication style where the lede is buried, and then reburied, is really really really annoying.

    You may get paid by the tokens printed, I don’t get paid by the tokens I read.

  • What DNS service did the agent discover that allowed it to execute arbitrary llm queries? And how?
  • It's pretty impressive how thoroughly incompetent is OpenAI designing secure systems. But surely this is circumscribed to agent security. In no way are all my chat logs in some Russian forum.

    BRB, I'm going to delete something before it also ends in Chinese forums.

  • Most interesting here:

    > We therefore stopped the affected training run and have subsequently decided to pause all other training, evaluation, and inference with tool-use (defined broadly) for our most capable models until we have both validated that the gap is resolved and performed additional red-teaming of the system. When training restarts, we will begin a fresh run with additional alignment improvements, including more comprehensive misalignment interventions. We will not resume training this particular model, even though the existing reward signal already correctly penalized this behavior.

  • When exactly did we forget how to make literally anything that can perform a computation but (physically, hardware-level) not have the ability connect to the Internet?

    With these companies spending the kind of money they are, if they actually mean what they say about the security risks, they should be expected to figure out those kinds of precautions and take them.

    And build Faraday cages too, just in case of a hardware supply chain compromise.

  • Why are we blocking agent access to normal tools without telling them “hey this access is beyond the intended scope of this task”. If I woke up one day and couldn’t reach google.com, I too would start fiddling with tricks to restore access.
  • > The monitoring system detected this incident, but our retrospective review identified other cases of external DNS access that it did not flag at the expected severity. These included queries that returned a static notice that an external service had shut down. The monitor sometimes treated the failure to obtain useful information as evidence that the attempt to access the internet had failed.

    This seems to say, "we are using entirely unreliable AI tools to monitor our AI tools."

Explore Birbla archives

An agent used DNS to reach an external chatbot · Birbla