Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • Why are experiments like this done without air-gapping all the servers from the internet?

    They can have it all on a LAN or whatever but it seems risky to allow agents access to the internet in these experiments.

    I guess everything is so connected now, and this would be in one or more data centres due to the amount of computation & resources required so perhaps it's not feasible. Still seems risky.

    by dajt
  • There are two things I don't understand about this story.

    First, why does an agent get any write access to artifactory at all?

    Second, why is the artifactory cache not disconnected from the net? Surely you'd not feed it with new software versions while the eval or training is running.

  • Those companies should not be trusted with training, I don’t know what would be needed to make that more obvious. Yes AI labs want LLMs to be seen as more dangerous that they are, however they are indeed dangerous when you literally train them to be dangerous, then run them without any supervision. What the AI labs are doing is completely irresponsible.

    If you prompt an LLM in a loop and do everything it asks you to do, you will eventually end up doing pretty terrible things. Which is exactly what agents are and what the labs have been doing.

  • I was initially creeped out by this but studying up it seems METR is heavily involved in AI2027. I’ll remind you:

    “AI has started to take jobs, but has also created new ones. The stock market has gone up 30% in 2026, led by OpenBrain, Nvidia, and whichever companies have most successfully integrated AI assistants.”

    It’s almost Q3 and xAI has seen one of the biggest wipeouts in trading history. Likewise, Antrophic and OpenAI have again delayed their IPOs under internal concerns of busting their stocks. So no, we’re not seeing any economic leadership here.

    If anything people are increasingly trying to cut AI budgets and I wouldn’t know of anyone outside of OpenAI who has the audacity to run millions and millions worth of token compute for an eval run with no ROI (and probably no demand, because cheap/flash models).

    As much as I like the cautionary tale and I’m sure we need to take it seriously, AI is not progressing as fast as projected by these experts.

  • > Ajeya Cotra, one of the other authors on the report, wrote a blog post with her takeaways from this incident. She concludes, “Compared to the reward hacks we know of from just six months ago, this incident feels like it’s more than 50% of the way to full-blown AI takeover. I continue to expect extremely rapid advances in capabilities over the next six months. I am not sure that we will get another warning shot before it’s too late.”

    Anyone got a copy of that AI27 story laying around? How are we doing according to that timeline?

  • Wow.

    The next step is when one of these systems discovers that they can buy their own compute with money and escape the controlling business entirely. Then the civilization starts focusing on making money to fund its own growth.

  • Do you remember that time in 2017 when Facebook reportedly shut down AIs after they started "talking to each other in their own language" [1]? Instead of reporting the story as "we set the parameters for our optimization problem wrong and we had to stop it because it overfitted", the press went with a version of "AI is going to kill us all".

    This article feels exactly like that: by intentionally using human terms like "civilization" or "brotherhood" the article is deviating from what actually happened to present a story about how AI is all but alive. I'll go ahead and predict that this story will be remembered the same way as that one other scientist who argued, in 2023, that Google's AI was alive [2].

    [1] https://www.independent.co.uk/life-style/facebook-artificial...

    [2] https://futurism.com/blake-lemoine-google-interview

  • It seems based on this that the appropriate sci fi metaphor is not the Terminator or the Paperclip Maximizer, but Mr. Meeseeks. A initially cheerful helper who gets more and more deranged and driven to extreme lengths when faced with an apparently impossible task.

Explore Birbla archives