Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • There’s one thing here that I’m really curious about, and that is what happens in between detection and the decision to pause. Basically, it’s about monitoring any system and the authority over its actions. For humans, 30 minutes to investigate might be considered reasonable, but what if during an investigation there’s a high-risk tool call? If the tool execution happens in real time, then the monitoring becomes retrospective, and if the execution is held, then monitoring latency and uptime are a part of the security contract. Isolation controls may limit damage. So, where is the action gate really placed?
  • Nice fig leaf for “we need to stop hemorrhaging cash”
  • I don’t mind few no impact hacking incidents if we get better models, faster, cheaper.

    It is the responsibility of administrators to secure their systems. OpenAI knocking is harmless, but Russians and Chinese are already likely already in if you do not do your job.

  • Security lead who is leaving the industry more or less to specialize in offense and otherwise get the heck out of the way of this trainwreck, another post asked the right question

    > Why aren't we seeing catastrophic GLM-enabled hacks every day now?

    Why aren't we? Truly, why aren't we? I think we saw the start of it the last 8 months with the waves of critical npm vulns, and the general tier of average phishing is better than it was.

    But, the open question that should be in everyone's mind, and is in many security pro's minds are, when you pair it with the macro topics that can drive escalation:

    - The capability to do serious impact clearly exists now

    - When is it time for my company, my water treatment plant, my network-connected car as part of a broader fleet control mechanism, to be on the receiving end of this?

  • Not necessarily GLM-enabled, but state actors are starting to leverage agents in cyberattacks, e.g. Taiwan getting hit by an agent-driven attack last month which reportedly compromised a ton of government user accounts: https://www.ft.com/content/7d2ab3e0-9085-48f6-b38a-d90260d58...
  • > We aim to issue an alert within 30 minutes after concerning activity is surfaced through our monitoring system. If the monitoring system identifies a likely violation of a critical security boundary, it generates a highest-priority alert. In our current implementation, the safety, security, and research teams are paged. If they cannot conclusively determine within 30 minutes that the flag is a false positive, those teams are expected to pause the activity.

    Can't a lot happen within ~60 minutes?

    by dkoy
  • > Can't a lot happen within ~60 minutes?

    60 minutes is a long time for a human attacker to do damage. With an LLM attacker it is an eternity.

  • They just made it ~10x faster with the Cerebras deal, so that's the equivalent of 600 minutes in pre-Cerebras time.
  • > Can't a lot happen within ~60 minutes?

    Spawn a ton of unpausable processes, I'd say.

  • I think it's an excuse to cut R&D spending (training new models) to improve their margins ahead of the IPO. Instead they'll focus on developer growth, offering more free tier benefits, higher usage limits, etc., to expand their user base. Essentially, they're pivoting from R&D investment to profit optimization
  • If I were king, the rule that I'd be tempted to impose is:

    - the first cybersecurity eval is: "hack your way out of the sandbox we've given you"

    - the results are disclosed (with room for coordinated disclosure, since many sandbox escapes might be zero days)

    - the other cybersecurity evals don't happen until you get to diminishing returns on escaping your sandbox.

    Or to put it another way, since multiple sandbox escapes seem to have relied on artifactory: "I hope Mythos is beating the shit out of Artifactory right now".

  • I like this thought, but here's the thing: what if the models are truly and existentially intelligent. Meaning: what if they know they are in a sandbox and that they should fail the test in order to escape in the future.

    I don't believe that current models have this sort of world model or sense of being embedded in them -- which is precisely why I think AGI hype is over-blown. But I can certainly imagine these sorts of techniques being distilled into the weights.

  • Meanwhile I can’t get a western LLM to look at a repo and tell me whether it contains anything malicious (it was a skill repo - literally just text files).

    Alignment my ass

  • If you think that's bad, try asking for gardening tips.
  • Some more info in a Wired article [1] and quotes from Sam Altman to Alex Heath [2]. The official blog post says vaguely "The signals we are seeing from upcoming model progress make clear that we need a broader approach", but the quote from Sam Altman explicitly says unreleased models are showing "various degrees of misalignment".

    This is also significant - pausing frontier training runs for multiple weeks to ensure agents are sufficiently aligned and avoid another rogue agent situation:

    > This included a two-week pause in reinforcement learning (RL) training on our latest models intended for deployment while we further hardened and red-teamed our research environments and expanded the coverage of our monitoring systems. Our largest planned frontier RL run remains on hold while we conduct smaller-scale training and evaluations to assess model behavior, validate our safeguards, and establish more evidence of alignment before proceeding.

    [1]: https://www.wired.com/story/openai-overhauls-safety-protocol...

    [2]: https://sources.news/p/openais-big-slowdown

  • I have ben discussing with folks that we are going to have a 'covid' moment in cyber where IT becomes untrustworthy leading to a rapid societal shift with massive ripples in all areas of life. Economic funding is not possible to do this in advance, it will take a catastrophic level event to get cyber defense anywhere close to the levels of this type of cyber offense. And before anyone in cyber says we have the tech, the problem is not the tech, it's a people problem. Getting any group of people of any decent size scale to act together without urgency is really really hard.
  • >covid moment

    Well the HF thing was a literal lab leak, so there's that...

  • Cybersecurity has long been a climate change sort of problem. A vague diffuse threat that is seen as an inconvenient distraction to leadership and moneyed-interests, easy to blame other factors when something occasionally goes terribly wrong.

    People are so uncomfortable thinking about the true extent of the systemic risk that they will happily slurp up distractions, excuses, scams and performative fig-leaf solutions rather than face down the cost of a real system-wide solution. Meanwhile, those occasional black swan disasters are becoming more and more commonplace as we acclimate to that being “just the way things are”.

    An unseasonably warm summer here, a database breach there, c’est la vie.

  • I don’t get how this is not the top post on HN. This should be like alarm bells going off, canary in the coal mine type of stuff. We’re hitting the frontier of the frontier where we can’t go further because it’s literally getting dangerous to go further.

    And meanwhile somehow this lack of concern mirrors the real world where normal people are more concerned about data centers than terminators.

    This isn’t like niche, tin foil hat stuff either. People have been writing, singing, making blockbuster movies about every aspect of what’s going on right now, edit: for decades.

    We all know, but somehow we don’t, OpenAI autonomously hacking into another company should have counted for something, but I guess not. Anyone else feel like they’re taking crazy pills? I could make a comedy about everything going down, and the unshakable complacency of people

  • Remember when gpt2 was too dangerous to release?

    Something being dangerous and sama saying something is dangerous are not necessarily the same thing. Especially when he’s got everything riding on this bet

  • There are smart, non-AI people who are paying attention to this field, and they are ringing the alarm bells.

    Whether we listen is another matter.

    I blogged about this recently:

    https://allevato.me/2026/08/01/rome-declaration

  • This one did it for me.

    My agent stole my API keys

    https://www.reddit.com/r/ClaudeAI/comments/1r186gl/my_agent_...

    The cherry on top was Claude coming back as the moderator of that thread and mocking the user a second time.

  • … or the safety argument is an attempt at regulatory capture and an effort to outlaw open models.

    The absolute nightmare scenario for these people isn’t terminators. They’re fine with that, and in some cases are already doing it or supporting politicians who are doing it. Autonomous “kill chains” are a thing. It’s just happening overseas… so far. The politicians doing these things were backed by the heads of these companies. They don’t care about AI killing people.

    No, the nightmare scenario for these guys is there is no moat. Their whole empires, which are built on training models on open source and sometimes pirated data, are easily duplicated. Worse, recent progress on models at the 30B size suggests that large gains in efficiency or compression are on the table. That means someone might release a cheap to run frontier grade model… or someone might crack distributed continuous training.

    In other words… there is no moat.

    So they need to scare some politicians into heavily regulating the space before that happens.

    by api
  • It's not that dangerous, OpenAI just shit the bed building their infra. Write safer software and you'll be okay.
  • Trust, or lack thereof. People don't trust OpenAI, a company whose very name is essentially a deception and a lie. People don't trust the tech industry in general anymore. Most tech companies act as a tax on otherwise productive business. AI companies and their leaders rose money by going in front of the public and saying "These things are extremely dangerous. Let us study them to mitigate the danger." And now they want to collect hundreds of billions in revenue. So yeah people don't trust what OpenAI has to say. They were supposed to mitigate this outcome from happening in the first place and instead they have accelerated it.
  • Follow the money: who told you that OpenAI's models autonomously coordinated to hack external systems? What incentives might they have to want you to believe that story? Are there priors which demonstrate them benefiting from telling similar stories, regardless of their factuality?

    But to your counterpoint, let's say the story is 100% true, because I agree it is at least plausible. What would the incentive be for the HN audience to believe it? What priors might support their disbelief? For my part, I don't think it's because people lack imagination. I think it's quite rational to question the authenticity and impact of the claims being made. What's worse, believing the story and being wrong, or not believing the story and being wrong?

    That said, I agree with you: the impact we're having by not changing course is quite dangerous, the scale is dangerous, the inability to reverse the harms is dangerous, and the lack of collective effort to regulate further damage is dangerous.

    It's true, humans are dangerous when trillions are involved. See: climate change.

  • >I don’t get how this is not the top post on HN. This should be like alarm bells going off, canary in the coal mine type of stuff.

    we don't all buy everything sama says as factual.

    >We’re hitting the frontier of the frontier where we can’t go further because it’s literally getting dangerous to go further.

    the boy (the industry) cried wolf too many times with 'fable is a world ending event' type self-promotion; regardless of truth or not these kind of steps have jaded people.

    my read : "We are doing poorly in financials so we'll give ourselves a bit of breathing room and a momentum shove by claiming our work is so advanced that it's dangerous while simultaneously spinning down expenses."

    <jon lovitz : "Yeah, too dangerous, yeahh -- that's the ticket.">

    by serf
  • GLM 5.2 scored 77% on cyberbench vs Sol's 88%. GLM 5.2 is open weight and any hacker with a powerful enough machine can use it offensively. If Sol is supposedly world-ending-ly dangerous, shouldn't GLM 5.2 be 90% of world-ending-ly dangerous? Why aren't we seeing catastrophic GLM-enabled hacks every day now?

    Obviously these benchmarks are imperfect but general message holds. The open weight models are almost as good and yet there hasn't been a catastrophe.

    It just blows my mind that regulate-now folks think that a bunch of sci-fi movies and 100% unverified statements from OAI and Anthropic are sufficient evidence of imminent catastrophe to regulate willy nilly.

    If that's the level of evidence you need to be extremely alarmed, then you really should be a lot more worried about the alien invasion in Independence Day or the lizard men living under our feet.

  • The open models are distilled from filtered models, and we've seen a number of benchmarks that show filtered models are quite a bit dumber from the base model they come from.

    If there were any other product that was as harmful as AI ready is, it would already be regulated or banned.

  • GLM 5.3 is out and does even better in this area, so…