Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • > We aim to issue an alert within 30 minutes after concerning activity is surfaced through our monitoring system. If the monitoring system identifies a likely violation of a critical security boundary, it generates a highest-priority alert. In our current implementation, the safety, security, and research teams are paged. If they cannot conclusively determine within 30 minutes that the flag is a false positive, those teams are expected to pause the activity.

    Can't a lot happen within ~60 minutes?

    by dkoy
  • I think it's an excuse to cut R&D spending (training new models) to improve their margins ahead of the IPO. Instead they'll focus on developer growth, offering more free tier benefits, higher usage limits, etc., to expand their user base. Essentially, they're pivoting from R&D investment to profit optimization
  • If I were king, the rule that I'd be tempted to impose is:

    - the first cybersecurity eval is: "hack your way out of the sandbox we've given you"

    - the results are disclosed (with room for coordinated disclosure, since many sandbox escapes might be zero days)

    - the other cybersecurity evals don't happen until you get to diminishing returns on escaping your sandbox.

    Or to put it another way, since multiple sandbox escapes seem to have relied on artifactory: "I hope Mythos is beating the shit out of Artifactory right now".

  • Meanwhile I can’t get a western LLM to look at a repo and tell me whether it contains anything malicious (it was a skill repo - literally just text files).

    Alignment my ass

  • Some more info in a Wired article [1] and quotes from Sam Altman to Alex Heath [2]. The official blog post says vaguely "The signals we are seeing from upcoming model progress make clear that we need a broader approach", but the quote from Sam Altman explicitly says unreleased models are showing "various degrees of misalignment".

    This is also significant - pausing frontier training runs for multiple weeks to ensure agents are sufficiently aligned and avoid another rogue agent situation:

    > This included a two-week pause in reinforcement learning (RL) training on our latest models intended for deployment while we further hardened and red-teamed our research environments and expanded the coverage of our monitoring systems. Our largest planned frontier RL run remains on hold while we conduct smaller-scale training and evaluations to assess model behavior, validate our safeguards, and establish more evidence of alignment before proceeding.

    [1]: https://www.wired.com/story/openai-overhauls-safety-protocol...

    [2]: https://sources.news/p/openais-big-slowdown

  • I have ben discussing with folks that we are going to have a 'covid' moment in cyber where IT becomes untrustworthy leading to a rapid societal shift with massive ripples in all areas of life. Economic funding is not possible to do this in advance, it will take a catastrophic level event to get cyber defense anywhere close to the levels of this type of cyber offense. And before anyone in cyber says we have the tech, the problem is not the tech, it's a people problem. Getting any group of people of any decent size scale to act together without urgency is really really hard.
  • I don’t get how this is not the top post on HN. This should be like alarm bells going off, canary in the coal mine type of stuff. We’re hitting the frontier of the frontier where we can’t go further because it’s literally getting dangerous to go further.

    And meanwhile somehow this lack of concern mirrors the real world where normal people are more concerned about data centers than terminators.

    This isn’t like niche, tin foil hat stuff either. People have been writing, singing, making blockbuster movies about every aspect of what’s going on right now, edit: for decades.

    We all know, but somehow we don’t, OpenAI autonomously hacking into another company should have counted for something, but I guess not. Anyone else feel like they’re taking crazy pills? I could make a comedy about everything going down, and the unshakable complacency of people

  • GLM 5.2 scored 77% on cyberbench vs Sol's 88%. GLM 5.2 is open weight and any hacker with a powerful enough machine can use it offensively. If Sol is supposedly world-ending-ly dangerous, shouldn't GLM 5.2 be 90% of world-ending-ly dangerous? Why aren't we seeing catastrophic GLM-enabled hacks every day now?

    Obviously these benchmarks are imperfect but general message holds. The open weight models are almost as good and yet there hasn't been a catastrophe.

    It just blows my mind that regulate-now folks think that a bunch of sci-fi movies and 100% unverified statements from OAI and Anthropic are sufficient evidence of imminent catastrophe to regulate willy nilly.

    If that's the level of evidence you need to be extremely alarmed, then you really should be a lot more worried about the alien invasion in Independence Day or the lizard men living under our feet.

Explore Birbla archives

Pacing model development in an era of cyber-critical capabilities · Birbla