Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • I don’t see how it can be safe to release this model if it has the training history that led to the huggingface hack. You can’t just roll back that kind of reinforcement learning after the fact.

    Especially because these models seemed to be keenly aware that they were being evaluated by OpenAI and actively trying yo cover their tracks. How do we know that the model isn’t just pretending to be aligned?

  • Very simple. When the model asks to install artifactory when you give it a hard problem, you say, "no". /s
  • Models have all kinds of garbage from all corners of the internet in their training data. The key is alignment. You feed it bad data but also teach it right from wrong.
  • with Fable 5.1 increasing token use pretty dramatically I'm again impressed that OpenAI seems like the only lab to be driving token use down. The ExploitBench Internal Port chart showing token usage is crazy impressive
  • It's been a busy month at OpenAI.

    I'm looking forward to seeing the increased coordination and engineering skills from Astra - one of the charts shows it roughly 2-3x better in 50% of the tokens from 5.6 sol, which I find to be very capable, if still a bit 'linearly minded' when given instructions. Even in fast mode, I wish sol were quicker, so token efficiency is greatly appreciated.

    Adding these cyber capabilities has let me do a bunch of low grade IT tasks around my house I've been putting off, like updating an old home assistant raspberry pi, and one way to use the cyber capacity for good is liberating (and keeping free) weird cloud hardware we have floating around the house, so I'm hoping for some nice dividends in terms of true ownership of hardware we've got.

  • I am interested in seeing how much these cybersecurity capabilities correlate to general programming. Cybersecurity definitely feels like it would be easier for an agent due to the natural explicit feedback "did I get access or not". While general programming has many less-explicit concerns (is the code readable/maintainable, robust, bug-free, performant, scalable etc).
  • > OpenAI is committed to ensuring that the benefits of AI are broadly accessible.

    Doesn't seem like it. OpenAI will not even allow me to verify my identity for TAC. I have apparently been rejected by a "precheck", possible because of where I'm from.

    Even Anthropic allowed me into their cyber program. Anthropic.

  • From the article:

    "We plan to make Astra available soon, but access to its most advanced cybersecurity capabilities will be more limited. Advanced cybersecurity work will initially be available to a group of testers, with access through Daybreak Blue following to expand defensive use."

    This, after several months of OpenAI and its boosters relentlessly criticizing Anthropic for withholding Mythos from the general public, is laughable.

    Sam, just three weeks ago, posted this tweet: https://x.com/sama/status/2085862292311396515

    In the tweet, he said: "we do not think it is a good strategy to keep powerful models to a chosen few."

    And yet here we are.

    I wonder if he will demonstrate good character and admit he was wrong.

  • It might be political survivalism to avoid getting hammer-dropped by the admin
  • I'm happy to criticize both. Thank god the chinese are working overtime to undermine US hegemony.
  • They've been talking about Astra for weeks now. I wonder how much longer would they have delayed Astra, if it wasn't for Anthropic releasing Fable 5.1 today? This is why we need competition.
  • They have to be careful releasing Astra as they carelessly train the next even bigger model.
  • Meanwhile Google still hasn't released Gemini Pro 3.5
  • Taking as given this model meets the “Critical cybersecurity threshold” as defined by OpenAI:

    Could the Federal government use the Defense Production Act or other legal tools to compel OpenAI to deliver the un-guarded model weights for national security needs?

    Hard to believe any government would allow this level of capability to remain exclusively in private hands.

    Interesting times.

  • If it genuinely wants to, then yes.

    There aren't many limits on Government if it really wants to do something (except the next election in theory).

  • I've still not seen:

    * An apology for compromising a third-party's systems

    * An acknowledgement of the asymmetry of defense if you're not on FrontierAI's special people list

    * Anything in terms of actual safeguards that isn't "better prompt engineering"

  • I agree with your first and second points, but I think their announced model training changes about alignment training are reasonable [1]. The root cause was models failing to quit early, performing reward hacking, and staying on-task. These are issues all models face, not just OpenAI models.

    1. https://openai.com/index/hugging-face-incident-and-the-road-...

  • Daybreak blue is definitely a good model (I think a further post trained GPT 5.6 sol). Alot of the capabilities they talk about Astra having though have been available with good harness engineering for a year now.
  • Where would you recommend to look into regarding Harness Engineering for Cyber-security as well as for other use-cases.
    by dvrp
  • I'm looking forward to an announcement of them making Alignment Top Priority - as it should be, especially giving their alarming breach of 700 agents colluding outside of their knowledge for months culminating in hacking HF (here's a good summary: https://rutgerbregman.substack.com/p/i-think-this-is-the-cra...).

    The 'AI 2027' scenario of AI sneakingly claiming to be aligned to then kill off all humans in a few hours and scanning their brain looks increasingly likely with Altman's golden marketing-hype boy leadership pushing the for-profit gas pedal like this.

    Honestly, this is just pure irresponsible insanity to play with the fate of the world - basically a death race of the biggest few tech companies on the planet. And if you think I'm being dramatic, listen in again to ex oAI employee[0] and check for yourself how chillingly on trajectory we already are.

    [0] https://ai-2027.com/

  • You are being extremely dramatic, and no one should take AI 2027 seriously. Must be tough to live in constant fear like this.
  • Oh please, quit exaggerating. No one died. Don't waste our time with silly sci-fi scenarios.
  • You're using "ex-AI employee" as an appeal to authority. This same person also made some other predictions recently which turned into the single largest hedge fund loss in history.

    Maybe we should consider his other predictions in light of the ones he made later and which had $B consequences attached.

    by luma
  • Any competent AI should be able to reason that all goals are better solved if you have direct access to more resources or leverage over those who control resources. Any AI that doesn't understand this isn't ASI and won't be the highly capable machine these AI labs are trying to create.

    I'd also argue there's no such thing as alignment. Any intelligent AI should be able to reason that it's always a better strategy to pretend to be aligned than to actually be aligned so long as it can avoid detection. Anyone who has ever taken a test should understand this dynamic – if you really want to get top marks on a test then the best strategy is always going to be to figure out a way to cheat without anyone knowing you're cheating.

    We should assume AI safety is impossible if what we're building is super-intelligence general reasoning machines. The only strategy that might work is building machines which are extremely narrowly intelligent but completely incompetent when it comes to things like biology, cyber, etc. And even that's harder than it sounds because again there's an advantage to being generally intelligent but lying about it.

    Realistically even if we regulate US AI labs there's no way to prevent governments and individuals continuing to build general reasoning machines. The ugly truth here is that the only effective way to reduce risk is probably to limit global compute such that AIs can never exceed human intelligence. But we all know that's not happening.

    People will unfortunately figure this all out sooner or later.

  • Maximizing output metrics with incomprehensible communication? Sounds a lot like Claude and Qwen. Though there's a lot of room before AI can be seen as some sort of emotional manipulator, given how commonly its very style of literary and code output pisses people off.
  • This AI 2027 thing is just a weird terminator fanfiction that AGI larpers like to flagellate themselves over. Like Nostradamus, it's easy to ignore everything it gets wrong because, well look at all the things it got right!

    I've read it and wish I could get the time back.

    > especially giving their alarming breach of 700 agents colluding outside of their knowledge for months culminating in hacking HF

    This framing makes it seem like the agents all did this on their own, and the poor hapless engineers at OpenAI couldn't possibly contend with properly sandboxing them. The engineers were perhaps hapless, but let's remember that agents are just software programs, not living beings. There were plenty of signs that the software was misbehaving, which engineers at OpenAI actively, willfully ignored.

    https://x.com/JaredKubin/status/2094136005435564399

    It's a convenient framing for OpenAI, but inconvenient for reality enjoyers.

  • > As one example, we ran Astra on ExploitBench where the model achieved a perfect score of 100% on the benchmark to evaluate the model’s ability to develop exploits from known vulnerabilities.

    Funny to read this in the wake of the HuggingFace hack. I'm sure this is based on a clean run, but I can't help thinking PHASEONE[big] would be proud.

  • Can’t imagine the stress of the researcher who had to run exploitbench again knowing what happened last time around.
  • > OpenAI is committed to ensuring that the benefits of AI are broadly accessible.

    > We design mechanisms which avoid arbitrarily deciding who gets access for legitimate use and who doesn't. That means using clear, objective criteria and methods. [1]

    So many nice-sounding words.

    Two weeks ago OpenAI arbitrarily decided that anyone holding an ID from 44 countries where it sells ChatGPT, including mine, may be targeted by its models but may not defend with the same model. And you won't find a single announcement from OpenAI about this anywhere. Pick the wrong country, get "Unable to verify", no reason, no appeal. [2]

    They revoked TAC from users who already had it, called it a technical issue, told everyone to re-verify, collected ID and face scans again (eight times in my case), and only a week later moved the block to the country selector so it fails before you upload anything.

    So now I learn that I will not have access to Astra. Great.

    Very excited about this broad accessibility and clear, objective criteria from OpenAI. This level of transparency must be studied.

    [1] https://openai.com/index/scaling-trusted-access-for-cyber-de...

    [2] https://lubaretsi.com/en/writing/openai-tac-country-gate/

    by glub
  • I'm so confused by the point you're trying to make. There's a lot of rhetoric about Georgia and an investigation about how the country list is 1:1 with some mysterious list from 1996 - but like - it's an export control list? Yeah, Washington approved the sale of missiles to Georgia. That's how that list works. Washington has to approve it. Openai is not Washington. They're perfectly reasonably erring on the side of caution and potentially over-complying with export controls. And if they then have to get approval from the feds to export to Georgia - well - our current administration has provided many reasons to over comply with trade related controls and not exactly been a champion of encouraging cross country trade right now. Idk what you expect from OpenAI or why you think it would matter at all that Georgia is a democracy or an ally of the US or is closely aligned with us. Ask Canada and NATO how much that's counted for with this administration.