Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • OpenAI's blog post: Our framework for reporting model misalignment

    https://openai.com/index/model-misalignment-reporting-framew...

  • This is a very smart move when you realize you have a commodity product. Get regulated. Be one of the only providers. Protected status
  • > OpenAI said it did not believe the industry “has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer.”

    Baffling. To my knowledge, they didn't properly airgap their systems. Keeping the genie in the box seems like 101 to me, and to "miss" that seems awfully fishy. This, among all of the Anthropic news, is an odd convergence.

    Maybe they're being truthful and it really is the end times.

    Maybe they've hit a wall in improvements, but I don't know enough on the topic to speak to that.

    Which is more likely?

    Either way, trying to sift through this can of worms is tiresome. I'm hopeful that this all comes to a head soon, what an exhausting few years it's been...

  • So is this 0 accountability applicable to just AI companies? Or can regular hackers also claim "misalignment" as in they tried to just google something but accidentally their hands typed commands on Kali linux, found a 0 day and attacked and hacked companies?
  • If you find six roaches, you've got more than six . . .
  • > Other A.I. executives have said no slowdown is needed.

    So the largest companies, the companies with the biggest budgets and most users, are pushing for regulations that only they have the resources to follow.

    And this is based on new disclosures that include, ~"used a key without asking permission one time."

    What a clever way to lock up a market before open models get better.

  • > The San Francisco company revealed what it said was the “unexpected or concerning” behavior of its A.I. models as part of a new framework for reporting “misalignment,” which is when the goals or actions of A.I. systems diverge from human intentions and values.

    Misalignment: "when the goals or actions of [...] systems diverge from human intentions"

    How about we stop trying to nudge the language towards implying sentience or consciousness and keep the same word that has been used for that definition for longer than I have written software, a bug.

    We should be talking about why the tools/environment keep getting overlooked. The software built around the text generator, forget the researchers and mathematicians discovering the math properties of language patterns -- why are we not talking about the software engineers building the LLM-pluggable tools that actually allow/cause real action to happen?

  • Is there any precedent from other industries where a company tries to frame their own product’s shortcomings appear to be society’s problem?

    Would nytimes cover a self driving car company disclose concerning ‘behavior’ of their cars the same way?

    For anyone who has had to remind a coding agent to not leave comments over and over again, not following instructions seems more feature than bug

Explore Birbla archives