Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • I often wonder if the people at anthropic actually believe in the existential risk stuff they talk about. If I was convinced AI could kill everyone, my actions would be almost diametrically opposite to theirs -- I would be committed to ceasing the operation of AI labs at any cost.

    It's easy to say they're just lying and it's marketing hype, but I also think it's possible they really believe the best way to stop an evil superintelligent AI is to work on making superintelligent AIs, but just do it smarter than people at OpenAI would. Humans are bizarre creatures!

  • > Once models can perform work that reduces AI risk at the level of human experts, AI(-assisted) output in the area might dwarf unassisted human output.

    Straightforwardly true.

    But doesn't this smell like asking an organization to design its own oversight and guardrails?

    Even before you get to alignment issues, LLMs are really good at generating content that "sounds right" to humans. That's basically what they've been hyper optimized for.

    At least with math (and to some degree software) we can verify the result. But with the intersection of science fiction, philosophy and ethics... not so much!

  • > For example, if we ask a model for the probability P(A) and another instance of the same model for the probability P(A&B), do the reported probabilities satisfy P(A) ≥ P(A&B)?

    So that's it? Are you saying they compressed risk reduction to high school level stats calculations and a greater than or equal to? To early for this

  • This is a modest start on an important direction for AI Alignment work; which is, as the authors observe, commonly comprised of tasks which are not readily empirically verifiable and not easily mathematically modeled - so it's hard to get at with normal RL techniques.

    I find the ACCoRD benchmark the most interesting, because you could theoretically scale it up from the baseline mode of testing two instances of the same model for their `P(A) ≥ P(A&B)` respectively, you could do `P(A)≥ P(A&B) && P(A) ≥ P(A&C) && P(A&B) ≥ P(A&B&C) && P(A&C) ≥ P(A&B&C) ...` etc

    i.e. a swarm of model instances could be collectively measured for consistency for even more confidence, right?

    At any rate, even the basic idea of measuring a model for consistency in beliefs improves our ability to bound the amount of trust we can put on it with introspection methods.

  • We invented a new benchmark and look we're at the top. Everyone else sucks compared to us. Especially those dirty open models.
  • We have totally come out with this idea of this index that will allow us to create policies to ensure only the best and safest AIs are used by the public....

    Absolutely no conflict of interest, no lobbying here

    And no this is not related to those chinesse models... it's not the same as HD vs honda thing...

    Trust us, this is the same kind of amazing thing as boeing doing their own certifications and inspections!

    -- First reply: those dumb models, who use them, they are good only for adding 1 + 1 -- another: They will do the same so who cares... -- The valve guy is a dick, and has a monopoly...

  • Opening line:

    > A core hope for managing AI risks is that AIs will help us understand our situation

    Gonna stop you right there and ask that you think deeply about that premise.

  • Ah, yes -- A closed source benchmark that Anthropic paid for that Anthropic ranked highest.

    0/10

Explore Birbla archives