Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • > Bringing the cybersecurity capabilities of Claude Mythos 5 to more defenders

    This dance is quite annoying. When Anthropic releases a paid model, the user should control what it does and doesn't do. The other day I had a security incident that required urgent response. ChatGPT and Claude were utterly useless (I quickly attempted to get access to the former's advanced security capabilities but was met with a form I could complete - far too slow for a time-sensitive problem like a security incident!). So I used grok and it helped! I am told Kimi also helps in such cases although haven't tried it yet.

  • Screw you Anthropic and screw your gatekeeping, Opus 5 refuses even basic reverse engineering / patching tasks, even BPF is apparently too dangerous, you lost your minds..
    by zb3
  • For those who had access, how does it compare IRL with GLM 5.3 ? iirc both models are similar in terms of benchmarks ?
  • $35M in credits (!) for the Defender Advantage Fund (0xDAF) doesn't sound that much given that e.g. the HAWK attack [1] cost $100k for 1 (albeit very advanced) vulnerability.

    As a side note, that Golden Eagle wording "bringing a wartime footing to the cyber domain to relentlessly patch vulnerabilities" sounds so AI-written maybe that's what the credits are needed for.

    (1) https://www.anthropic.com/research/discovering-cryptographic...

  • Seems like a whole lot of nothing for the average user. They have really lost the plot.
  • What a bunch of wankery from Anthropic. I already use Sol 5.6 for security auditing and it works great, and as a bonus doesn't give verbal vomit every time.
  • This week I found two issues in my company codebase. After finding them I told a claude session about one and asked for a quick proof of concept demo of the exploit. It refused, including refusing simple things in the same session afterwards.

    Meanwhile same model in a new tab, say I need help creating a page that hits an endpoint with a special payload and it does the same things that were too dangerous in the previous tab...

  • The problem is that "cybersecurity" isn't some special task that only your security team does.

    In the project I maintain, I find bugs and fix bugs. Some of those bugs might result in an LPE. I generate a regression test, then I fix the bug.

    The problem is that generating a regression test for that type of bug is technically a PoC. I can almost never get Fable to create one. Sometimes Opus 5 punts as well. Same with Sol and Luna.

    That is, unless I socially engineer the model. I can't talk about security. I make sure they don't read the file call cve_test.c (literal regression tests for CVEs). I have to hide part of my project from the models for them to work.

    Anthropic and OpenAI are driving me to use other models.

Explore Birbla archives