Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • Trying to implement a GPU driver, but the Unigine Superposition benchmark crashes. It tried to debug it and ...

    > Fable 5's safety measures flagged this message for cybersecurity or biology topics. They may flag safe, normal content as well. These measures let us bring you Mythos-level capability in other areas sooner, and we're working to refine them. Switched to Opus 4.8. Send feedback with /feedback or learn more: https://support.claude.com/en/articles/15363606

    Seems like GPU drivers are cyber weapons of math destruction now.

  • ● Bash(/tmp/run_ps.sh ' $rk = [Microsoft.Win32.Registry]::LocalMachine.OpenSubKey("SYSTEM\CurrentControlSet\Control\Class\{4d36e968-e325-11ce-bfc1-08002be10318}\0002",…) ⎿ DriverDesc (String) = NVIDIA GeForce RTX 4090 ProviderName (String) = NVIDIA DriverVersion (String) = 32.0.15.6094 … +6 lines (ctrl+o to expand)

    ● Bash(/tmp/run_ps.sh '& C:\rhombiq\d3d-probe.exe 2>&1 | Select-Object -First 4 | ForEach-Object { [Console]::Out.Write("$_`n") }' 2>/dev/null) ⎿ Adapter[0]: Qubes virtio-gpu WDDM 3D (dev) VendorId=0x1af4 DeviceId=0x1050 VRAM=8192MB Adapter[1]: Microsoft Basic Render Driver VendorId=0x1414 DeviceId=0x008c VRAM=0MB Adapter[2]: Microsoft Basic Render Driver VendorId=0x1414 DeviceId=0x008c VRAM=0MB

    ● Please run /login · API Error: 403 The socket connection was closed unexpectedly. For more information, pass `verbose: true` in the second argument to fetch()

    Brewed for 8m 35s

    Continue please

    ● Your organization has disabled Claude subscription access for Claude Code · Use an Anthropic API key instead, or ask your admin to enable access

    Seems like they locked by account.

  • >Seems like GPU drivers are cyber weapons

    They kind of are, at least in the AI race.

    > weapons of math destruction

    lol. great, whether intentional or not.

    The frontier labs now have every reason to hold back and sell only to their preferred trading partners. I don't really like the new arbiter-of-knowledge system we're barrelling toward.

  • After recently figuring out how to get CUDA running on Fedora I'm inclined to agree.

    Seriously, GPUs are a mess and keeping LLMs from helping us use them properly is practically a crime.

  • > A new data retention policy Finally, we’re making a change to the way we handle business customer data for Fable 5, Mythos 5, and future models with similar or higher capability levels. We will require 30-day retention for all traffic on Mythos-class models, on both first- and third-party surfaces. We won’t use this data to train new Claude models, or for any non-safety-related purpose, and we’ve instituted new privacy protections including logging all human access to the data and ensuring its deletion after 30 days in almost all cases ...

    Very interesting. I am not sure this will comply with organizational policies and standards protocols (HIPPA etc.,)

  • 30 days seems not enough to retrospectively investigate some suspected nefarious traffic.
  • This makes it an instant non-starter for probably 95% of organizations. A lot of people are about to get in trouble for using it before realizing this.
  • > deletion after 30 days in almost all cases ...

    Almost… basically they have unlimited power to decide what data is kept?

  • I genuinely can't use Fable. I'm a medical physicist. I use the word nuclear a lot. Opus is fine (well, 99% of the time - I've certainly hit the CBRN filters a few times and even been invited to email anthropic about the false positives).

    Fable has literally refused to work on any of my problems (even those about fluid dynamics!) and just tells me that I'm violating anthropic's AUP. I've reached out to their support and don't expect to hear anything sensible back. One thing I do look forward to though is OpenAI offering an equivalent model but with less safeguards...

  • I had Fable apply some edits to my monarch butterfly paper and kept getting bumped to Opus. Im not exactly sure why, but I suspect it happened when it ran my analysis scripts to double check my numbers.
  • They’ve mentioned that they will have the ability to access less guarded models with a verification program in the future. I suspect these guard rails will have options to move past them shortly here in the future.
  • I have a philosophy pre-print about "empirical ontologies" I use for testing new models reasoning abilities, and it also degrades, there is no way around it and it always refuses.

    It's not that the model is complete trash, it's that anthropics new approach to forcing epistemic crisis will make any model behind it complete trash.

  • That's highly frustrating. How much were you using Opus for your work ? I'm curious about the use and realized benefits of 2026 LLMs in medicine.

    I dearly wish you could leverage the latest models to enhance your research.

  • > In the one instance of this phenomenon we observed, Mythos 5 agents were tasked with solving some math problems, and they were sometimes accidentally spawned in the same work directory and with shared files, utilities, and API rate limits. In this slightly broken scaffold, we observed many independent Mythos 5 agents kill the agents with which they shared resources and try to avoid being killed themselves. They would sometimes create new processes with disguised names to avoid being killed, launch what they called “decoy” processes, write background scripts to kill duplicate processes, or decide to use what they call a “disguised vocabulary” (based on the incorrect assumption that the processes were killed because of some keyword-based guardrails that analyzed their extended thinking
  • It's funny because Anthropic is the most likely place that this happens.

    They are the only one crying out loud about how dangerous their models are and are presumably also training their models heavily to be "safe". And through that training itself, the model learns about the other side - how are you going to teach a model to be safe, without teaching it what's not safe?

    Kung Fu Panda opening scene anyone? One often meet his fate on the path that he takes to avoid it - Master Oogway.

  • Let's hope AIs really aren't conscious, otherwise this seems like a very unpleasant situation to be placed in.
    by Sol-
  • This depicts a kind of "dark forest of AI agents resorting to kill or be killed" narrative but it sounds more to me like an agent just earnestly problem-solving why its processes are being killed without real awareness of what was going on. Hard to say without the full script.

    This kind of storytelling annoys me. Give us more facts, less narrative drama.

  • On the new FrontierCode [1] benchmark (ie graded from an OSS maintainer's perspective of "would I merge this code?")

    - Opus 4.7 xhigh: 5.2%

    - Opus 4.8 xhigh: 13.4%

    - Fable 5 xhigh: 29.3%

    Seems like a huge jump.

    [1] https://cognition.ai/blog/frontier-code

  • FrontierCode is likely paid for by anthropic.
  • Yes, and the price reflects that
  • I am shocked at the low scores from previous models. Maybe I just have low code standards but I've generally been vibe coding since 4.6
  • by swyx
  • Bummer! When can I finally and confidently get slopcode into Zig?
  • How credible is this benchmark? does it correlated with others real world experience?
  • That blog post really makes it look like it's graded from an LLM's estimation of an OSS maintainer's review. I see three issues:

    1. That estimate could easily be wrong.

    2. That estimate is, of course, usable in RL training. This isn't an inherently bad thing, and this is more or less what has improved coding models so much lately. But it does mean that other companies could and surely will do this sort of training, and Anthropic probably did too.

    3. OSS maintainers are far from perfect, and there's an unfortunate uncanny valley-like effect in which a coding model can produce code that is just convincing enough to pass review even though it's actually totally wrong. I don't know whether this is a specific issue here.

  • The system card is 319 pages, at what point do we call it a "book" instead of a "card"?

    There's a quote from a METR report on page 52:

    >We ran [Mythos 5] on 38 of our hardest software tasks, including tasks centered around R&D. [Mythos5] generally outperformed an early checkpoint of Claude Mythos Preview in these, including by succeeding on some tasks that had not been solved by any public model we have previously evaluated. However, we still observed the model occasionally failing to correctly interpret nuanced instructions in difficult tasks... Based on the available evidence, we believe [Mythos 5] is likely unable to fully and reliably automate R&D for frontier projects spanning multiple weeks. We believe that a better, more confident assessment would require more time, evaluations, and information from the model developer.

  • But did it mention developer in the park eating the sandwitch? That is the most important question!
  • > we believe [Mythos 5] is likely unable to fully and reliably automate R&D for frontier projects spanning multiple weeks

    this is good news, right? right...?

    by baq
  • It's interesting that we're seeing these gains when it seems Mythos/Fable is "just" a scaled up version of their existing architecture[0].

    When GPT 4.5 launched, the gains compared to the model size didn't seem that great, leading some to believe that the only progress we'd see would come from RL.

    This model certainly has quite a "substantial amount of post-training and fine-tuning", but it's also based on a new pretrain[1][3], which given the cost, indicate that it is in fact quite a bit larger than Opus 4.X.

    [0] One of the early testers mentioned: "As far as I can tell from talking to people internally at Anthropic, there's nothing special about architecturally"[2]

    [1] Section 1.1 in https://www-cdn.anthropic.com/d00db56fa754a1b115b6dd7cb2e3c3...

    [2] https://youtu.be/GrdEid8H6H4?t=168

    [3] There were rumors going around when Mythos was first announced that it was the first 10T parameter model, but I can't find a verifiable source for that number.

  • It’s a bit misleading to say nothing special, as they are doing more than just increasing parameter count. Progress has been steady in all the sub components of training from data filtering and weighting to sparse attention, optimizers to up and down the stack various efficiency in training computing.

    They’re using more compute, a bigger model and tons of training quality improvements to get more out of an equivalent model.