Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- My friends and I, and the teams I'm a part of, just want to build and create fun, cool things. I am so tired of being preached to by Anthropic like they're some arbiter of 'ethics.' So, so tired.by aquarious_
- You’re allowed to switch to the competition, including open modelsby dgellow
- > However, we are less confident in this assessment than we were in prior risk reports, since our most concrete task-based evaluations have “saturated”—i.e., no longer capture increases in models’ capabilities—and because we are seeing early signs of acceleration.
I totally understand this is a subset of alignment-related evals, but if Anthropic of all is running out of evals, doesn't that also means we are running out of things to scale?
I mean. I totally believe they have a model that is better at Kernel Optimization, creating new matrix multiplication algos, than Mythos. But it's clearly no generalizing, rightw
What am I missing?
by MP_1729 - > 5.3 Benefits from Anthropic’s operating as a frontier AI company
It does feel they are trying to ask the government to lock the market for us.
by hartator - a mystery “model 2” is mentioned alongside mythos/fable.by visiondude
- > More capable than Mythos 5 in some areas, less capable in others; overall slightly more capable.
This sounds like it might be a Mythos finetune for some specific task.
EDIT: After reading some more reading, it looks like model 2 might be an AI research fine tune based off the section 3.4.3 CoBench
by lwarfield - Meanwhile I can't really tell the difference between Fable and Opus for my tasks. I kinda think Fable does a better UX work so I keep using it for that because I couldn't be bothered to A/B them, but otherwise it's all the same and the model and effort are just feel good knobs I twist to still remain a load-bearing element. At least that's my honest take.by flyinglizard
- > Model 2, which is somewhat more capable than Mythos 5. Our rough qualitative sense is that this model is a noticeable improvement on Mythos 5 for many tasks relevant to internal use but does not display a capability jump of the degree observed from Claude Opus 4.6 to Mythos Preview. We do not currently have plans to release this model externally, and have not run all of our typical suite of predeployment assessments, so we have somewhat lower confidence in our beliefs about its capabilities.by merksittich
- Does anybody have any good reading on how the Chinese labs approach risk vs. the American ones?by int32_64
- by erwald
- If US model hacks US government, that's Very Bad. (China did this last year with Claude Code.)
If Chinese model hacks US government... free marketing?
by andai - It's crazy how Anthropic talks so much about their "AGI risk" and not enough about the risk of bankruptcy.by _ache_
- Are you surprised? Why would any private company spend time publicizing their financial risks?
Seems like a strange expectation.
- "all traffic through our systems for collecting human feedback data from contractors evaluating our models ran without blocking biological classifiers"
"totaled around 133M exchanges."
While this wound up being relatively benign, I still find this concerning, amidst numerous sandbox escapes, and previously, unreleased models being accessible via a custom URL. I don't think these companies are giving the responsibility they possess enough weight. How many more issues like this exist?
by internetter - AI companies are flooding the zone like Steve Bannon. Leave no one time to develop thoughts.by 12ahGA
- If they leave Steve Bannon with no time to develop thoughts, I'll tip one out for them.by esafak
- So as of a month ago their best internal model was "somewhat more capable" than Mythos "but does not display a capability jump of the degree observed from Claude Opus 4.6 to Mythos Preview." I thought they would have a significantly more capable model by then, more than five months after Mythos finished training. They'd better have one by now, or the Chinese competitors are closer to catching up than I thought.by modeless
- I'm still uncertain if mythos is real. Subsequent model releases have been lackluster, no one has claimed to verify mythos performance and it's silently vanished from most comparisons.by lumost
- You thought they were gonna double the model size again?
Also it occurs to me that they're somewhat incentivized to downplay cyber risks after what happened last time...
by andai - "We believe our internal AI R&D efforts are significantly faster than they would be without AI assistance, but not yet by a factor of 2 (though we are uncertain and measurement is difficult)"
So Anthropic thinks their productivity is not even doubled by AI. Interesting data point.
- >thinks
They can’t measure even measure it, it’s just vibes. They may not even be more productive.
by what - R&D productivity. Pretty sure they claim more for e.g. Claude Code.by OJFord
- It is also difficult to measure this in the face of ever shifting baselines. Anthropic is built on AI from the ground up. I'm not picking a side, but Anthropic getting a 2x increase is different bar than some random enterprise shipping legacy apps with clunky processes getting a 2x increase by adopting.by dangets
- I think LLMs are best for ideation, experiments, small things.
They've probably already settled on most of the architecture and the big ideas, so they're details in big things instead of how to make complete small things.
The thing LLMs really speed up is how some ordinary person-- a PhD student, or similar, can whip up a miniature synthetic experiment that turns out to be horrid and needs to be fixed by hand, but which at least gave him a plot on the same day he had the idea. That's, I think, where LLMs shine: prototypes. Anthropic probably doesn't need that to the same degree as the small experimenter.
- AI R&D efforts != productivity. I think it's fairly obvious that SOTA research is less affected by AI than writing another boilerplate react frontend.by T0Bi
- Well, it is a data point but AI R&D at a frontier lab is not really a representative stand-in for a regular workplace.by bonoboTP
- > So Anthropic thinks their productivity is not even doubled by AI.
I find it hard to imagine launching this criticism at a new technology.