Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • Refusal: "Opus 5.5's safeguards flagged this session. You may be seeing this for the first time on an Opus model: Opus 5.5 is more capable and has stronger safeguards as a result, which can sometimes flag non-cybersecurity work. We're improving these safeguards to reduce the amount of incorrectly flagged messages."

    My appeal: "This is my own code, my TST and PRD environments and I am concerned about the hardening the hand-rolled BasicAuthHttpModule function.

    I asked DeepSeek V41 Flash to review this code already and worked in his recommendations.

    Now I want to ask you for a second round of review, a second opinion audit.

    This is an ASP.NET 4.8 application facing internet and I want to make sure I handle the edge case, HTTP error codes, and have no logic gaps in my code."

  • Anthropic is the new Microsoft. Just my gut. I'll be staying away from their products. Hopefully it will benefit my career the same way by focusing on open standards, instead of some proprietary bullshit that changes every 3 months.
  • First thing it did when I tried it, was roaming through files in directories way outside of the project. I tried to get it to explain why it did it multiple times, but I never got anything resembling an explanation.
  • > Fourth, if long tool-calling turns still go quiet for longer than you want, have your harness ask for an update

    I'm not sure I understand this complexity. In all harnesses I've ever used, tool calls themselves are surfaced to the user as an indication of progress. When the UI/UX around this is engineered well, the user should be able to infer roughly what is going on. Different tools have different ideal presentations. You can't reduce everything to plaintext blobs.

    If I absolutely needed intra-turn progress updates, I'd accumulate a separate per-turn transcript and feed it into a cheaper model at deterministic intervals.

  • "the biology safeguards are the same as Claude Fable 5.1's ... Everyday health and educational questions are unaffected"

    Yet here we are, "why my calves hurt more than any other muscle after training" being classified as a naughty question.

  • Opus 5.5 is a good model, but I've tried to understand the extreme hype about it on social media about Opus' ability to do 2d work, as we got with Astra doing 3d work. In both releases, the models required extensive access to third party apis to generate assets for it, and a lot of the models work was essentially coordinating everything.

    There's so many "x generated this in one shot, this is agi" stuff that gives you the impression that you can vibe operate modern models the same way you operated last year's models. There's so much more to it than that. It requires you to put a faith in the leap in the capability of models, one that would've surely been a waste of time in previous models.

    Not sure where i'm going with this other than I think most can relate that it's exhausting keeping up with. I cant imagine what it'd be like parenting a kid that went from toddler to puberty in the span of a year and planning for them to go to college the next year. This industry is moving so fast that it's becoming fact that it's the user that's "holding it wrong" every six months.

  • This is a failure of the AI foundries; if we have to use totally different prompting techniques for every model, this wont work.

    AI is rapidly saturating it's ability to be useful and these products need to start to mature.

    It's not 'fun' to manage 50 different broken MCPs and their variety of ways in which they are broken.

    It was 'fun' at the start, now it's just 'broken technology'.

    Astra and Opus 5.5 are the 'starting point' for the next era of AI where we expect robust tooling.

  • All that keeps jumping out at me is how they've set it to refuse giving users thinking tokens and prompts for full reasoning in output. Just drives me further away; I may not stop using Claude completely for now, but I'll be moving even more of my primary workload to Chinese providers. That's where openness and freedom is now at.

Explore Birbla archives