Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • Are people still using deepseek-v4-flash everywhere? I found after the price increases, mimo-v2.5 seems far more attractive.
  • It's been cheap again on openrouter for the past few days. No idea how long it will last, but I've been using it from Baidu over the weekend, and it was about half the cost of the old DS prices, before the increase. Looks like people are figuring out how to offer it for peanuts.
  • I have really good experience with GLM-5.3 The subscription limits are generous, code quality is comparable to old (good) version of Opus 4.8 Some people report issues with it’s being slow, but I didn’t feel it. I use OMP harness (Pi derivative) and Matt Pocock skills.
  • glm 5.3 gets awfully slow during peak hours. But you might not hit them to frequently.
  • Is the real LLM revolution the fact that every piece of news and opinion is now phrased as if it's the end of the world though?
  • Apocalyptic doomsaying as marketing strategy. Not great for the Zeitgeist honestly
  • Seeing a lot of people in here say that they need Fable for the tasks they're doing and Opus just isn't enough. My experience could not be more different. I seriously feel like Opus-level performance is totally adequate for most of my use cases, if not all of them. And it's probably been this way since, like, realistically, Opus 4.6. On the other hand, Fable I've observed getting into verification loops that just burned so much of my token budget. Combined with the higher cost of tokens from Fable to begin with, I just pretty much never use it for anything.
  • I agree. I've been perfectly happy since Opus 4.6. I remember thinking at the time if it never improved I would have been fine there.

    The vast majority of people, eg vibecoders, do not need Fable or Sol tier intelligence for their slop To Do app.

    by dbbk
  • As models train up the intelligence ladder, many common tasks will hit fully diminished returns, and instead it'll just get progressively cheaper to do that task. But the tasks that AI is capable of doing are also expanding. I'm not sure 'Some tasks don't require the peak of the frontier' is worth worrying about, from an AI finance perspective.
  • It’s worth considering for companies paying API prices, and not relying on a subscription quota
  • Exactly; so far, we've only replaced the need to design algorithms and hand-write code; what if we apply the same effort towards the skill needed for system architecture, project management, and the rest of the SDLC? Or even outside of software!

    Right now, it feels like all of that is today where coding was a year or two ago, and we're on the cusp of some massive improvements outside of coding. It'll be interesting to see what these companies decide to automate next.

  • Yes. I use Opus for tasks that Sonnet could probably handle, but I'm not hitting my quota. Whatever minor incremental gain is "worth it", since marginal cost is zero.

    Even now, I use Fable as the planner and coordinator, with it farming out to agents. I don't hit my Fable limits either.

    Which means I could accomplish more, but these are side projects so I don't need 30x productivity. Still, claude is constantly churning away at something.

    by tyre
  • >> GLM 5.2 is worth focusing on. It came out the same week as Fable and is roughly 1/9th the cost (and ~1/5th the cost of Opus 5). Is GLM 1/9th the quality of Fable? Perhaps, for certain classes of tasks. But for most rote coding it’s more than sufficient. Especially when provided with great context. I frequently chat with Fable to interrogate and shape a design, before handing off a brief to GLM.

    People say stuff like this a lot, but I have a different take.

    The whole "such-and-such model is 90% as good as Fable at 1/10th the price" assumes that the value increase of intelligence is linear. But I think it's exponential: that last 10% makes a massive amount of difference. It can result in a key insight that helps you strategize more effectively, a novel approach that saves a huge amount of time, a feature design that is lot more user-friendly (because top models like Fable also possess substantial non-software domain knowledge that help bridge the gap between user and software), or the depth and breadth of engineering expertise that helps avoid a nasty bug that would otherwise have cost you users and revenue.

    Yes, it is totally possible to use Fable as the planner and delegate implementation to lesser models. I do that. But, my theory (which I unfortunately do not have the money to test and prove) is that a codebase designed and implemented by Fable would be substantially better than one that is designed by Fable and implemented by Opus 5, GPT 5.6 Sol, GLM, Qwen, Deepseek, etc. The reason I believe this is because I read the code Fable writes and compare it to code that any other model writes and the difference is night and day. It's not just 10% better. It's mid-level engineer vs. principal/staff-level engineer. And the thing is, even for rote tasks, a more senior engineer is going to be more likely to come up with a clean design than a mid-level engineer. They will also be much more likely to take a step back and ask important questions or propose different approaches.

    So if you're using Fable and everyone else is using lesser models, sure they might be saving a lot of money, but there's a higher likelihood that your product will be higher quality, perhaps to a significant extent. And models that are released in the future will benefit from it as well.

  • > It can result in a key insight that helps you strategize more effectively, a novel approach that saves a huge amount of time, a feature design that is lot more user-friendly

    My brother, that's my job.

  • > my theory (which I unfortunately do not have the money to test and prove) is that a codebase designed and implemented by Fable would be substantially better than one that is designed by Fable and implemented by [others]

    I don't have proof, only my anecdotal experience: I leave plenty of Fable usage on the table because I do not think its implementations of code have been better to Opus 4.8, not even close. It overengineered, obscured and picked awkward constructs all the time over plain, simple, perfectly clean and performant code patterns. Code was smarter AND worse in the kind of way that a brilliant and overeager recent grad often does. (I know I did)

    by Jare
  • Something I’ve found comparing between Fable and Opus is that Fable has impressively good analysis skills, but both of them seem to go way way overboard with “present state” comments “# We’re making this change here because of this issue blah blah, here’s what you need to know about np.percentile, blah blah” that I end up significantly pruning before making a PR. I let it do the same style verbose commit messages (because a contextual history is cool there). I haven’t actually noticed a ton of difference in the code that they write personally, but have found that Fable does find nuances during data analysis that Opus misses.

    In that light, I often go the other way: let Opus (and Haiku subagents) do most of the heavy lifting and then give Fable a shot at finding holes, especially if there are holes or unanswered questions or unearned assertions that I’ve caught on my own in Opus’ output. This, so far, seems like a clean tradeoff that doesn’t burn my Fable credits as hard and still gives solid results.

  • Why are we not just in a free lunch moment but with harnesses, rather than models?

    Right now a lot of people have a lot of opinions on which model to use for which task. They get better results for less money by judiciously switching between Fable and Opus and whatever else. Spending my time learning this skill would have an immediate benefit for me.

    But on the other hand, maybe the harness vendors will just solve it in 6 months? I'll ask a question, something in Claude Code (or whatever we're using by then) will figure out the most effective model based on the question and the context and my apparent willingness to get it right. I'll get billed X or 10x as appropriate, and I'll be happy with that, because that's what I would have paid if I made my own choice of model every time.

    Claude Code already does this a bit, sometimes it will tell me it picked Sonnet for such and such a sub agent, or some other detail I'd rather not care about. The best humans seem to be better at deciding what model to use than any of the tools is, but surely that won't last long.

  • Hard to inagine the final outcome being anything other than the smartest model + cheap subagents. The only problem with that is user requests being pasted directly into the model's context, but that's got to be temporary.
  • Reading this as someone who switched over to ChatGPT after (and largely because of the changes made in) the Fable release, it reads a bit naive. Not only do I find Sol to be as good, if not better than, Fable it is also faster, better behaved and has a much more coherent writing style. You also don't randomly get the Opus downgrade. OpenAI seems to be pulling this off due to their partnership with Cerebras so I wouldn't make any comparisons to Moore's law just yet considering it seems like we're just getting started in that department. Anthropic could (and should) do the same thing. It certainly feels like model development is at a point where it would be worthwhile building special purpose silicon for the models we have now since they are capable enough that they would still be useful even when/if further advancements are made. If anything, I think Anthropic's problem has more to do with their micromanagement of what users can do with their models, they're creating an undue amount of overhead for themselves by over-policing usage and capabilities.
  • Anthropic -> OpenAI switcher here.

    I fully expect I’ll switch back to Anthropic, or another model in the next 90 days. The fact that we are switching indicates that the models aren’t ready to be baked into silicon.

    I wonder if they will ever been that good, or if the lifespan of silicon is longer than the lifespan of a model before it needs to be retrained.

  • Etched is doing this. it seems like in the near future they'll actually ramp up production. not sure how much faster/economical compared to Cerebras but..
  • The Cerebras version of 5.6 is available only to select customers
  • I find it somewhat funny that the author starts by talking about how Moore's law enabled inefficient software and that we then had to make it more efficient, when almost all software today is horrendously inefficient compared to even 10 years ago, let alone 20-30. We've somehow even achieved a state where it doesn't matter how fast your CPU and memory are, the software will just perform horribly on any machine.
  • Its slow because of human slop from the early to mi 2010s, and now slow because of AI was trained on the slop that existed.

    Bottom line is the slop used to be manageable, but now there is 100x more code pushed, so that train has departed.

    In the end its more bad code for features no one will use.

  • One of the funniest parts of the LLM wave is discovering that cron was so annoying to use that we will burn the planet to put an interface on it that people can actually work with.

    The tendency for absolute inefficiency is effectively unbounded until scarcity is imposed.

  • Most of the things I work on are at least security adjacent. At some point chatting with Fable inevitably leads to it thinking about the security related aspects, tripping the safeguards.

    Maybe Fable can do the same things better than other models, but having to tiptoe around to avoid tripping safeguards makes GPT 5.6 so much easier to work with that I don’t even bother with Fable (or Opus 5) now.

  • That's completely valid. But worth noting that most of the stuff I work on is not security adjacent (mostly UI / layout / rendering related), and I almost never run into this.
  • I don't even need to tiptoe ! Not being there and not prompting anything is enough to trigger safeguards.

    Having not asked a single security question it will write wildly vulnerable code, go back and fix it, and guardrail itself out of existence after charging me a large sum with no refunds for no output and having not fixed it because that might be secuirty adjacents.

    And if it doesn't do this you end up with code that has such holes, store xss , no authz ... if it does not go back and notice it has written bad code.

    Since they hide thinking and reasoning from the user (who is also paying for those tokens) it is a black box what is triggering it, has the LLM this time thought of "Oh, this has XSS" and used a bad dangerous word such as XSS, while the previous conversation did not ?

  • > At some point chatting with Fable inevitably leads to it thinking about the security related aspects, tripping the safeguards.

    It happens to me all the time with things that have nothing to do with security, Fable spawns a subagent that then adversarially checks the code Fable just wrote and hits guardrails, with zero prompting from me.

  • There are whole classes of things I can’t thought exercise or really learn about because the “safeguards” keep tripping me down to haiku.

    Like middle school level genetics stuff from a guy who hasn’t been in school for decades.

    They need to fix that. It’s just broken. Nobody is making bioweapons if they’re asking the dumb sort of questions I’m asking.

    Also, it refused to identify an actor in a popular tv show from a photo. Apparently the policy is it won’t identify ANYONE from a photo, now. Even publicly listed cast members from a very popular show, from a photo of a scene in that show.

    It claims that’s a fixed security policy. Nevermind how that makes absolutely no sense… argue about it enough and it terminates the chat.

    I don’t know what the Anthropic clown car is even doing anymore, but I won’t be surprised when the others eat their lunch.