Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • > Once a model is built, the biggest cost is inference

    Something i cant find any reliable data for, but would help for a sense of scale: How much use before its equal to training?

    I.e. assuming you have the training data and setup, and we only care for compute - How hours of using eg Kimi K3 / Fable, for it to equal the compute required to train it?

  • if you assume that training requires about 3x the compute of inference (one forward pass, one backward pass, parameter updates), and we take DeepSeek-V3 since their numbers are public.

    they used ~14.8 trillion tokens with about 2.66 million GPU hours. 14.8 * 3 = 44.4 t inference tokens.

    obviously, this is back of the envelope math, but at 100t/s you would need like ~14k years. scale this to >100k GPUs and your in the hours to a couple days range.

  • OAI feels far more cooked than anthropic.

    One of them heading to IPO and the other opting to not show their books should tell you everything

  • I have a problem with the cost per task metrics of Artificial Analysis. We don’t know how they calculate it exactly. But recently, cost per task has become the most discussed topic. The logic is basically: if model A achieves 55% on benchmark X and model B 60%, but the cost per task of A is 50% cheaper, people would choose A instead of B.

    But that implies that all output of the less intelligent model A is usable, perhaps only a bit worse than the output of B. But what if the output of A is unusable, or it can only deliver usable results in 1 out of 5 tries? In such cases, the user will have to rerun the task and it will very quickly double or triple the cost and makes the old average number misleading! I would argue the retry and flaky cost will be many times bigger than the average token cost and that is the true cost the users have to bear.

    AI-Benchy [0] (admittedly a one man benchmark) shows a much different figure than the numbers of Artificial Analysis. Opus 4.8 cost per task according to AA is $1.80 and Kimi K3 is $0.94$. According to AI Benchy, however, the *cost per successful task* of Opus 4.8 is 10.7 cents vs 19.4 cents of K3. The number of correct tests and pass rate of Opus 4.8 is also higher than Kimi K3.

    So on a cost-per-usable-result basis, Kimi K3 is actually pricier than Opus 4.8 — the opposite of what AA’s headline number suggests.

    Thus, I don’t know if I can believe the numbers of AA or we need to track the cost ourselves.

    [0] https://aibenchy.com/compare/anthropic-claude-opus-4-8-mediu...

  • Definitely feels like there’s a particular narrative being pushed on HN today.
    by a13n
  • Dario can reverse this by going on the podcast circuit again and threatening everyone with 75% job losses due to (his) AI this time

    Ramp the number up to 85% if that doesn’t work

    If it still doesn’t work, go nuclear and target 100% job losses language

  • Upper bound of AI progress - recursive self improvement. In this case AI will be responsible for building better models, making people who own datacenters the winners. Anthropic/OAI is cooked.

    Lower bound of AI progress - plateau. Progess is slowing, focus is on serving a meaningful peak capability at the lowest possible price. There's been news today that Google is building a Gemini chip with weights baked into silicon. Considering a chip's lifetime of 2-3 years at minimum, and that a 2-3 year model today would be useless today, they're expecting they wont make a similar amount of progress in the next 3. Game is about selling at the lowest margin. Anthropic/OAI is cooked.

    So their survival rests on the presumption that AI progress will fall between these two extremes.

  • > So their survival rests on the presumption that AI progress will fall between these two extremes.

    That feels like a very generous framing. There's very little opportunity between the extremes that would paint a convincing outlook of survival for either company.

    Perhaps if they were able to scale down their spending drastically they could survive, but that requires acknowledging their current valuations are BS. Doing so is a major risk, that will piss off all share holders. There's also the employees they would need to fire or reduce salary. The shift of focus internally to sustainability would be a major challenge.

  • Baking weights in makes a lot of sense for inference speed and power efficiency and has the added benefit of putting many end-users on the hardware refresh treadmill.
  • >. In this case AI will be responsible for building better models, making people who own datacenters the winners.

    We need SETI@home for Open Weight models yesterday...

  • I think a big question is whether any of these labs can produce a model that is _ahead_ of Anthropic and OpenAI.

    A related question is how much they're dependent on the APIs of Anthropic and OpenAI to achieve their results - whether through distillation or other uses.

    If these models are derivative of Anthropic/OpenAI I would expect performance to be more narrow and progress to be limited.

  • They dont need to be ahead on performance alone.

    Its value per unit of currency spent.

    Financials will ultimately drive decision making.

    We are already seeing that more intelligence does not correlate with more revenue, for the firm purchasing tokens.

    If I was OAI/Anthropic I'd be brown and yellow in the boxers.

  • > More importantly, as sustainable long-term businesses, model-only providers are particularly at risk. Knowledge Atlas, Moonshot Labs, and Anthropic face defensibility challenges versus OpenAI, Alibaba, SpaceX, Meta, and Google.

    Hm. How is OpenAI not a “model-only” provider just like Anthropic? Seems like they are vulnerable in the same way.

  • They've been doing hardware stuff as well. Both "consumer facing" bs like wearables / portables, but also more importantly chips for inference.

    Having a good model is one thing, being able to serve that model at good speeds and match demand is another. See Anthropic ~6months ago. Or Moonshot, they've already suspended subscriptions to their coding plans, because they can't meet demand.

  • They are trying to diversify into consumer hardware and also are further along in owning data centers. The consumer hardware product, assuming it launches and does decently well, puts them in a very unique position relative to pretty much every other frontier lab. It's more speculative, but I'd argue it can change things quite a bit for them if it works out.
    by cl42
  • ChatGPT is synonymous with non technical/work related LLMs. They're amassing a ton of user history. That history improves the product for the user because it has more context into the person. They can feed it back into model improvements and for advertising.

    You can see a future where a user types in "plan a vacation for me" and ChatGPT coordinates everything from there. Those sorts of users aren't going to switch because model X is 10% cheaper or better.

  • The article does address what they see as the difference ("its investments in product, consumer experience, site publishing, voice, and hardware are all directions that have clearer moats")

    That said I think it is pretty easy to make a case that these would-be differentiators are either currently underwhelming or completely unproven (as in the case of hardware).

  • It's amazing how quickly Fable went from 'Game-changing model that needs to be banned' to 'Yeah it's alright, but OpenAI is also just as good and there are a couple of good open weight alternatives that are equivalent for almost everything'

    The hype cycles are shortening, perhaps we really are reaching some kind of plateau this time (famous last words)

  • Almost like the company has no credibility with respect to its safety claims.
  • Fable is still the same model, it’s still a great model, and to be honest all these articles writing and speculating on how the LLM industry is going to evolve are not that insightful nor interesting.

    I don’t think one should pay much attention to them.

  • The plateau is inevitable because their rapacious training methodologies are only viable when there are no defense in place, but information continues to evolve, which means the models will have to be continuously updated, but will be doing so with less and less freely available data.
  • Open-weight models were lagging 4 months behind OpenAI/Anthropic at the beginning of the year. They are now just 4-6 weeks behind.
  • I think the risk is overstated.

    For one, on the margin people are willing to pay a lot for slightly better models. I know personally the value the LLM adds to my workflow is considerably more than the $200/m I pay the frontier labs. I have no interest in optimizing that to get it slightly lower. There are a very vocal minority that optimizes this or companies whose LLM expense is marginal, but I think that's the minority (correct me if I'm wrong, curious what their customer base looks like)

    Also the actual LLM is a tiny portion of the value added. Anyone that tried to build agentic solutions from LLM apis quickly realizes that a huge value is the Claude Code / Codex harness. There are open source implementations like OpenCode but they're not nearly as good.

    Think about it another way. Consider how much money Microsoft spends on maintaining Excel. There are open source alternatives that have >90% of the functionality, they'll even work w/ Excel files and generate them. Google sheets is probably 99% and available to everyone and better in a lot of regards. But the immense value spreadsheet software produces workers above the $100 or whatever a year makes it so that there is a real moat and no one bothers exploring alternatives.

    by bko
  • Not true. Enterprise pay for api. If i am allowed to use fable withouta budget, i can spend 3k per day. Using combination of cheaper models (and better at least in my experience) can lower this to 300. Considering the layoffs, thw companiesdo care a lot of this cost. Price matters more than quality for many use cases. For example, i can writea game where npc are all k3/5.6 level intelligence if the price is cheap enough.
  • > There are open source alternatives that have >90% of the functionality, they'll even work w/ Excel files and generate them

    Excel has network effects. If N people are using excel, the incremental N+1'th user is forced to use it or risk not being able to open documents from the N users. Excel is also sticky, in the sense that a user used to Excel UI is hesitant to switch to a new UI. Tokens dont have network effects. If N people use fable, the N+1'th user could use any other model, and have no impact.

  • I don't think subs are what keep their lights on. And their API is so absurdly expensive. For the harness hard disagree but ultimately its up to each ones taste. You may want to check this tho https://harnessrank.net

    On our side we use Claude/GPT/Kimi (it replaced Antigravity) for development. But we build our systems around a cheaper denominator (Deepseek previously, recently we added GPT 5.6 which have good prices as well). We offer BYOK for Claude but its def not an option to build something on top of it (for us).

  • > on the margin people are willing to pay a lot for slightly better models

    That margin is getting smaller and smaller. I would have been with you a week ago; paying for Fable was worth it compared to all other models. But with K3, the difference has shrunk to the point where, for me, it's not worth the cost anymore.

    In other words, it may be worth paying five times as much to get 10% better real-world outcomes for a lot of people, but a lot fewer people will pay five times as much to get 2% better outcomes.

    > a huge value is the Claude Code / Codex harness

    For me, it's the opposite. Having to use Claude Code instead of the harness I prefer is a point against Anthropic, not for it.

  • > a huge value is the […] Codex harness. There are open source implementations

    Like Codex https://github.com/openai/codex

  • > considerably more than the $200/m I pay the frontier labs.

    This is a very rich / developed country privilege perspective.

    Where I live, it's not unusual for a monthly wage to be around $200. Of course developer wages are much higher, maybe as much as $1000 a month, but $200 is still a huge chunk of that so it doesn't really matter how much "value" you get out of it if you're no longer able to pay rent or buy decent food.

    Even in developed countries, $200 a month is out of reach for all kinds of people who would benefit from it (students without rich families, entrepreneurs, etc.).

  • Here, people tend to forget about enterprise customers. Enterprise is excluded from using these heavily subsidized subscriptions, I know of orgs that are spending around $500k/month for teams of ~100 developers actively using AI.