Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • I've posted this anecdote before, but i feel its worth posting again because I've seen this spread further.

    >amusing side note: >Was in a meeting reviewing a potential new product, it was going well until they showed us that they had added AI to it (of course they have). It was pretty obviously just shoehorned in, and one part of that obviousness was that they had a column that showed how many tokens it took to make each query.

    >I asked who is paying for the tokens, they said its included in the license. I said, so is there a budget or is it all you can eat. they said good question they didnt know and would get back to me. I said the reason i asked was just one query there had a 250k token burn on it. and it was a fairly simple query about one device.

    >then, one of the execs on their side was heard saying out loud "Why are we even showing this to the customers?"

    >it have us quite a chuckle. But lesson learned... the cost of adding AI to anything isnt really being accounted for let alone the true cost of actually running the AI.

    >all things AI are going to get more expensive. even if you dont want the AI aspect.

  • > We are seeing improvements with each model release these days but it’s clear that the improvements are getting smaller and smaller.

    This is obviously untrue, both with GPT-5.4, and Claude Fable as examples in the last 6 months.

  • gpt 5.5 regularly wastes tokens on wrong commands, requires lots of handholding. I highly doubt there's substantial improvement
  • > the improvements are getting smaller and smaller

    The AI haters have been saying this for 2 years now.

  • I would struggle to ascertain the day-to-day difference between GPT-5.4 and GPT-5.5 tbh. Also, imho, Fable is highly hyped, I don't think it is dramatically better than Opus 4.8. Maybe my tasks and interaction with AI is relatively simple (i.e., lots of Rust programming, Linux system engineering stuff).
  • Prices will go down one way or another. That is of course unless the market gets cornered by restricting model use, restricting supply of essential hardware components or raw materials to make this hardware, etc.

    In terms of running the model locally vs a service provider, that will be down to convenience more than anything else for the same reason why not everyone is hosting their own website at home on their own box.

  • Token prices will go down for sure, but i watched a video interview on yt from cloudflare ceo and apparently the internet traffic of agentics increased and took over human.

    If we continue this year with a2a, agentic layer and co, there is probably a huge bulk coming up with a lot more agents running a lot longer and talking to each other to solve issues which will increase token usage significanlty.

  • It's weird to see people claiming that model capabilities are plateauing. It wasn't until late last year that we even had strong coding models. Imagine if, less than a year after the first iPhone launched, people claimed that smartphone capabilities were "plateauing" because Apple hadn't yet launched a new phone. And it seems the issue is less than "models aren't getting better" than, "models are good enough to handle 99% of the coding tasks people give to them".
  • People claim what they see. I see no improvement since opus 4.6, quite the opossite.
  • I wonder how's model improvement/dollar invested ratio going. If gains are made by simply spending increasingly vast amounts of cash that's not going to last.
  • I am convinced that the combination of capable open weight models and specialized hardware will mean that Apple (and other hardware providers) will start shipping computers with built-in, hardwired, "LLM-on-a-chip" cards that are capable enough to meet 90% of your AI needs.

    I really believe that in the near-term future we will run our LLMs in hardware, not in software. Hardwire a capable model into a device the size of a graphics card, embed it into a laptop, and you have something that uses less power, does faster inference, doesn't require additional CPU or memory, doesn't cost a monthly fee, and will probably eventually be available for under a (few) hundred bucks.

    by akie
  • I’d put money on Apple buying/acqui-hiring the Talaas people to further this.
  • The article's own example makes the point I'm about to make:

    > To give an example, just doing Typescript type fixes with this model across 50 files cost me $54 this afternoon.

    That's all because it ran through the most expensive frontier model for a mechanical task that a cheaper model could easily handle.

    What hardly gets mentioned is that most people don't actually measure what each task costs them on a granular level. A lot of the waste comes from running everything through one expensive model without considering breaking the big task into smaller tasks farmed out to cheaper models.

    Whether the labs' economics hold is above my paygrade. It costs what it costs. What I can control is my own usage.

    Like everyone else leaning in heavily on AI usage, my tokens started running out mid week...sometimes within a couple of days. I had to do something about it or double my spend. So I started tracking my cost per task type a few months ago and it completely changed my workflow. The lowest hanging fruit was the mechanical stuff. Moving that to the cheapest models was a game changer, and much faster to boot. Mid tier models take the workhorse tasks. The frontier heavy hitters are now only used for judgment calls like reviewing and planning. Spend dropped dramatically.

    Freeing up all those tokens made me even more ambitious to explore parallel ways of working, to get even more out of what I was already paying for.

  • Would prefer not to offend the author, but I do believe this article has very little for the HN audience. No new insight, and no numbers or new information.
  • Is there any place with better curation? I notice quite a few articles summarizing the state of AI that feel redundant with one another
  • There is a wave of users switching over to DeepSeek Flash. There are Reddit threads of users sharing billion token spend for $20.

    If all of global spend on Anthropic/OpenAI/Gemini APIs just switches over to DeepSeek then easily we can decrease total AI spend by 10x

  • I am not sure if that is wise. It’s a hostile superpower after all
  • Probably won't be too long before the government decides to block deepseek's website based on "security" concerns.
  • DS is restricting the "expert" model usage already, because they do not have enough compute.
  • v4 flash was really, really good in practice for me. While on openrouter it's around 1/100th of what the "SOTA" models cost.

    But billions? A bit exaggerated.

  • I have already seen a number of people doing the math on what it would take for hardware to self host a Q8XL quantization of GLM5.2 shared between N numbers of people.

    There's additional advantages that everything you query, all of your context cache and everything it outputs stays private and can't be arbitrarily turned off by external interference.

    Personally I think it would be a fairly good bet that something with the 1TB of RAM needed to properly self-host GLM5.2 will still be a very usable piece of hardware in 4 to 5 years from now. There will be even larger, newer models available, sure. But there will also be better models that continue to fit in the same size.

  • the same argument was made 2 years ago: "in 2 years we'll be able to run GPT-4 level models on an expensive laptop, most people will be using this instead of the fancy cloud models".

    we are there, Gemma4/Qwen3.6 are GPT-4 level models runnable on a fancy laptop.

    but expectations shifted, nobody wants a GPT-4 level model anymore

  • Back in the earlier days of the internet, when "dedicated servers" were a competitive advantage, hobbyists and small dev shops definitely shared dedicated hardware.

    So you could see small LLM co-operatives working out, yeah.

    But my thinking is that this four-to-five-year scenario just won't come to fruition, because the whole concept of needing to run these massive, massive models will slightly more likely be rendered moot by smaller models with better reasoning capacity, and possibly even in that timescale by hardware innovations.

    One of the biggest problems I have with the whole "we won't be profitable until 2030" model is that 2030 is almost exactly as far into the future as the launch of ChatGPT is in the past, and in that time, models far more capable than that first ChatGPT have been made available to freely download and run on desktop hardware that existed before it launched, and the entire non-model surrounding functionality of that original ChatGPT plus many more functions is now not much more than a routine weekend coding project.

    I don't know why the market would entertain the idea that no upset like that is possible in the same period of time again.

    by dofm
  • I am using perhaps 15% of usage count on Claude with just the normal subscription. And I do full time software engineering and would say I use quite a lot of AI input on thoughts, designs and code drafts.

    So how these companies and people manage to use these absurd amount of tokens is a mystery to me. It feels like this are just running huge amount of non-vetted data to the LLM's and or running loops against the LLM's which only produce fractional results if not wasted results for insane cost.

    So really it is the equivalent of just burning money, or heating your house in the winter while having all your windows open.

  • I mean have you tried to tokenmax?

    It is not that hard. Just launch 10 different windows and make sure to loop back in after every turn and you will be burning billions of tokens per month in no time.

  • coding harnesses loading entire code bases for every task - at least that's my theory because I also never even get close to the limits of my 20-sth bucks level subscriptions.
  • What is a "normal" subscription? Are you using Claude Code, or just Claude??
  • If you don't reset sessions eagerly or compact regularly it is easy to consume billions in input tokens while Claude churns away.
    by oezi
  • > So how these companies and people manage to use these absurd amount of tokens is a mystery to me.

    Fire and forget. They run multiple agents in parallel 24/7. AI isn't just a rubber ducky for them, its their main (only) tool at that point.

  • >>So how these companies and people manage to use these absurd amount of tokens is a mystery to me.

    Absolutely!

    I know some colleagues who are routinely spending thousands of dollars worth of tokens, I can't see to even max out the subscription limits even if Im working all the time. Curiously enough their output is lower too.