Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • On a purely monetary basis it probably never will.

    You're competing against companies that get tax breaks, locate themselves optimally, and have large economies of scale.

    Also, if it did, the hardware would be bought up, raising the price until there was no economic profit again.

    If you can find a unique application for it then maybe?

  • I want the autonomy but local models of the size I would have the means to host wouldn't be capable enough. What usecases tend to suit these smaller models that tend to produce incorrect or otherwise flawed responses often? Could they work for anomaly detection and what would a rough architecture look like?
  • The math is wrong, the tok/s is at least 2x that, at least with MTP and Q8 KV which you should always use. And the default tokens a day is ridiculously low at least for coding.

    Having said that, it will never pay for itself. A simpler more absolute math is, if I buy a Mac and use it to sell tokens on OpenRouter, will I make a profit? And the answer is no.

  • This tells me that the max throughput for the models I'm running on my hardware is lower than it actually is. Please allow us to tweak all the variables instead of locking me in to whatever rate you found by searching
  • Thanks for all the feedback. You can now enter your own measured tok/s for any machine and model.
  • Fun feature: can you show some sort of list of the best combos? Eg shortest payoff time for best capability in various situations.
    by bix6
  • Good idea, pretty crude but it's up: https://sunkcost.ai/best/

    For each usage level, it lists the quickest pay-back in each capability class, with each model on its quickest machine and one click into the calculator to change the assumptions. Short version: at 1M tokens/day the best Sonnet-class option is Qwen3.8 27B on a Mac mini M6, 8.3 years. It only drops under a year if you're running agents at around 20M tokens/day.

  • Local LLMs are not really about saving money, they're about autonomy. Choose the exact model you want, fine-tune it if you want, and no one can take it away from you.
  • True - definitely agree!
  • Not a fair comparison really. If you can run a model locally then you can somewhat train out the guardrails, censorship, and brand-safety. That has value a subscription does not.

    Idk about the quality of this setup but just pasting it here as an example. https://explainx.ai/blog/heretic-llm-abliteration-guide-2026

  • That can also be done with neoclouds.
    by xnx
  • > If you can run a model locally then you can somewhat train out the guardrails, censorship, and brand-safety.

    When does the average person actually need to do that?

  • Claude Code is $100+ or else be constantly throttled. My usage on GHCP was gonna be $300+ a month.

    I paid $1350 and threw an R9700 in an existing machine. That's a 4 month pay off or so.

    Plus, I can feed it sensitive data all day and not be worried where it's going.

  • today, the r9700 is 1800$ at micro center :-( so payoff is around 2 months?
  • An R9700 has 32 GB RAM. Is your comparison against a similar size model? Or shouldn't you be comparing it against the cost of a hosted model matching the one you’re using locally?
  • I doubt it will ever be cost effective for the foreseeable future. The AI companies have astonishing amounts of compute and they’re effectively dumping it on the market.
  • "If they are selling it for less than it cost to make, buy as much as you can."

    -- Warren Buffett

  • More than that, running hundreds of conversation streams at once is essentially the same cost as running a single conversation. And then you add on the secondary benefit of having the GPUs running nearly all the time rather than mostly idle...

    Local inference makes sense for speciality needs, or very small models. But if your model is bug enough to span GPUs its excessively wasteful to hoard those GPUs for yourself without piggybacking hundreds of other conversations on top of all that memory bandwidth and matrix multiplies.

  • The comparison is about how many tokens you buy vs how much hardware you could buy with the same money. It's as saying "if you have rib eyes at Applebee's every day, how long until cooking your own rib eyes pays for itself".

    If you don't consume many of tokens, it will likely never pay for itself. If you do, though, it will have trade-offs, but you'll probably save money in the end.

  • The idea that you need a new machine is pretty ridiculous. I bought a used HP Omen with a 3090 last month for $2k. 57t/s with Qwen 3.8.
  • I've not heard of others running HP with it. Hows much RAM do you have?
  • I'm so happy for the two used 3090s I bought for $500 each after Ethereum mining ended. I even saw them for like $430 at some point lol.
  • Agreed, I was also annoyed that the only params on the site were mac products. I run qwen 3.8 on a 12 year old asus and a 3090, 50tok/s. It's not even the only guest running on the box. For my usage profile (not running it 24/7) it's actually less expensive per-month than claude subscriptions.
  • 43 years to break even on Qwen 3.8 at 25% the speed of the API, lol. I like the idea of local models for really small tasks like automation/toolcalling, but it will probably never make sense for coding. I tried them and it was just excruciating compared to what you get for $100 a month from a subscription.
  • Yeah it's surprising how long it would take to get back on those local models!
  • It pays off instantly, because OpenAI/Anthropic can no longer see what I'm doing and that's worth a lot of money to me. If I am offloading some of my thought processes to a machine, I want to own that machine. And if I finetune the model, I can gain access to parts of thought space that are cordoned off by OpenAI/Anthropic/Alibaba/whomever due to their "alignment" efforts (i.e. alignment to the AI company rather than me). Otherwise, it's like if someone else owns a part of my mind and has a backdoor into my mind.
  • Also, whatever your doing won't be at the whims of cloud providers; it won't fail because they decided to quantize your customer $ into a shittier model.

    Some how, _instability_ has gained valuable currency, so now we all act like the constant change of whatever is actually good for us. FOMO is just like breathing guys. That anxiety induced by tech culture constantly churning is healthy.

    In reality, these people churn for their own self worth and nothing else.

  • I’m very sympathetic to this point of view but I also can’t remotely afford the hardware required to get in the ballpark of Fable.
  • I'm curious what people are sending to Claude that is so secret. Claude knows about my interior decorating, questions about light bulbs, curiosity about what the Galactic Empire was even trying to do, unpacking SCOTUS decisions, shoe trees, Fed inflation history, etc.

    What part of my brain is contained here? Sure, the conversations have back and forth (some have dozens of exchanges), but, like, that's not the secret to me. I don't think it can replicate me, and even if it could… okay?

    Are you worried they're going to target ads? That the government will steal something? What?

    Claude Code has information about my home server, but google or DDG would also have the broad strokes (torrents). I don't know. Maybe others are working on more sensitive things at home.

    by tyre
  • Agree. As the meme/old-ad goes, "Running it on my own machine? Priceless!"

    Some of us get a weird thrill that we can actually do this. Mind-boggling time we live in.