Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • Cloud > Local

    I have a stack of ten or so 3090s sitting in boxes, but it's not worth the hassle to use them. You can easily run models as cheap as water in the cloud.

    Sitting around 15 minutes for local Minimax is stupid when you're trying to be productive. You can spin up parallel job instances and multitask in the cloud.

    If you want freedom, build open source cloud infra.

    You rent your ISP line. Why isn't renting GPU compute seen the same way? You still have compete ownership over your stack, you're just letting someone else deal with the capital outlay and headache.

  • Do you really keep ten cards in boxes just to brag on Hacker News? Geez, our industry has gotten pathetic.
  • > You rent your ISP line. Why isn't renting GPU compute seen the same way?

    Is it mostly seen the same way. What is unacceptable is removing the freedom (which you mentioned) of people who prefer to run models locally.

    Just as people have the right to tinker at home with DIY and electronics, fully knowing they won't compete with the latest ASML machine, people are free to use open-weight models at home.

    BTW Minimax H3 running local can generate amazing short vids quite fast: about 50 seconds to generate a 7 seconds vids on a 4090 (depending on the settings). I've got a friend who spams my Telegram daily with such (no censorship and NSFW btw) vids.

    I don't run local but I'll defend the rights of people who want the freedom to do so and I won't look down at them from my high-horse talking about "electricity" and "productivity".

  • Why do you have ten or so 3090s sitting in boxes?

    I mean obviously it's worth it just so you can flex on HN. But curious whether there was any other reason? Retired scalper?

  • Why not rent those out on vast.ai and make a little money each month?
  • Why are you hoarding the 3090s T.T They've jumped from $700 to almost double on eBay.
  • The value prop really depends on what you're doing.

    If you're just vibe coding with giant frontier models, yes, the value will be worse. Especially now, where GPU prices have spiked another 20% last month.

    For some tasks where owning the setup and full kv cache matters, the payoff calculation is ridiculously in favor of running your own deployment.

    For instance for some batch classifications jobs where the prefix cache hit rate will be >95%.

    The calculus also changes if you just use AI as a light tool while coding and don't need the giant models; qwen3 27B runs at 80TPS on a 5090 properly deployed.

  • Deepseek will run fine on a single m5 max, cutting that 24 years in half.
  • also worth remembering API price includes power, just dividing price of the hardware by usage doesn't
  • Isn't DGX already legacy? I mean 128GB in 2026?

    Clearly nVidia and others are gatekeeping technology from the pleb so that the rich who own the datacentres can charge us massive margins.

    Oh the debt or not making an even they are supposedly "suffering from" is just a classic mechanism to avoid paying taxes.

  • AMD AI Halo machines are similar cost and RAM. I'm in the market and comparing both. The new 192GB AMD box will be tempting but I'm leaning 2xDGX's for now.
  • It is worth mentioning DGX is about a third up to a half performance of a 6 year GPU Rtx3090...

    I prefer to stay with my 3090s.

  • The 3090s also don't have enough VRAM to run larger models too. It really comes down to how you like to develop (synchronously with lots of steering vs async agentic)
  • Appreciate the heads-up. Even if some of this is overblown, the transparency thing is a legit concern. Gonna double-check my config before I touch OpenCode again
  • Rippling had a writeup on this. 40% of R&D payroll going to tokens, one engineer at $50k/month. They got it down 37% just by routing cheap stuff to cheap models, no usage cut.
  • Fellow readers, would anyone please mind sharing their current experiences? qwen3.6-35b-a3b for local inference, GitHub Copilot Chat was previously worth it, and no longer is, tried OpenRouter and still read through their rankings to see what the industry is actively using, wholesale migrated to OpenCode Zen/Go.

    Does this mirror what other people have been experiencing in waves?

  • I've been using Qwen3.6 models locally for a couple of weeks. Both the A3B moe and the dense variant. The moe works well in Librechat combined with my local search/Web retrieval system. All components use open source projects such as SearXNG, Crawl4AI, MetaMCP, Jina rerank, but all needed quite a bit of coding to work nicely together.

    I get 140 tok/s on short prompts on an rtx3090 on the qwen3.6 moe which makes is easily 4x the speed of Chatgpt or Claude doing Web research.

    But it is a much simpler model. It is only good for simple queries, usually I search for cheapest product in stock in my country available online and stuff like that.

    I use the dense model for planning and such, but on its own it is much inferior to for example opus. It needs careful pipelines that check facts and such and in such harness it can be used for mamy tasks.

  • That's only true if the value of keeping your data and code private is zero. And in that case, Anthropic and OpenAI subscription plans may be even cheaper per day.
  • If you’re doing real work you will hit limits quickly on anthropic and codex paid tiers
  • I'd really rather just pay Deepseek directly. Why wouldn't I want to support the company that trained the model?

    It isn't even really worth the (minimal) ops to stand up rented MI300Xs to sell excess capacity to them even if it was minimally profitable, when I tried I was content to give API keys to friends to beat on it.

  • There is some fear mongering, because - China. I wouldn't want to send even the dumbest of my half assed of my ideas to their servers... but then, could be a vector for some shenanigans... but, then as if we have a reasons to trust either Sam or Dario. IDK, but interesting to think about.
  • According to OpenRouter, DeepSeek trains on input. At least when disabling all providers that train on your data, DeepSeek gets disabled.

    Other providers don't.

    You can also choose to route to ZDR only.

    https://openrouter.ai/deepseek/deepseek-v4-flash-0731#provid...

  • You can do this but you still need a harness. I like open code because I can use literally any model through it and can connect directly to the provider (although I use open router)
  • The reason to run local models is not for coding mostly it's for learning how to deploy models and tinker with self hosting. It's also for massively crunching data 24/7. Imaging having an agent analyzing constinous log streams etc .. that could be a usescse where even deepseek could add up cost.
  • Also privacy and offline capabilities.
  • For me its stability and workflow.

    Models are all non deterministic and local is the only gauntee that your investment can continue to pay.

    Cloud models will continually change nondeterminism ontop of the model.

  • Privacy, no need to pay for tokens, offline usage, endless possibilities - you do not have to pay for tokens (subscribtion payment), so you can do more. The big issue is that local models are not "ready yet" compared to frontier and paid services. It is hard to run decent model without decent hardware. And to be honest even if you can buy hq hardware and spend a lot of money then it is not the same quality.
  • That's almost half way to their stated 6X usage goal! Just from DeepSeek.

    > With Go, you pay $10/month and we aim to give you 6x that in usage.

    For most models, we make this work through bulk discounts and reserved GPU capacity. We then pass those savings on to you through the 6x multiplier.

    https://opencode.ai/docs/go/#why-some-models-have-lower-usag...

    I think it's back to 2X usage, meaning it's cheaper token-burn than usual to use. Which is lovely.

    OpenCode Go has been so nice to have. I love having access to DeepSeek, Qwen and MiniMax M3 when doing design work, to see what different models cook up. I've been very surprised with MiniMax M3, not as a particularly good architect, but at it's very good ability to state the problem elegantly & to frame the different decision points very well. That's been a fun ongoing surprise.

  • Don’t forgot a number of models they have offered to you will train on your data
  • Since the announcement of DeepSeek price hike, I have been using MiniMax M3 and I am really surprised by the quality this model spits out. Subjectively speaking, the bullshit MM3M produces is waaay less than DS0731. Could it be a result of less hallucinations than DS0731? https://artificialanalysis.ai/evaluations/omniscience#aa-omn...

    Maybe hallucination is good for prototyping and creative work, but maintaining and debugging code might be better done by a boring model?