Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • I think the frontier will command premium for sometime just as slight better software developers were 10x's vs their peers as their architecture & development strategies and code approach compounded quickly. One less error per block of work compounds quickly.

    Sure, there may be some cases and reasons for local models and industry is so large they will continue to make progress and gather economic value and users for specific use case; but frontier will command vast majority of the economic value distinct from Linux and open source where the model created better than proriatary economic incentives around development

  • Youre clutching at straws.

    Ultimately its a financial game. Open source is far cheaper so it already has an upper-hand. Frontier models have to justify financially why they are worth the additional spend.

  • 10x developers were not slightly better than their peers, they were vastly superior and faster. OTOH, the lead of frontier llms is diminishing as training is getting diminishing returns.

    Also, on that note. Not every company needs 10x developers, just as not every task needs frontier llms. Ultimately, operating costs will be the largest contributing factor.

  • It was easy to be a rebel and use Linux when it was clearly competent, but needed hacks and extra elbow grease to get it polished for use. IME, the open models are “not there yet” in terms of capability or operational needs. Sure, GLM5.2 looks competent, but I will only be able to get it to run that competent if I had a huge cluster of GPUs.. if I am accessing an open model via hosted API, I might as well run a closed model via hosted API. The incentives fall apart in comparison to using Linux 15 years ago.

    Don’t get me wrong. I wish I could run a local model and be happy about it. At the moment, I’m not.

  • > if I am accessing an open model via hosted API, I might as well run a closed model via hosted API.

    uh.. no?

    The whole thing is that it cannot be enshittified, because there's not just a single party having control over it.

    As it has happened, is happening and will happen.

    With open weights, you cannot easily be rugpulled or locked out or any of that stuff. If the corp attempts that, someone else with an server farm will gladly take you as a customer with absolutely 0 changes to your workflow other than swapping out the API URL + Key.

    You'll be talking to the same model with the same personality and same knowledge.

  • What makes an open model worse is ultimately the budget : you have access to worse data, not SOTA models, less GPU compute time, and having a good fine tuning team is extremely expensive. Linux works because the entry barriers are purely on a software side : a lot of contributers all around the world can outclass any OS by contributing on their scale to Linux. All you need to contribute is a computer, and your brain. Open models don't have the same community push, they rely on core ressources that not anyone owns. And injecting them in the model costs too much money. If there are no public breakthroughs in the way we train large open models that makes community led models 10x better, the shift to open models will never happen on a large scale.
    by GL26
  • While I agree with some of the gist of the article, 2 remarks:

    1. Unfortunatly in my tests the open models do not (yet?) rival, at least Claude Opus, for software development/engineering and adjacent tasks.

    2. Enjoy while it lasts. I'll be genuinly amazed these open models will not be declared 'illegal' under some security pretense by the end of the year. And I say 'pretense' because the primary driver will be regulatory capture and industry protectionism.

  • Banning models in US just strengthens competing states, ie. China.
  • I’ve been wanting to get better acquainted with local inference but I don’t have the hardware, which has made me think about something I haven’t seen discussed, which is local collaboratives. The economics makes it seem like a group of people joining together to run good hardware and an open model might make sense, but I haven’t seen anything like this mentioned. Have I been missing it?

    I think it would be pretty neat to launch a service helping people who wanted to participate in something like that locate one another.

    by bnj
  • Open models hosted in Cloud???
  • There are plenty of providers of open models that offer very affordable rates. Generally, I recommend looking at OpenRouter since they track various metrics for the various providers.
  • The reason you don't see more of this is because everyone does the math, realizes it's not a good deal, and then gives up on the idea.

    There's a post at the top of /r/localllama about this exact math right now: https://www.reddit.com/r/LocalLLaMA/comments/1ubrcwj/tokenom...

    TL;DR: Running GLM 5.2 is going to cost about $20K minimum, and that's going to be painfully slow compared to the cloud hosted versions. Even the estimates where the server is computing tokens 24/7 you can't break even for several years.

    The only reason to run locally is if complete data privacy is your top concern. You pay a high premium for that.

  • Have you read about Opencode Go? They are great provider for open model, like GLM 5.2, Deepseek v4 Pro, Kimi 2.7 Code. You should give it shot to them :-)
  • The amount the HN community, at least from what I’ve seen, is sleeping on OpenCode Go (and zen) is kind of amazing.

    $10 a month gets you generous usage with the best open weight models and they claim to have zero retention and not to train on your usage.

    It’s unclear to me what the advantages of openrouter are but it seems to be a default I see many people talking about here.

  • Sure. But OpenAI is the same price. Why would I pay $18/month for z.ai when OpenAI is $20/month?
  • OpenCode Go is $10/month and the limits are much more generous than those or Codex
  • the pricing page doesn't seem to call it out anymore, but the claim on z.ai coding plan used to be 3x the usage of the equivalent-price claude plan. whether that's accurate i don't know, but just based on api pricing GLM is way cheaper.
  • Why pay a monthly fee when you can pay for exactly the # of tokens you actually consume?

    The API rates are very affordable once you start to optimize for the fact that prepaid tokens seem to massively outperform other kinds of tokens.

    I can often do with 1 million tokens what my peers have failed to do with 100 million. For me to spend more than $200/m in prepaid API tokens I'd have to pull a 007 work schedule.

  • One reason might be request limits. OpenAI's ChatGPT Plus w/Codex ($20/month) provides a worst-case 5-hour-request-limit of 15 for GPT-5.5, 20 for GPT-5.4, 60 for GPT-5.4-Mini. Whereas Z.ai Lite ($18/month) provides a worst-case of ~80 for GLM 5.2 (off-peak; on-peak is 2am-6am New York time). So Z.ai can provide higher limits for a cheaper price. (https://codeberg.org/mutablecc/calculate-ai-cost/src/branch/...)
  • One big advantage I’ve found — people get attached to models (including me). With open models if you find one that works perfectly for you but the next version doesn’t, you can run the old one forever (or someone will for you)
  • Claude started becoming useful for my coding purposes after it hit version 4.6. After that sure some nice to have additions but I think if I had 4.6 sonnet & opus as open weights, I would not need something more.

    Having played a bit with Fable, reinforced the above.

  • I agree and I'd love for local models to hat the sonnet 4.6 level but nothing seems really all that close, and I'm not particularly excited about giving money to deepseek.
  • Yeah for me the coding inflection point was relatively recently (GPT 5.3 perhaps). There's just a threshold they have to hit to be consistent enough to avoid having to redo work and only the later models started delivering it.

    This certainly seems feasible for open weight models eventually, but I'm still extremely skeptical of the claims about reaching this level with any open weight model that can be run locally (nevermind the hardware costs to do so practically).

  • What's amazing about these models is they are effectively a distillation of the internet in something that can fit onto your local machine [1] and be queried via natural language.

    [1] It seems inevitable that decent local models will be possible as the technology and the hardware is improving at a rate beyond the growth of the knowledge base to be distilled.

  • The headline says one thing, then the article text says this:

    > I’m hoping it’s going to be minimal.

    I have multiple subscriptions and I pay per token to try out different LLM providers through OpenRouter. I also run open weight models locally.

    I just can’t agree yet. The models from Anthropic and OpenAI really are that much better than anything else. The open weight models must be universally benchmaxxed across the board because my real world experience with them is very different than what the benchmarks imply. I get downvoted a lot for speaking about my experience because I don’t think it’s the reality that people want to hear right now, but it’s true for complex work.

    I do think there are a lot of easier tasks that can be handled appropriately by the open weight models in the hands of a skilled operator. If an entire job is simple enough that you wouldn’t hesitate to hand it off to a junior with a little supervision then any model will do. However for a lot of the work I do, even Opus 4.8 on Max requires a lot of attention and extra steering and review to keep it on track. Fable did, too, though to a lesser degree. When I try to use the big open weight models (hosted, because they’re not running at reasonable speeds locally at a quantization I can tolerate) it feels like I spend more time waiting while they burn tokens for output that I probably have to reject anyway, at least for the bigger tasks. I wish they were there, but that’s not the case yet.

  • Do you have any example?
  • The article also contradicts itself halfway through:

    > There remains a clear penalty for being an open LLM user.

    The conversation here _around_ the article is interesting, but the article itself boils down to “I’m going to try using open models and hope for the best.”

  • I find the attitude shown in this post very surprising. On the one hand, the post starts with a story of adopting Linux and other FOSS. The core of FOSS is giving its users the ability to understand and modify software they run. On the other hand, the rest of the post is about using a tool (LLM) that the author has no way to modify and no way to understand. Huge matrices of floats are at best comparable to compiled code. But the reality is even worse - it’s actually easier to decompile and understand proprietary software. Not to mention the fact the most of the time users can’t even run the “open” models since it requires hardware that most can’t afford.

    How did we get from prising software freedoms to this?