Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- although i initially thought it didn't make sense financially to run this kind of model locally, i did run the numbers and for heavy users this could justify buying $10k worth of hardware with a ROI over a few months, less than a year.
I was looking at my token usage, mostly from subsidized codex/grok subscriptions and i'm a somewhat heavy user. The thing is i would actually use even more tokens if it wasn't for the weekly quotas.
In the end, with a $10k investment and running this kind of model, estimating a 2x increase in token usage because i wouldn't have weekly quotas and comparing to glm api prices, this thing could pay for itself in less than a year.
Obviously i'm paying subscription price right now, so the math doesn't work. Although using local ai removes all weekly quotas. Keep a subscription to have access to frontier models for planning work, and local hardware + glm-5.3 flash for implementation, e2e testing, qa work 24/7.
It's not that crazy of an idea and the numbers aren't that bad.
by guybedo - You aren't going to get nearly as much token usage locally from DGX Sparks or even M5 Ultra (though it might be close, unsure would need to get my mittens on it to clarify).
You will get around 2-4 concurrent streams of aggregate tokens at best for such a model and around 0.5B output tokens per month assuming you use loops and run it when you are sleeping. That's 500 (per mill) * 0.5$ = 250$ only at most.
Then there is maintanence and efficiency costs due to electricity usage and such, any down time, etc.
You will be lucky if you can squeeze more than 200$ of value out of it in a month.
I don't think people should buy local hardware for money reasons, by the time you will pay off a 10K USD machine, 2-3K USD machine will catch up and beat it by a significant margin.
Unless your expectation is that we will be in hardware winter for the next 10+ years. At 200$ per month it will take around 200 * 50 = 10k, that is, 50 months, so around 4-5 years.
Again assuming you are making the most of your hardware somehow, very hard to do in practice.
I don't recommend people to use compute as investment or payoff thing, but if you have the money to burn and can afford it why not, maybe with some software optimizations it will be cheaper but then again Z.ai is currently offering 50% discount and providers will offer cheaper rates for sure.
But either way you will never be able to burn more than 200$ worth of token on a cheap hardware device, because inference becomes more profitable the more you scale it up, you have separate prefill and decode engines/systems, and a lot of nuance, but assume for every 10x increase in infra you increase margins by 5-10%.
So from 10K to 100K to 1M to 10M to 100M.. I don't think this curve continues beyond 100M but I have no idea about that scale unless some AI lab is interested in hiring me lol.
So a 100M infra will have ~30% better margins than you at 10K, then there is software optimizations but that's cheap enough, though some of it is only viable at scale.
Either way assume 10K is the price of privacy if you really want to buy it. Don't worry about making the most out of the usage, you will always be in a net loss but I would assume for you 10K doesn't matter.
by minraws - Nice, finally they fixed the huge reasoning tokens count.
Now it's similar cost to DeepSeek v4 flash, but smarter.
My tests: https://aibenchy.com/compare/z-ai-glm-5-3-flash-max/deepseek...
by XCSme - > 320B total parameters and just 18B active parameters
This is pretty hefty for a "flash" model, even a 256 GB setup is insufficient at q4 - and q4 is already the worst-but-still-acceptable quant in my experience. The benchmarks look great, especially since GLM tends to be more honest than the average Chinese lab, but you’ll need to splurge to run it at home.
@edit: so many releases that I forgot to math. This fits just fine in q4, realistically the minimal hardware would be 192gb - so blazing fast on double rtx 6000 pro and usable on 256gb unified memory. You could even go with 5bit quant on 256gb.
… you’ll still need to splurge, though.
- Speaking as someone who isn't really well versed in this, does 18B active parameters mean that you could potentially hold only the 18B parameters in RAM and stream the rest from a fast NVMe SSD for acceptable performance similar to how Colibri works?by yonatan8070
- Looks like the M5 Ultra Studio wait times are going to increase again. Already at 10-12 weeks, I wonder how long it'll go?by dannyw
- That's 160GB-ish for Q4...how is 256 insufficient?by colingauvin
- So the vagueposting by googlers about Ox Alpha was just... what exactly?
Like I get that they have to be careful about comms, but surely senior members of the team can clarify when something is NOT them, when everyone is gosspiing it is them.
by preommr - On Twitter they mentioned that it was unfortunate timing as the 3.7 flash release collided with ox alpha.by asar
- Trolling. GLM is heavily distilled from Gemini.by qeternity
- The lack of measurement causes existence of such claims or discussion. Last Friday I built this Model fingerprint calculator and I tested between OxAlpha with all other claimed models, the only match was GLM. It generates, or measures regardless of the model-weights or its training data. No need guessing when one can measure it. I tested with Gemini family too, far different. Here is the link to my experiment https://github.com/unclecode/modelprintby uncleocode
- If we fast forward say 5 years, I don't see how we don't end up in world where people (and enterprises) are more savvy with how they use LLMs. Meaning, more models, smaller models, weirder models, more specialized models, etc. And all of it running on a variety of hardware (edge devices, personal computers, on-demand cloud compute).
I don't see how NVIDIA can keep their spot as belle of the ball. If LLMs and friends are truly to become as useful and ubiquitous as everyone thinks they will, then commoditization is the only option.
by cootsnuck - > I don't see how NVIDIA can keep their spot as belle of the ball.
FWIW, people were saying "ASICs will kill CUDA demand!" since the crypto mining boom. Then a few months later, CUDA found another niche application in LLM applications.
With the mounting demand for robotics, surveillance and autonomous weapons, I don't see how Nvidia couldn't keep their spot. They have their pick of the litter with hundreds of market segments, and unlike the rest of FAANG they're not afraid to branch out.
by bigyabai - We need to figure out what the real pricing is for a going concern. Right now, everyone is subsidizing and discounting to grow (or maintain) market share. The big question is whether the steady state, market derived inference pricing is above or below what we’re seeing today. I honestly don’t know. Anthropic had said that inference is profitable, but they’re clearly not yet profitable overall with training and buildouts still happening.by drob518
- > with all of this traffic served on Chinese AI chips
RIP Nivida shareholders
by sunbum - Not really. Chinese AI companies were never using NVidia AI chips.
This announcement doesn't really mean anything at all. It means the very few people who are already using Z.ai's API will continue to do so, but the vast majority of money going to Nvidia is through the massive amount of business going to Anthropic, OpenAI, and other western cloud providers and inference providers, who are mostly using NVidia chips for inference.
Also, NVidia chips are still sold out and supply constrained.
by saberience - God I wish I could’ve shorted NVIDIA right now
- This is no surprise [0] [1].
>> "They are already there on open weight models and Jensen knows that it is only a matter of time until China catches up with GPUs or other AI accelerators."
It is also why Nvidia becoming a bank for other AI companies who are unable to find VCs to fund them isn't really a good thing and that is bearish.
by rvz - Most US companies that have anything to do with government, finance, medical, etc. already have contractual or regulatory obligations which prevent them from using Chinese hardware or services, even before the AI boom. That's a huge market.
Nvidia will do just fine. (Disclaimer: not a shareholder. At least, not directly.)
by bityard - Ox Alpha is a smaller model and it was running very slowly. Chinese AI accelerators are coming along, but nVidia’s lead is huge.by Aurornis
- I don't see a situation where subscription payers move outside American LLMs (chatgpt, claude, gemini)
And I don't see a situation where serious API payers are OK with handing the Chinese state all their data. Like manufactures of decades past did and learned a hard, even existential, lesson for it. The state mantra has been "Collect and Copy" for a long time now, tech just hasn't had that moment to experience it yet.
So that leaves local hosting/leasing, but one of those has totally non-practical economics and the other doesn't have enough compute to meet any kind of real demand.
I also have yet to meet a single person who isn't neck-deep in the tech space mention a Chinese LLM. It's 100% the big American three.
If anything it's custom chips from the labs that threatens Nvidia.
by WarmWash - Another self-inflicted own courtesy of US government policy.
While I think China would always get to hardware self-sufficiency eventually, all export controls have done is (1) accelerate China's development, and (2) divert revenue that would've otherwise gone to NVIDIA/AMD/etc instead.
by dannyw - This is the takeaway here: That's how they have been serving it at scale as Ox-Alpha. This is a definitional moment.-
Further quote:
"Compared with our initial baseline on the same hardware, we achieved a 3× improvement in end-to-end serving performance, reaching hardware efficiency and per-token cost comparable to mainstream NVIDIA GPUs. This demonstrates that Chinese chips can support frontier-model inference efficiently and economically at scale."
by Bluestein - > Before release, we tested GLM-5.3-Flash anonymously as ox-alpha on OpenCode and OpenRouter to gather user feedback. It quickly became the most popular model of the week — with all of this traffic served on Chinese AI chips.
Just like that we are witnessing an open burial. It's now in everyone's interest to keep the valuations in the 'A.I' economy as they're though it's apparent they're not justified.
whether it's the cost to develop models, cost of hardware, cost of serving ie inference.
by dzonga - Weren't they giving free access? Not exacty a meaningful heuristic if soby micimize
- If you're on opencode's go $10/mo plan and want to use GLM-5.3-flash right now on pi, you can add this to models.json until pi updates to support it:
{ "providers": { "opencode-go": { "models": [ { "id": "glm-5.3-flash", "name": "GLM-5.3 Flash", "api": "openai-completions", "baseUrl": "https://opencode.ai/zen/go/v1", "reasoning": true, "input": ["text", "image"], "cost": { "input": 0.15, "output": 0.5, "cacheRead": 0.03, "cacheWrite": 0 }, "compat": { "supportsStore": false, "supportsDeveloperRole": false, "maxTokensField": "max_tokens" }, "contextWindow": 1000000, "maxTokens": 131072, "thinkingLevelMap": { "off": null, "minimal": null, "low": "low", "medium": null, "high": "high", "xhigh": null, "max": "max" } } ] } } }by bel8 - For openrouter in pi:
{ "providers": { "openrouter": { "models": [ { "id": "z-ai/glm-5.3-flash", "name": "Z.ai: GLM 5.3 Flash", "reasoning": true, "thinkingLevelMap": { "off": null, "minimal": null, "low": "low", "medium": null, "high": "high", "xhigh": null, "max": "max" }, "input": ["text", "image"], "cost": { "input": 0.075, "output": 0.25, "cacheRead": 0.015, "cacheWrite": 0 }, "contextWindow": 1048576, "maxTokens": 131072 } ] } } }
by Kholin - You guys read Z.ai's terms of service, right?
Broad and perpetual license over inputs and outputs, and even your name and profile picture.
Vague prohibitions on whatever may harm Z.ai’s "interests" or even the "national interests" of any country.
Vague prohibitions on "disturbing" or "inappropriate" content, whatever that is.
Vague prohibitions on discussing Z.ai, even my posting this comment violates it.
Can ban you if you, in the "sole and absolute opinion" of Z.ai, have violated these broad terms, and if you paid for the discounted yearly plan kiss your money goodbye.
- Yes, the terms are dubious. But they are also reasonably lenient with enforcement. They also don't require persona id verification, witch is wat turned me away from openai.by gunalx
- I blocked Z.ai as soon as they were loading 10 different external providers including Alibaba who was just proven to execute silent sound fingerprinting mechanisms.