

Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- Pretty impressive to see the amount of performance they can squeeze out of the same hardware. I suspect the same process will play out for all combinations of LLMs, inference providers and hardware stacks. This should bring down the cost of inference for the providers by an order of magnitude in the next year and lead to fantastic margins for inference providers.by a11r
- Time to tackle consumer GPUs next, since I’m not getting that Intel Arc B770.by KronisLV
- If only this infrastructure could handle all the traffic. I've tried using glm via z.ai - and it's a snail kind of slow.
And at the same time you have pretty strict limits to your usage, so in many cases you can't even let it work all night, as you will reach your limit faster than that.
by konart - Interesting that the tone of announcements between US and Chinese providers is converging.
GLM has in the past been more technical rather than speculation about future development on RSI etc.
Also curious whether those 100k accelerators are entirely locally made. If that's genuinely end to end on all components including lithography, memory, design etc then that is quite a feat.
by Havoc
This whole thing sounds like industrial scale auto-research, but done by people who actually know what they are doing."We implemented a series of aggressive memory optimizations, including..."by throwa356262- We built a complete production-grade inference service from scratch on a cluster of more than 100,000 Chinese-made AI accelerators. All production inference for GLM-5.3-Flash runs on this system.by dada216
- US chip export restrictions may actually be an advantage for China's AI Infrastructure. Chinese companies are forced to speed up developing their own AI chipsby zicohacks