Discussion summary

Discussions highlight the importance of tools for LLM usefulness, the availability of GLM-5.2 from multiple providers mostly outside China, and concerns about hardware and model migration in AI development.

What the discussion says

  • Good tools are essential for LLM utility.
  • GLM-5.2 is accessible from many providers, mainly outside China.
  • Migration to open weights models increases flexibility.
  • Hardware limitations may impact future frontier models.
  • API strategies can hide proprietary technology.
“LLMs are quite useless without good tools.”
— jillesvangurp
“Most GLM-5.2 providers are outside China.”
— kristianp

Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • > Of course, this was a hugely poor read of where the costs actually lie in AI. Training - while no doubt capex intensive - is a fixed, up-front cost. You spend hundreds of millions to train a model, then you are "done".

    I don't understand this point that people make. If you're consistently needing[0] to train new models and the cost of training relative to the % improvement seems to go higher, isn't this just a constant cost that you continue to bear? The footnote seems to allude to this, but then sort of waves it away anyways. Also are there continuing incremental training costs to keep models relevant? Or do they only have knowledge of events up to the day they were trained?

    [0] needing, because you have competitors and people expect more and more.

  • Last month, I cancelled my Claude Pro subscription and instead used those 20$ to purchase Openrouter Credits. Most of my knowledge-seeking questions can be answered by Gemma4, for basic code editing, Qwen3.6 27b is enough, and for really difficult tasks, GLM5.2 doesn't leave me hanging. I'm by no means a heavy AI user, so I'm even saving money going the API Credit route and relying on the smallest possible model depending on the task complexity.
  • Unlike the belief that frontier AI is expensive due to a high margin, and going to be expensive if there is no competition. My understanding is that, under certain circumstances (which is most likely true), the price will be driven down just because of profit seeking.

    The frontier LLM labs run on a huge fixed cost and very low marginal cost. They need the economies of scale to make sense of the business (an incentive to expand their user base as large as possible). Imagine that you want to buy a few B300s to run GLM 5.2 and rent the service out to other people. How could this business be viable and sustainable in the first place? You need as many customers as possible. If you charge everyone $1000, you find fewer customers who can afford it. It rots the ROA if the servers are not utilized 100% (you would better buy less compute instead).

    Also, the marginal cost for onboarding a new customer is low. And it's getting even lower when you have more customers. You wouldn't leave money on the table (especially for your competitors) if you want to maximize your profit.

    By this logic, all frontier AI labs are incentivized to lower the price to maximize their customer base, profit, and ROA.

    by typ
  • It’s important that none of these entities can collude to price fix. Having China be the competitor ensures that.

    Basic microeconomics is still the easiest way to understand token economies. How is it not a competitive market (where profits go to zero?).

    Anything A or O does to keep more margin, any competitor can copy or choose to undercut, and undercutting has the benefit of collecting training data. So what is going to stop gross profit of tokens going to zero except for collusion/price fixing?

  • They have a vision MCP to make up for the model itself not having the capability natively: https://docs.z.ai/devpack/mcp/vision-mcp-server

    I also found their web search to be mostly okay.

    Furthermore, in case this is of interest to anyone, if you use their ZCode harness then you get bigger Coding Plan quotas: https://zcode.z.ai/en

    Used it for a bit, it sits somewhere between OpenCode Desktop (still new but nice) and Claude Desktop (recent versions are good).

    As for GLM 5.2 as a model - with max thinking it’s generally satisfactory, somewhere between Sonnet 5 and Opus 4.8, better than DeepSeek V4 Pro for sure.

    Pricing wise, the subscription doesn’t seem as good as expected. I spent like 60% of the weekly limits of the Pro (50 USD) plan in one day, only because each 5 hour limit only gave me 20% to spend, otherwise it’d be 80-100%. Not even doing anything crazy, just parallel long form work on 2 projects with about 96% cache rate and at most 3 parallel code review sub-agents.

    Their Max (100 USD) subscription would last me the whole week, but so does Anthropic for the same money and so would OpenAI. Off-peak is more palatable but I can’t just twiddle my thumbs at 9 AM to 1 PM local time.

    Proper savings would show up with the Max plan and yearly billing, but that’s more of a tough sell.

  • Meanwhile:

    > China’s Ministry of Commerce has led meetings over the past month with major AI companies, including Alibaba, ByteDance, and http://z.ai/, to discuss measures that would restrict overseas access to cutting-edge AI models, including models that have not yet been released.

    > The discussions reportedly include not only closed-source models but also open-weight models.

    > Future regulations could take the form of a tiered framework based on technological capability. Basic open-source AI models may be managed through a filing system, high-performance models may be subject to security reviews, and the most sensitive frontier models may be banned from public release or restricted to use within China

    https://www.reuters.com/world/beijing-is-looking-curbing-ove...

  • I'll agree but from the other direction. AI continues to absorb my job as a senior systems software engineer (c/c++) and after a couple months I've only spent a few hundred dollars using gpt-5.5/5.6 and codex. I have no idea what people are doing to burn so many tokens but for me this is laughably cheap and every day I discover new capabilities. I don't care if costs go up or down, it's so cheap for what I get that I don't care.
  • I'm not convinced raw costs matter:

    1. Compute costs collapsed since the advent of Cloud and yet hyperscalers still have fat margins.

    2. Many open source office suites exist yet none compete with the ubiquity of gsuite or office. GitHub, Slack are similar examples.

    3. Both Windows and macOS dominate the home desktop space despite free alternatives existing for a long time.

    4. Many formerly open source infrastructure components like Redis and Elastic Search have Apache equivalents, but they still command healthy margins.

    I understand the arguments for a margin collapse, but I don't see any historical analogues. It seems that enterprises will pay top dollar for service guarantees, integration, and someone they can sue.

    It's nobody gets fired for buying IBM all over again.

    by fny

Explore Birbla archives