Discussion summary

Discussions centered on rising RAM prices, the cost of AI compute, and the increasing expense of tokens for AI models. Participants debated hardware options, market manipulation, and the impact on employment.

What the discussion says

  • RAM prices have increased, affecting costs.
  • Nvidia's market dominance is criticized.
  • Token prices are rising despite hardware improvements.
  • Some suggest companies subsidize token costs to reduce workforce.
RAM prices skyrocketed, impacting AI costs.
shevy-java
Token per dollar is more expensive, indicating higher costs.
calin2k

Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • The compute-in-memory and neuromorphic paradigms are likely to push this much, much farther over the next decade as more radical improvements make it out of the lab. Sooner or later it will involve new materials and new nano devices and providing multiple orders of magnitude better efficiency. And just scaling up existing things like MRAM.
  • Slight criticism of the headline there, you can't get cheaper per dollar.
  • I was hoping they would be discussing some path to improving things faster and cheaper. But in this post it looks like they offer quantized version for the same price as full version, and a fast version at much higher cost.
  • Do these providers have 80+% gross margins or is something eating into them? Maybe utilization?
    by oDot
  • hi i work at wafer. no the margins are lower averaging at about ~40%. utilization is one of the highest order bits in determining margins here, yes.
  • The 2600 tok/s is an "aggregate", not the actual throughput.
  • yes it is 213 tok/s single stream (so per user)
  • Not a new phenomena - performance per dollar has been fairly steadily exponentialling since 1900 or so

    1900 - 2010 https://www.thekurzweillibrary.com/exponential-growth-of-com...

    1939 - 2023 https://medium.com/@timventura/kurzweils-law-for-the-ai-age-...

  • I'm not surprised to see competition with Blackwell. Rubin is 5x faster than Blackwell at inference - Blackwell is the last generation Nvidia didn't optimize specifically for inference.

    If I'm missing something, please let me know!

  • how do you get 5x faster at inference when inference is memory bandwidth limited? getting 5x the memory bandwidth of a h100 seems physically difficult.
  • It's very unclear what's special in Rubin to be optimized for inference? I can see disaggregated bit (with having separate prefill and decoding nodes), but what else?
  • Agentic coding drivers for different architectures is a massive unlock for the world

    So much compute is under utilized waiting for a savant or company to prioritize an architecture, and now all the other engineers can tackle this at any time if they get inspired on the right prompts

  • this is exactly our thesis at wafer :) thank you for the support
  • Personally, I can't wait till something like this starts getting to consumer level. https://www.anuragk.com/blog/posts/Taalas.html
  • There’s noticeable accuracy degradation when they switched from fp8 to mxfp4
  • And somehow they claimed that it is "lossless".
  • Wafer discontinued their own "Wafer Pass" flagship coding plan within weeks of launch and had to issue prorated refunds. Now they're bragging about squeezing costs down even further via quantization, even though their implementation is clearly lacking.

    [1] https://www.ycombinator.com/launches/Q9i-wafer-pass-flat-rat...

  • I think we should make it illegal to not specify the quantization in the headline for these types of posts.
  • A nice filter is checking for the `.ai` in the end. It is very likely slop if you see that. Slop meaning low-effort/clickbait/shallow/useless/scam etc.
  • And to use the heading "Why this matters".
  • Its MXFP4
  • I don't know what you mean by "quantization". I guess you're asking to explain how they measure "performance", but that kind of thing often won't neatly fit in a title.

    I do think the phrasing is weird, though. The performance may be improving, but it isn't the thing getting "faster" (e.g. responses to queries might get "faster"). And the dollars aren't getting cheaper; the performance is. "Performance per dollar" is a rate; it is not getting faster or cheaper. It should just say "Performance per dollar is increasing".

  • While cool, quantization to FP4 is practically never lossless in actual use. A lot of providers are advertising high TPS on Kimi and GLM, but the models are functionally lobotomized and no longer close to frontier quality. Would love to see this not be true.
  • from memory, it is like 96-98% of the accuracy.
  • First thing I noticed as well
  • Doesn't Nvidia with their NVFP4 claim that it's lossless?

    I haven't tested enough models Nvidia has converted to NVFP4 besides GLM 5.2 but it seemed fine to me.

    My own luck has been hit or miss with it.

  • MI355X can perform FP6 operations with the same speed as their FP4 (unique to AMD) - people should be making MXFP6 quants which would be pretty much lossless, and much closer to FP4 performance than FP8
  • Kimi uses INT4 as its native format, there's no such thing as "better than 4-bit precision" for that model. This is in contrast with GLM for which 16-bit precision is native and 8-bit is in common use.
  • Can you folks add performance per watt as a metric to these comparisons, I honestly want to understand where AMD fits in the stack in terms of actual performance to dollars. I have had talks with companies wanting to build data centers outside of US and find it hard to source anything Nvidia in sufficient capacity and scale.

    If AMD is competitive performance per watt and roughly reliable in terms of software support which is what most folks outside of US prioritize above all else, since outside of China and US electricity tends to at a relative premium.

    Maybe if they make smaller data centers viable at the right price, AMD could be part of the stack outside of US where ever Nvidia is more limited in supply. Though I have genuinely no idea what sourcing an AMD GPU looks like.

    I have never seen a company use AMD outside of wafer and a couple others mostly in US.

    Genuinely intriguing or maybe not really (could be this stuff is common knowledge) and I am just stuck in my Nvidia bubble here.

  • Typically any company that can’t get Nvidia to fill their orders will have at least some AMD.