

Discussion summary
Discussions centered on rising RAM prices, the cost of AI compute, and the increasing expense of tokens for AI models. Participants debated hardware options, market manipulation, and the impact on employment.
What the discussion says
- RAM prices have increased, affecting costs.
- Nvidia's market dominance is criticized.
- Token prices are rising despite hardware improvements.
- Some suggest companies subsidize token costs to reduce workforce.
“RAM prices skyrocketed, impacting AI costs.”
“Token per dollar is more expensive, indicating higher costs.”
Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- The 2600 tok/s is an "aggregate", not the actual throughput.by AussieWog93
- Not a new phenomena - performance per dollar has been fairly steadily exponentialling since 1900 or so
1900 - 2010 https://www.thekurzweillibrary.com/exponential-growth-of-com...
1939 - 2023 https://medium.com/@timventura/kurzweils-law-for-the-ai-age-...
by tim333 - I'm not surprised to see competition with Blackwell. Rubin is 5x faster than Blackwell at inference - Blackwell is the last generation Nvidia didn't optimize specifically for inference.
If I'm missing something, please let me know!
by Schiendelman - Agentic coding drivers for different architectures is a massive unlock for the world
So much compute is under utilized waiting for a savant or company to prioritize an architecture, and now all the other engineers can tackle this at any time if they get inspired on the right prompts
by yieldcrv - There’s noticeable accuracy degradation when they switched from fp8 to mxfp4by p1esk
- I think we should make it illegal to not specify the quantization in the headline for these types of posts.by nxtfari
- While cool, quantization to FP4 is practically never lossless in actual use. A lot of providers are advertising high TPS on Kimi and GLM, but the models are functionally lobotomized and no longer close to frontier quality. Would love to see this not be true.by hassaanr
- Can you folks add performance per watt as a metric to these comparisons, I honestly want to understand where AMD fits in the stack in terms of actual performance to dollars. I have had talks with companies wanting to build data centers outside of US and find it hard to source anything Nvidia in sufficient capacity and scale.
If AMD is competitive performance per watt and roughly reliable in terms of software support which is what most folks outside of US prioritize above all else, since outside of China and US electricity tends to at a relative premium.
Maybe if they make smaller data centers viable at the right price, AMD could be part of the stack outside of US where ever Nvidia is more limited in supply. Though I have genuinely no idea what sourcing an AMD GPU looks like.
I have never seen a company use AMD outside of wafer and a couple others mostly in US.
Genuinely intriguing or maybe not really (could be this stuff is common knowledge) and I am just stuck in my Nvidia bubble here.
by minraws