Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • I've been very excited with the most recent speed improvements for GLM5.3-Flash on DGX Spark clusters. It really feels close to what Opus ~4.5 was like to talk to. It's not quite there yet on consistency, but it's a really nice experience. Less guardrails and high quality abliterated versions further enhance its usefulness.

    Though, Qwen3.8-Flash-Next is very close to that level while requiring fewer resources to run, so I'm really looking forward to Qwen4.

  • That’s just a horrible article, very vague, wrong reasons, waste of time really.
  • GLM5.3-flash has been fantastic for me to make minor fixes in ambigious ways. "Fix x feature, whats going wrong. " It does the job.
  • Was this post generated with LLM, did he properly mention anywhere why exactly did it fail with example or i have trouble reading.
  • This wasn’t the author’s big point, but it did hit me when he said “this isn’t what we usually aspire to, but it works well for prototypes”. It captures how I feel about vibe coding - you can deal with incredibly complex tasks! But, boy, am I not about to use (say) an AI produced GPU driver on a daily basis! Perhaps we should normalize (again) “you’ll throw away your first attempt”.
  • I am excited and worried at the same time when it comes to these open weight models coming out of china..
  • > Unfortunately there are still consequences to it. I chose the 'wrong' model for the prototype, and we spent 450M tokens / $150 / 5kWh of energy use almost overnight. The MCP server itself works well and we now have a great demo of the capabilities, so it’s not for nothing:

    > Nonetheless, it’s a good reminder to be careful with model selection and with agentic patterns. We could have achieved similar results for most likely 5x less cost with not that much more effort. Lessons learned! We need to budget for this, and be more careful. Could have seen it coming, but now we know.

    I don't get it. Why was it wrong? Which one would have been better? What was the lesson and how could you have foreseen it?

  • One thing about these numbers that's absolutely shocking to me is how low the energy use is:

    > That model’s usage was well within our budget ($68, about 4kWh of energy use / 365 grams of carbon emissions).

    The energy cost is literally 1% of the total cost. For context, 4kWh of energy would drive you about 15 miles in an EV, about half of the average person's daily driving miles. It's boiling 10 gallons of water.

    With the talk of AI data centers' impact on the world, you'd think this would be 10x to 100x the amount of energy in order to get the effects they're using here.

    My takeaway: the AI data center buildout is an overbuild probably at least as large as the fiber buildout that left us with so much dark fiber. If not even bigger. The only thing that will save the economy is the inability of NVIDIA and chip fabs to produce enough chips to match the buildout planned.

Explore Birbla archives