Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • 600W * 8 just for the GPUs when maxed out (besides the cost). Def nothing for my home lab.
  • Utterly awful article. R1? Llama3.1? Not being able to serve larger llms on 8(!) RTX pro’s? You can literally run open weight SOTA models with relative ease. Even 4 GPUs get you there with a bit of elbow grease and compression. Pure slop.
  • Pass. When articles keep mentioning models like DeepSeek R1, or Llama 3.1, or Qwen3 32B, it is a pretty robust indicator of AI slop. LLMs love to suggest DeepSeek R1, etc. - training data cut-off?

    No person with real practical experience and real use cases will be using these ancient models as examples, when talking about local LLMs.

  • 4x RTX 6000 Blackwell cards is a good place to be if you can't swing 8 of them, or if you don't have the power or cooling to run that many. A system based on 4x RTX6K can run GLM 5.3 at NVFP4 precision [1] from a US-standard 120V 20A circuit when derated to 300W, and give you a better pelican than Fable 5.1 [2]. What's not to like?

    (Edit: I'm mistaken here, the pelican didn't come from Flash on 4 cards but from the full GLM 5.3 model on 8. But the Flash model is still crazy good for its size.)

    1: https://huggingface.co/local-inference-lab/GLM-5.3-NVFP4

    2: https://crimson-jeri-74.tiiny.site/

  • These people have zero idea what they're doing. Not a single mention of pipeline parallelism that would actually make the setup useful to run a big model.
  • For those who can't afford RTX 6000's you can unlock around 20% increased card to card speed on consumer GPUs using this library:

    https://github.com/aikitoria/open-gpu-kernel-modules

    The hardware supports it, but Nvidia disabled it if the driver detects cheaper cards.

  • > We currently have 14x nodes of CG480-S6053 ready to ship.

    Oh, okay, so this is an ad.

    I do still think it's well written and interesting... But if anything, it's just making me more curious about the newest generation of M5 Ultra. (and less and less interested in PCI-E Gen 5 anything)

  • Makes you realize how insane the M5 Ultra Mac Studio is. 1.2TB/s bandwidth 512GB memory. Its rated max power draw is just 480W. And it also has amazing M-series CPUs. It costs less than just one of these GPUs which each take 700W to run.

Explore Birbla archives