Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- Now if only I could afford 8 RTX PRO 6000'sby robotnikman
- Start with 4, see my other comment. The recent GLM, Qwen and DeepSeek releases are amazingly promising.by CamperBob2
- > (~9.2 million tokens node-wide at 4k context).
stopped reading after that. What 4k context would be usable for?
by varispeed - Not long-horizon coding but for a lot of other things like batch processes with structured outputs, quick checks/fixes, making sense of unstructured data etc..by erdaltoprak
- So much more than you'd realize!
That's a solid 2200 words to spend on operating parameters and conveying state, leaving a generous 700 word window for them to decide and respond in.
When the bonsai/prism 1bit models dropped and I saw how many prompts a minute I could get from a dusty m2 mini I started hooking it up to all sorts of shit, like a traffic simulator that translates the car state/surroundings/immediate goal into text, it responds with a seqeunce of actions defined in the system prompt, which then get translated back into NPC input.
What I was hoping for here was that it would result in fucking chaos, all sorts of stupid decisions and epic car accidents. I cannot overstate my disappointment (and terror) when they were perfectly reasonable, safe drivers. I had to cut the tire grip by 75% without telling them and make them control twice as many cars to delay their ability to respond before I saw anything resembling an enjoyable traffic accident.
- SLI is relevant againby xyst
- did you try p2p enabled driver and proper nccl env vars ?by Avlin67
- 600W * 8 just for the GPUs when maxed out (besides the cost). Def nothing for my home lab.by christkv
- I'm curious whether actual inference workloads actually push to 600W (and not 350W) and what the last 250W get you. Rare is the (generic gpu) workload where I get >5%, some rare light inference benchmarks up to 10%...by touisteur
- Utterly awful article. R1? Llama3.1? Not being able to serve larger llms on 8(!) RTX pro’s? You can literally run open weight SOTA models with relative ease. Even 4 GPUs get you there with a bit of elbow grease and compression. Pure slop.by proxysna
- Pass. When articles keep mentioning models like DeepSeek R1, or Llama 3.1, or Qwen3 32B, it is a pretty robust indicator of AI slop. LLMs love to suggest DeepSeek R1, etc. - training data cut-off?
No person with real practical experience and real use cases will be using these ancient models as examples, when talking about local LLMs.
by kmike84 - 4x RTX 6000 Blackwell cards is a good place to be if you can't swing 8 of them, or if you don't have the power or cooling to run that many. A system based on 4x RTX6K can run GLM 5.3 at NVFP4 precision [1] from a US-standard 120V 20A circuit when derated to 300W, and give you a better pelican than Fable 5.1 [2]. What's not to like?
(Edit: I'm mistaken here, the pelican didn't come from Flash on 4 cards but from the full GLM 5.3 model on 8. But the Flash model is still crazy good for its size.)
by CamperBob2 - > A system based on 4x RTX6K can run GLM 5.3 at NVFP4 precision
It actually runs fine at FP8 on this hardware too, with the full 1M context.
by nojs - These people have zero idea what they're doing. Not a single mention of pipeline parallelism that would actually make the setup useful to run a big model.by jimmoores
- I feel like you would want to run 8 smaller models separately for quantity of raw output. 1 big model is slow and isnt guaranteed to make no mistakes.by kelmoran
- can you point to a write up that discusses what you're talking about?
because I would read it.
by schaefer - I can’t stand it. Very engineering-y over specified formal language around a complete lack of core understanding. Is damaging other people read this and try to learn things from it.by itkovian_
- For those who can't afford RTX 6000's you can unlock around 20% increased card to card speed on consumer GPUs using this library:
https://github.com/aikitoria/open-gpu-kernel-modules
The hardware supports it, but Nvidia disabled it if the driver detects cheaper cards.
by RachelF - > We currently have 14x nodes of CG480-S6053 ready to ship.
Oh, okay, so this is an ad.
I do still think it's well written and interesting... But if anything, it's just making me more curious about the newest generation of M5 Ultra. (and less and less interested in PCI-E Gen 5 anything)
by schaefer - You must be the only one 'round here without a stack of RTX PRO 6000s, 8 high, that you're not sure how to use.by baron3dl
- Makes you realize how insane the M5 Ultra Mac Studio is. 1.2TB/s bandwidth 512GB memory. Its rated max power draw is just 480W. And it also has amazing M-series CPUs. It costs less than just one of these GPUs which each take 700W to run.by srcreigh
- *TB/sby qeternity
- If you want to see how impressive Nvidia is, serve 32 concurrent request on it and compare the same with the mac.
No comparison.
None.
by segmondy - These GPUs are extremely inflated in price because Nvidia effectively has a monopoly on hardware that is used to train models. Apple Silicon tends to have good inference software available but as soon as you want to train even a YOLO model bits and pieces fall back to software implementations. Try to train an LLM and it'll get even worse.
The M3 Ultra's GPU performance is around a 4070 Ti. The M5 Ultra more like a 5080. They're both amazing deals compared to Nvidia for local inference because of their massive pool of high bandwidth memory. But a single RTX PRO 6000 should be 2 or 3x the compute of an M5 Ultra.