Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • What's GPU-5?
  • by sheo
  • Absolute treasure of a website with these graphs, thank you for sharing this. Huge help for me to find a faster model (smaller quantization) for my VRAM.
  • I’m on a strix halo @ GPU-5 with MTP and I get 600 prefill and 30 TG which pushes it into a very usable range. The odd thing is that Dflash2 is really slow for me, like sub 10 TG.
  • Well I tested it on 7900XTX with the same prompts and their draft model gave me about 30t/s. Their own snippet of code with regular MTP model gave me 60t/s.

    Also model with their draft answered incorrectly. With MTP it answered correctly.

    Question was: "Does MikroTik CRS312-4C+8XG-RM have combo ports?". The answer is Yes.

  • From my own test. It's not faster than the unsloth model.

    Disclarer: I'm unsing Vulkan on an AMD GC.

Explore Birbla archives

Shapelearn Qwen 3.8 27B (13.1 GB VRAM) · Birbla