Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • I did a similar thing running Q6_K model and Q8_0 DFlash2 (draft=7) quants:

        DavidAU/Qwen3.8-27B-Cold-Fusion-GAIN-V1.1-NM-DAU-NEO-MAX-MTP-GGUF
        incoai/Qwen3.8-27B-DFlash2-GGUF
    
    using llama.cpp PR/commit https://github.com/ggml-org/llama.cpp/pull/27342 on an AMD R9700 (32GB)

Explore Birbla archives

I Rented a 96 GB GPU and Took Uncensored Qwen3.8 From 44 to 125 tok/s · Birbla