Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • the fact that this author cannot get qwen3.8-27b run at the same speed as qwen3.6-27b, says the article is not worth reading. the author does not know anything about how to run local AI. 3.8 and 3.6 are the same model with different weight.

    both tg and pp speed are so terrible on author's machine.

  • The cheapest card that will run this model very well is a ln unlocked CMP 170HX. But you can run it on a 3090. I run it on an old spare A6000 Ampere. I think I wouldn’t use anything lower than 60 tok/s though, which you can get with MTP etc. I just use a full vllm stack but some people see a lot of speed with ninfer (there are non 5090 ports).

    The large RAM Macs are unusable for inference of dense models as of now. Token generation is too slow.

  • I am really surprised you’re seeing half the generation performance of 3.6 with 3.8 at the same parameter size and quantization (and same prefill performance to boot) - is there just an optimization in the stack somewhere for 3.6 that hasn’t landed yet for 3.8?

    Curious if other folks are also surprised by this apparent discrepancy or have a ready explanation.

  • Pretty impressive how the Mac ends up less expensive than the Strix Halo boxes, at least here in Ireland. A 128GB Mac Studio with an M5 Max (the Ultra can only have 96 or 256GB) still costs less than the "GMKtec EVO-X2" or the Nvidia DGX Spark with similar performance. Is it the same in the US?
  • Yea Qwen3.8 wasn't fun to use on my 64GB M4 Max either (better than these numbers though), so my new daily driver is Ornith-1.5-35B-A3B-MLX-4bit. I recommend giving that a whirl if you're on similar hardware, it's definitely better than Qwen3.6 35b-a3b which was my go-to before.

    https://huggingface.co/ornith-ai/Ornith-1.5-35B-A3B

  • I'm encountering the same behavior. I've tried 4-8bit quants and get 14-17 tok/s with one run that achieved 19. I'm eagerly awaiting dflash2 support in Unsloth or LM Studio, as allegedly that should increase throughput to around 30tok/s, which is the baseline for what I consider at least somewhat interactive.

    Jealous of the folks with 5090s running ninfer and getting >100tok/s. At those speeds it's a true frontier replacement IMO.

  • Qwen 3.8 has the same architecture and the same parameter count as Qwen 3.6. Something is not right with the GGUF if it's 2 times slower. The post says "The hybrid attention architecture is new" and says the author's older Llama build from a "couple weeks ago" failed to run Qwen 3.8 because it did not support Qwen35 architecture, but both 3.6 and 3.8 are based on Qwen35 which was released in February 2026. The post doesn't make any sense.
  • I've been thinking about buying a system to run LLMs locally but the price for one that'll run Qwen3.8-27B well is quite offputting to say the least.

    What I've been looking at instead is inference providers that use TEE and E2EE to provide cryptographic guarantees that my prompts and responses are only visible to me and the GPU itself.

    Despite their docs and assurances of what their guarantees mean, I'm having trouble getting to a point where I'm actually comfortable trusting them with secrets though. Phala for example seems to be E2EE only to the gateway and will then forward prompts to (potentially third party) providers.

    Has anyone been down this path and found a provider they feel safe with?

Explore Birbla archives

Run Qwen3.8 27B locally: real numbers from my Mac Studio · Birbla