Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • At least I'd be in control of model quality vs. when Anthropic decides to randomly drop the quality of their offering
  • Comments are mostly showing off M5s and 5090s without addressing the article.
  • I’m running Qwen3.8 aggressive uncensored Q4_K_P on a 4090 in a loop against the 2026 CrackMe CTF challenges.

    Using oh-my-pi in a prebuilt environment that I let Qwen build too.

    Codex wouldn’t even look at the files - literally, as soon as it read something with CTF it shut down. Didn’t even offer to fall back to a dumber model.

  • Much of this is why I stick to the rule of:

    a) Don't quantize your KV cache

    b) Don't run quantizations of the LLM that are worse than the best available Q8 (the largest possible file size unsloth GGUF for a given model like qwen 3.8 27B as an example). I would rather things go slowly but I have confidence that it's doing things more accurately.

  • Even a 4-bit quant of Qwen3.8 27b is indistinguishable from Gemini 3.7 flash in our internal tests. With an RTX5090 card and ninfer, you can get ~800 TPS token generation (c=8) and ~140 Tokens per second single stream.
    by a11r
  • I saw some guy streaming how he was deploying qwen3.8 37B on his local setup. Well, he was asking Claude to do it. It took him two hours of passing errors to Claude for the endpoint to start working, he then started testing it against DS4 Flash when Qwen had thinking disabled and Claude messed up sampling parameters, it was an absolute pain to watch
  • I just got qwen 3.8 27b mlx running on my Macbook Pro and honestly I’m pretty blown away by how not-dumb it is.
  • There's quite a few tangential features that must be implemented correctly or risk affecting the LLM output in significant ways.

    Parsing/encoding is one example: A couple of months ago I've debugged a reasoning loop bug in Step 3.7 Flash on llama.cpp that was caused by the parser capturing an extra `\n` as part of a reasoning block. It was something that only manifested at longer multi-turn agentic sessions, and the extra linefeed was steering the model into making reasoning self corrections that only got worse with longer sessions (more details about this issue: https://github.com/ggml-org/llama.cpp/issues/24181#issuecomm...)

    No inference engine is perfect, but I feel that llama.cpp is the most reliable way to run language models locally.

Explore Birbla archives

Why your local LLM feels dumber than it is · Birbla