Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • how much energy does it consume?
  • Good one! I haven't measured this. I'll include it!
  • For what it's worth, I've spent some time with Claude to develop a local runner for `llama.cpp`. I run Qwen3.6-35B-A3B-MTP (fast!) and Qwen3.8-27B (20 tokens/s).

    This was definitely worth the effort. Benchmarking and testing various approaches and various options really paid off. For example, one thing that surprised me was that MTP made things slower, not faster for Qwen3.8-27B.

    I use a 64GB MacBook Pro (M4 Max).

    by jwr
  • I find mtp=3 does well with that model, only at 4 it becomes unprofitable.

    Check your quants, its worth having the mtp layer be a bigger quant if it leads to 2x throughput from more accepted tokens.

  • I’m not an expert, but my understanding is that MTPs are smaller LLMs fine-tuned to "mimic" / predict a specific model’s response. It’s possible that the MTP you’re using isn’t trained well enough on Qwen 3.8. What accept rate are you getting?
  • I tried this in my mac mini m2 16GB, unfortunately I have to use an usb disk for the model weights, and I’m getting 0.5 tok/s. Still, being able to run (heh maybe crawl is more accurate) a 100B model on this computer AT ALL is pretty cool.

    I see disk maxing out at 400 MB/s, this disk should be able to hit 1GB/s (it hits that eg when verifying the check sum of the weights), so there might be some optimization to be done there (I’m guessing it’s because the weights access is not pure sequential reads but involves some randomness depending on which expert)

  • Brilliant! Have you had any success integrating it into a MacOS Swift app. I'd love to see it in action before I consider adding it to a future build. My biggest issue is getting these opensource models to use tools well enough for production.
  • This is next! in the works rn.
  • As someone who is just looking at the theoretical benchmarks of each of these models I'm curious if anyone could share what are the problems (maybe around code) that flash-next was able to solve which 27b was not able to
  • This is the best I could find: https://huggingface.co/Qwen/Qwen3.8-Flash-Next?utm_source=ch...

    About the specifics, I have only anecdotal evidence, but I guess this info can be found somewhere

  • For a local non-coding agent, instruction following and tool use are the most important gains
  • I love these efforts to get proper models running on lower cost hardware and I think this is where the next real breakthrough will come from. The more efficient this sort of thing can be done the bigger the chance to democratize this tech, 'good enough' is what you need and as long 'top of the line' gives a competitive edge even if it is at a cost there is a substantial risk of the door closing on general computing at some point in the near future. Keep in mind that there is no guarantee that the pendulum has to swing back, it can swing one way and get stuck, and then you're going to have to beg for crumbs from the haves.
  • I think there's a very good chance that history will rhyme a bit.

    DOS/Windows and PC clones were by no means the best available, but they were cheap, ubiquitous, and versatile compared to alternatives that were either much better at one task but more expensive or better at everything but wildly expensive. They were "good enough" and represented a solid improvement over what many existing computer users had as well as a good entry point for new users. As such they spread like wildfire and became the standard while the expensive alternatives either became hardcore niche or vanished.

  • I am trying to make it easier to use LLMs on older, cheaper, smaller GPUs. I'm taking a similar approach (move MoE expert weights to disk, avoid wasting VRAM on these). My goal is also to run models that do not fit. My work also suffers from AI documentation issues. Where my approach differs is that instead of running an LLM that doesn't fit slowly, run many agents in parallel sharing the streams of MoE experts weights, to increase throughput. I envision a team of AI agents sharing a pretty-good-at-coding LLM that does not fit to collaborate on a set of related features, being developed in parallel.

    My work is showing promising results (if you can get past the way the AI tries to describe what I am doing). https://sw-ml-study.github.io/emufpga/index.html

    I am doing this work initially on a 6-Xeon-cores Linux workstation with an RTX5060-16G to run MoE models larger than that. Then I will be moving this to a server with a lot more cores (Dual 32-cores) and a mix of SAS HD and SSD drives, using older GPUs.

    Ultimately, I hope to build some FPGA/MCU "accelerators" that process the expert weights on systems with not enough CPU cores to offload the experts. If I can enable large capable models to run on older hardware, keeping the limited GPU VRAM for context and things that must be in VRAM, I can get useful work out of my old refurbished systems without paying today's RAM and VRAM/GPU prices.

  • "Disk is the gate that bites first"

    AI;DR

  • The never ending gate bites.

    How I have come to detest certain phrases.

  • Ha! AI;DR is a great phrase. Had not seen that before
  • Not a mac/UMA discussion point, but is it time to add additional, installable, DDR5 to GPUs? I can see this as a win/loose. PCIe 5x16 is close to maxing out the bandwidth available from high end dual channel DDR5 now, but not quite. I'm not a hardware person but I suspect putting it on the card could lead to significant performance improvements over using system ram so allowing systems like this, where MOE weights are shed, to get even higher performance than just adding that DDR5 to the system. Bigger models become closer to reality and it provides more of a pathway for developing technologies that take advantage of it. Of course the loose side is that you just put a lot of specialized ram on a card instead of into the system where it could be used for other things. I could see a place for a 16GB card with 64GB(or more) of DDR5 especially if we start seeing MOE and similar technologies really start being designed for this concept.
  • Probably not with DIMM modules, as longer traces mean higher latency (speed of light is ~30 cm in 1ns). GDDR typically uses larger buses (more wires) for higher bandwidth, even more so for HBM, so DIMM would be hard. Maybe CAMM would be up to the task?

    It certainly seems feasible from an engineering perspective (though it does make cooling harder), at least for mid-range, not H100-class HW, but it prevents market segmentation, so EOMs may not be too interested (as long as no competitor does it).

  • I'm hoping to see progress in this space.

    Folks talking about how 32G is not enough for local use, but then there's been work like this to empower it.

    My hope is that the new 32G M6 will be "useful" locally, possibly because of work like this.

  • yes! I'm bullish on this. there is a lot of work to do. I've been experimenting with pruning, distillation, and retraining too. I'm sure your 32gb m6 will run a badass local model!
  • 32GB is simply too tight; you need 8 minimum for the OS and you need about 4-8 more for the LLM you’re visiting and kv cache.
  • Yes, but also 12 tok/s versus Claude is so far from comparable. I know that it’s not exactly 1:1, but it’s a long way from an easy trade-off, especially considering hardware prices for high levels of RAM.
    by tyre
  • It's hard to believe 16GB unified memory will give you 5 tok/sec unless you are ignoring the thermal warnings. I am running Qwen3.6-35B-A3B on my 16GB M3 and get 7-8 tokens/sec with all the optimizations while keeping the peak memory and thermal warnings at check. https://github.com/deepanwadhwa/samosa-chat
  • interesting! Yes, thermal is important. Pretty cool project man! Starred and checking it out!
  • Anything smaller than a 16” runs into serious thermal problems; even an identically equipped 14” just can’t dissipate enough heat.
  • The laptops definitely can't hang but the minis don't really care. I threw mine down in the basement just to put the heat somewhere else, can tell when the dehumidifer next to it is on because it's a few C lower but that has no impact on performance. I don't think it's ever seen anything north of 70
  • Now I'm feeling pretty good about getting 10-11 tokens/sec running Qwopus 3.6-35B-A3B Q6_K on an old Mac Pro 2013 (trashcan) with 128GB RAM (DDR3), 12 core Xeon, dual D700s. Arch Linux and llama.cpp.
  • I have a 48GB M5. I don't need to run larger models. I want more context. I've managed to set the context window at 71,680 using Qwen3.8-27B-oQ4e-fp16-mtp. But I want more. Is anybody, with similar specs, able to set their context window higher?
  • yes, with qwen3.8-27b-4bit run via rapid-mlx i can get to about 200k