Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • Tip: include pictures of your machine. Not sure if I’m alone in this, but I love to see other people’s setups
  • Apple did great work convincing people their unified memory was good at AI. Even AI says Apple is the best of all time at marketing.

    Meanwhile the stock market has Nvidia at the top... Until everyone gets cuda.

  • > The main reason to run local: cloud APIs are rented land. They can change their pricing, hit your usage limits, or swap the model being served behind the scenes whenever they feel like it.

    Yes but it's easy to replace them.

    The main reason should be privacy.

  • My biggest problem with running local LLMs on my M4 Max/128GB RAM is the prefill latency.

    I've since acquired two DGX Sparks, and it feels so much snappier.

  • Recent performance data on my M1 Max 32GB MacBook using oMLX. I have been working on identifying suitable model and config for my use case and system. Using a refactor and suggest improvements prompt for a specific Django code block using VSCode Cline extension.

    ---

    Qwen3.8-27B-4bit, Prompt Processing (PP) 66.3 tok/s, Token Generation (TG) 11.8 tok/s

    Ornith-1.5-35B-A3B-MLX-4bit, PP 379.7, TG 45.8

    Ornith-1.5-35B-A3B-MLX-4bit, PP 381.5, TG 46.4

    Qwen3.6-35B-A3B-mxfp4, PP 389.6, TG 47.6

    Qwen3.6-35B-A3B-OptiQ-4bit, PP 342.6, TG 44.4

    ---

    Qwen3.8-27B-4bit generally runs out of output token before completing the task though excellent partial results.

    Ornith-1.5-35B-A3B-MLX-4bit seems to get in the loop often specially with tool calls.

    Qwen3.6-35B-A3B-mxfp4 seems to be optimal with speed and quality output.

    I am going to test Qwen3.6-35B-A3B-4bit soon with same code block just to check my intuition that any derivatives don't seem to perform better than the originals.

  • Nice setup; but, for simple tasks or questions, AI is currently free? And it will probably stay free, as I don't see Google starting to charge for using AI on its search engine? So costs can't be a motivation for running small models locally?

    For more complex or important tasks, costs, autonomy and privacy matter, but then so does performance/quality.

    So I'm not completely convinced it's really worth it; but it's tempting!

  • No mention of the performance of the models? I'm able to load a bunch of different models on my little mini-PC with 16GB RAM, but the performance is terrible. I always wonder what performance people are getting with local models that they find is acceptable?
  • I like the in-depth description. Everything from the naming convention of the models (and how much RAM they require) as well as all the components needed underscores just how complicated this all still is.

    I suppose I am waiting for AI-in-a-Box to come along so I can (painlessly) join in.

    (I'm sure wrangling with all these esoteric aspects of LLMs though is fun for some people.)

Explore Birbla archives