Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • Can you share more about the architectural/design tradeoffs you considered or decided upon? Particularly for me, why is a model that is intended mainly to just make tool calls and marshal the results back focusing on speed? Speed as an inherent result of small size, I get, but speed as a design focus confuses me because it’s simply not going to be dealing with large outputs as a rule, wouldn’t it be better to trade some of that raw speed for better smarts?

    For example, I mocked a dumbed down version of what would be a reasonable intermediate tool call prompt:

    > It's currently 58 degrees. User asks for house to be 8.5 degrees warmer. What temperature to set thermostat to?

    The reply?

    Reasoning: “User asks for temperature to set thermostat to 8.5 -> set_thermostat with temperature=8.5.”

    Sounds like something Siri would do!

  • Specifying units seems to be unreliable; I tried adding a description to the set_thermostat temperature:

        "temperature": {
          "type": "number",
          "description": "degrees Fahrenheit"
        },
    
    Set the living room temperature to 70 degrees Celsius

        {
          "function_calls": [
            {
              "name": "set_thermostat",
              "arguments": {
                "room": "living room",
                "temperature": 70,
                "mode": "cool"
              }
            }
          ],
          "confidence": 0.6045
        }
    
    Set the living room temperature to 70 degrees Fahrenheit

        {
          "function_calls": [
            {
              "name": "set_thermostat",
              "arguments": {
                "room": "living room",
                "temperature": 70,
                "mode": "heat"
              }
            }
          ],
          "confidence": 0.4536
        }
    
    Set the living room temperature to 70 degrees

        {
          "function_calls": [
            {
              "name": "set_thermostat",
              "arguments": {
                "room": "living room",
                "temperature": 70
              }
            }
          ],
          "confidence": 0.8517
        }
    
    Trying "in degrees Fahrenheit" for the tool description had similarly counterintuitive confidences.

    Edit: to be clear, the counterintuitive behavior is that the confidence ended up higher for the wrong units.

    by wky
  • Could someone please share how such open source micro-LLMs might have been created?

    Do the creators take something like DeepSeek, and then delete most of the neurons to whittle down the size?

  • That's really cool - I was already thinking of compressing `functiongemma-270m-it` down to 1-2 bits so it would work flawlessly in the browser. Your `Fine-tuning` feature is even much more convenient.
  • Funny result from the web demo. I'm well aware that it's an extremely small and, well, stupid, model, but even so:

    Query: HN

    Result:

    { "function_calls": [ { "name": "lock_door", "arguments": { "door": "front door" } } ], "reasoning": "User wants to lock the door. No specific door mentioned, so use 'front door' as default.", "confidence": 0 }

    I'd expect it to at least ignore (call no tools) for the queries that it doesn't understand. And it seems like it does do that, just not consistently.

  • My first query:

    > Make it a little warmer in here.

    The reply:

    > "name": "set_thermostat", > "arguments": { > "temperature": 65, > "mode": "cool", > ... > "reasoning": "'warmer' implies need for cooling; set_thermostat with temperature 65 (typical warmth) and mode 'cool'.",

    Maybe I'm doing it wrong?

  • It's definitely cool that you can get any reasoning whatsoever out of such a small model. That said, its reasoning is "interesting":

    Query: "Make the living room dark" Agent: "User wants lights on in living room. 'dark' implies dim. Room 'living room', action 'on'." (And on every test I did, it just completely ignored the "brightness" parameter)

    It also appears to have no concept of what a door or light actually is, whenever the query diverges from "Lock door X" or "Turn on light X", it tries to shoehorn whatever additional context is given into the device name:

    Query: "Lock out the vacuum salesman at the front door" Agent tries to lock "front door vacuum salesman"

    "The way you talk really makes me appreciate silence" is classified as "positive" with 82% confidence.

  • This is cool. I definitely think the "micro" sized LLM space is underappreciated, so it's always good to see work like this. I foresee a paradigm in some contexts where you have a hierarchy of LLMs, with more competent models actively training smaller models to solve specific tasks very efficiently, and something like this could be the smallest layer in that stack.

    With that being said, the web demo is not particularly impressive. It really doesn't like anything I throw at it. I'm fine with accepting that fine-tuning is the solution to this, but I wonder if there's anything to gain from a bigger model? I know it's completely counter to the whole point of this, but a 14MB binary using 28MB of RAM seems unnecessarily small and pretty arbitrary.

    Like, what does a 28MB binary get you? Or a 140MB binary? Or a 1.4MB binary? I'm guessing the choice of 14MB came from minimizing the size as much as possible while meeting certain requirements/performance expectations, but even a Pi 5 has plenty more room to spare. Curious if there's a good explanation for this (which I may have missed in my skim of the post).

Explore Birbla archives