Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • by tosh
  • It's strange that they don't include reasoning training (RLVR). Their justification doesn't sound convincing:

    > While reasoning models have grown in popularity in recent years, their abilities aren’t always the most efficient way to get a result. In enterprise settings, token costs and speed are often as important as performance. That is why turning to less expensive, non-reasoning models with similar benchmark performance for select tasks like instruction following and tool calling makes sense for enterprise users.

    I guess they currently don't have the ability to do proper RLVR.

  • I may have misunderstood: is not reasoning training (RLVR) independent from the use of the "<think>" tags - is it not a method that improves results in reasoning? How do we know that it was not carried out?

    Incidentally: I am trying to spend some time researching in the progresses in the area (the jump from parroting, to inconsistent apparent reasoning, to reliable reasoning).

  • The Granite 4.1 3B model is only 2GB from Unsloth: https://huggingface.co/unsloth/granite-4.1-3b-GGUF

    I ran it in LM Studio and got a pleasingly abstract pelican on a bicycle (genuinely not bad for a tiny 3B model - it can at least output valid SVG): https://gist.github.com/simonw/5f2df6093885a04c9573cf5756d34...

  • Do you have any reasons to believe that granite is more immune to the effects of quantization than other tiny models? Otherwise it seems odd to judge a tiny model true capabilities by using its 4bit quant.
  • Very impressive series of SLM by IBM here.

    I have been using it with their Chunkless RAG concept and it is fitting very well! (for curious https://github.com/scub-france/Docling-Studio)

    I convinced that SLM are a real parto of solution for true integrated AI in process...

  • Nah, I ain't reading that. If they can't be bothered to get a human to write it, it can't be that important. I'm glad for them though. Or sorry that happened.
  • This is the official announcement: https://research.ibm.com/blog/granite-4-1-ai-foundation-mode...

    It is not the researchers' fault that some slop got posted here instead.

  • > Full stop.

    Why people don't edit out obvious sloppification and expect to still have readers left

  • Are you referring to the literal use of the expression "full stop"? I don't see it anymore in the article, maybe they edited it out?
  • So are we saying it's fine that the article is written by an LLM as long as it doesn't have the tell-tale signs of LLMs?
    by cbg0
  • Third line in to the article: "But there’s one result in the benchmarks I keep coming back to."

    I hear this sort of thing all the time now on YouTube from media/news personalities:

    “And that’s the part nobody seems to be talking about.”

    "And here's what keeps me up at night."

    “This is where the story gets complicated.”

    “Here’s the piece that doesn’t quite fit.”

    “And this is where the usual explanation starts to break down.”

    “Here’s what I can’t stop thinking about.”

    “The part that should worry us is not the obvious one.”

    “And that’s where the real problem begins.”

    “But the more interesting question is the one no one is asking.”

    “And this is where things stop being simple.”

    It doesn't really worry me but I think its interesting that LLM speak sounds so distinctive, and how willing these media personalities are to be so obvious in reading out on TV what the LLM spat out.

    I've never studied what LLMs say in depth is it is interesting that my brain recognises the speech pattern so easily.

  • People complain a lot about LLM-written articles, but the human comments here on HN are far worse. Mostly a bunch of people extremely proud of themselves for not reading an LLM-written article, and then a bunch of people who take it at face value and make the model seem almost useful, and one comment that actually looked at other benchmarks. Good 'ol humanity, good at.. being emotional... and not doing analysis.....

    The article makes some good points about model design (how different size models within a family can get similar results, how to filter out hallucination, math result reinforcement), so that's worth understanding. It's analyzing a paper, which only discussed 3 sizes of the same model family. But what the article doesn't say is, compared to other model families, Granite 4.1 8B sucks. The only benchmark it does well at compared to other models is non-hallucination and instruction following. Qwen 3.5 4B (among other models) easily outclass it on every other metric.

    This article teaches a valuable lesson about reading articles in general. You can take useful information away from them (yes, despite being written by LLM). But you should also use critical thinking skills and be proactive to see if the article missed anything you might find relevant.

  • > the human comments here on HN are far worse

    I already assume some comments here are LLM written.

  • > But what the article doesn't say is, compared to other model families, Granite 4.1 8B sucks.

    Right. This just says that Granite 4.1 8B is better than a previous version, Granite 4.0-H-Small, which has 32B, 9B active.

    So, they made a less bad model than before. But that doesn't tell you anything about how it compares with other models.

  • >Mostly a bunch of people extremely proud of themselves for not reading an LLM-written article

    I'm not sure it's proud as much as people voicing displeasure with the uncertainty about what went into the LLM prompt. This may have been a 1 sentence prompt, or it may have been some well researched background that simply reformatted it. Why waste minutes-hours on verifying it if it's possible someone could have spent 10 second on it? It's very easy to see their point.

    People seem to indicate people they disagree with voicing their opinion about anything lately is some auto-fellatio, I wonder what causes them to think this way.

  • "The article makes some good points about model design"

    But how can I tell if those are good points or not?

    I don't want to invest time in reading something if the presence of those "good points" depends on a roll of the dice.

  • > people complain a lot about LLM-written articles, but the human comments here on HN are far worse.

    No, they aren't.

    You are comparing writing produced with little to no effort to writing produced with the minimal effort required to communicate.

    It's reasonable for people to complain that they are presented material that not even the author thought was worth the effort.

  • >> The only benchmark it does well at compared to other models is non-hallucination and instruction following.

    I think instruction following is going to be the most useful thing these models do. Add a voice interface and access to a bunch of simple, straight-forward devices or APIs and you have a mildly useful assistant. If that can be done in 8B parameters it will soon run on edge devices. That's solid usefulness.

  • The problem is the signal/noise ratio in these articles. If the AI has written the article, then this same info could have been generated by my own AI, but tailored to my needs. So what, exactly, is the new info that this article is generating that I can use to consult with my AI? That's what I want to get out of this interaction.

    Maybe my point is something on the lines of "Just send me the prompt"[0]

    [0] https://blog.gpkb.org/posts/just-send-me-the-prompt/

  • The pro LLM rant is weird, LLMs "hallucinate" in creating detailed elaborate lies, the frontier models still do this egregiously, an LLM written article by default has 0 value since every single line could be true or it could be a convincingly crafted lie, every line has to be fact checked

    I'm using Gemini 3.1 pro to help me research my thesis, it still with search enabled and on pro mode, invents entire papers that don't exist, and lies about the contents of existing papers to relate them to the context or to appease me, if I submitted an LLM written article based on the results its given me 80% of the article would be lies

    Commenting to complain that the article is LLM written is helpful too since some people aren't able to distinguish

  • On the topic of local models, is there a good equivalent to something like Claude's chat interface? I've recently started transitioning to open models after getting fed up with Claude's usage limits (I'm not in a position to drop $200/month), and for coding tasks Kimi 2.6 has been about the same as Sonnet in my experience. The only thing I've found myself missing is a nice interface to ask it questions and have it help me with my math assignments.
  • I've been mostly using LM Studio for this recently. Ollama has an OK chat UI now too. 'brew install llama.cpp' gets you 'llama-server' which provides quite a good web UI.
  • I re-created Claude's interface closely here, feel free to fork https://github.com/mudkipdev/chat
  • You can try Open WebUI. Its genuinely useful when it comes to running open models locally with a clean interface
  • Ollama does this, as does llama-server from llama.cpp
  • Open WebUI or Jan (https://www.jan.ai/). Work well with Ollama.
  • With Ollama* you can use Claude Code with `ollama launch claude`

    * https://docs.ollama.com/integrations/claude-code

  • Most of the common ways to run local LLMs include a chat interface. llama.cpp's `llama-server` stands up a chat interface on 8080, as well as an OpenAI compatible API. LM Studio is a desktop app with a chat interface and API, as well. unsloth Studio, too.

    LM Studio is nice in that it makes it easy to add tools, like search. Qwen 3.6 is such a small model that it lacks a lot of knowledge of the world (so it can hallucinate at an uncomfortable rate, which is a common failure mode of very small models), but it can use tools, so being able to search lets it research before answering. It has pretty good reasoning and tool calling, so it's actually pretty effective. I've been comparing Gemma 4 (31B at 8-bits, also very good with tools and reasoning for its size), Qwen 3.6 (27B at 8-bits), against Claude Opus and Gemini Pro lately. And, obviously the frontiers are better, but most of the time, I find the tiny models are fine. I'm still not quite at the point where I'd be willing to code with local models, as the time wasted on hallucinations and logic bugs and sloppy coding practices are much higher, as is the cost of security bugs that make it past review.

  • Yes but not exactly.

    - A lot of people suggesting llama-server's web ui, but that requires you use local AI (llama.cpp), it's persisting content into your browser rather than the server (so you can lose your chats), and it doesn't support much functionality.

    - There are some pure-browser chat interfaces that are like llama-server but you can use remote LLMs. This is closer to what you want, but everything is stored in the browser, so backup is harder.

    - There's LocalAI, which is like the llama-server option, but more stuff is built in and it persists data to disk. It's flashy and very easy if all you want to do is local AI.

    - There's LM Studio, which is another thing like LocalAI, but a desktop app.

    - There's OpenWebUI, where it's like LocalAI, except you don't do local inference, you use remote LLMs. It sucks to be honest, just stops working a lot of the time, UX is terrible, lots of weird bugs.

    - There's OpenHands, which is more like Codex/Claude Code web UI. You run it locally and connect to remote LLMs. Kinda clunky, limited, poor design. Like most coding agents, it doesn't support all the features you would want, like LocalAI/OpenWebUI do.

    - There's OpenCode's web UI, which is like OpenHands, but less crappy.

    - There's Jan, which is probably what you want. It's a desktop app rather than a web UI.