Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • Why is there still a hard VRAM requirement that's dependent on the model size? Isn't that exactly what this project is supposed to solve?
  • Looks cool, I'll try this when I get home.

    I have a couple of comments about https://trysoup.dev

    > Get Started for Free

    Does this mean that this will not be free at some point?

    The website is difficult to read (gray on black doesn't work well for me).

  • How much data do you need to fine tune a model?
  • This seems really interesting - I was curious about this line from the website.

    “The whole post-training stack in one CLI. Soup doctors your data pre-flight, picks the method, writes the config, derives evals from your own data, gates every save, and self-corrects reward hacking mid-run instead of just halting.”

    How does soup auto tune the hyper parameters and make some of these more complex training decisions?

  • There are some samples of training data in this folder - https://github.com/MakazhanAlpamys/Soup/tree/main/examples/d...

    They're all very short though. Anyone got a good rule for how much data of this nature is needed to successfully fine-tune a model of this size?

  • I run a fine-tuned 4B for AML compliance at community banks — the ROI math is exactly this
  • Tangential/meta: Holy shit, I've never seen a thread where almost half the comments are dead (and LLM written), especially for a post that's (currently) at 86 points and 20 comments (4x ratio is "pretty good quality" post signal generally for me).
  • Small open weight local models are the future.

    While hosted mega models make headlines for doing cool stuff, the vast majority of applications for AI simply don't need all that power, and thus cost. That’s a big part of why businesses are screaming that there’s no ROI from AI.

    Brining this tech down into small local models is likely where this all converges for the vast majority of use cases and what solves the present ROI crisis for LLM-based AI.

Explore Birbla archives