Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • The big news is that it's going to be open weight - https://x.com/finkd/status/2095232032896946311.
  • That tweet says “Muse Spark open weights releases coming soon”, not that this specific model is going to be open weights.
  • Nice, Muse Spark is so good and keeps improving, but it's still not the best choice for any use-case. The Sol models are in their own league currently in terms of cost/speed/performance.

    Good improvements from 1.1 and 1.2[0], but when I tested 1.3 it was very slow (through openrouter).

    [0]: https://aibenchy.com/compare/meta-muse-spark-1-3-high/meta-m...

  • Cannot agree more with the Sol models. Everything else I try just seems "dumb".
  • “contributor” pricing at $0.10/$0.20 is crazy cheap if it’s measuring up to Sol.

    Definitely shows how important a user data flywheel is for RL and model improvement.

    by wxw
  • The "contributor" pricing is the standout here at a ~20x discount, if you allow training on your data.

    The model seems on par with Sol and Opus 5 on paper (admittedly on some older/saturated benchmarks, but very competitive for $).

    Stats:

    1M context, $0.10 input/$0.002 cached, $0.20 output (Mtok)

  • Not to mention, this is hyper competitive against even Chinese providers given its multi-modal support.

    Muse Spark 1.3 supports Text, Image, Video, File, Audio inputs. We've only started to see models from China include image and video inputs recently.

  • I have a feeling that Meta is not gonna like what people actually use the contributor model for lol.

    (It's probably going to be a bunch of repetitive batch jobs like web search that have no training value)

  • Used Muse Spark 1.2 and was not impressed at all. Fast and cheap but even GPT 5.6 Terra felt much more capable. Also not really looking to support a company that was just forced to pay $18B for mental health damages.
  • I'm party using 1.2 to reverse engineer and re-implement an old game binary and it has been quite good and fast. The contributor pricing is very attractive, excited to try 1.3 and see if I feel a difference. 1.2 can get stuck outputting similar sounding thought summaries with no apparent progress when asked to solve bugs. Then I've switched to GLM-5.3-Flash which for this use case has been clearly better at finding suspected causes and following tracks.
  • Practically free for "contributors" at 0.2 usd/mtok. That's going to be hard to say no to for hobbyists.
  • Price segmentation at its finest
  • they key pricing is cache reads at $0.002 per M (same as old deepseeek v4 flash prices)
  • I'm wondering whether anyone has yet extracted AWS keys from a model trained on user input. Because users are definitely feeding secrets into these "contributor" models
  • I like the approach of providing a discounted version of the API that is used to train vs. the full price version. Seems reasonable and transparent.
  • DeepSWE scores 75.4 - that's the best score so far. And it's crazy cheap! Google held the top a few hours today with Gemini 3.8 Flash, but now second to Spark 1.3. All this competition will drive prices down!
  • how much of it is from reallocation of staff to ai training and labeling
  • But is the score really reflective of the quality or are both models benchmaxxing?
    by cbg0
  • Gemini 3.8 flash has better rates. $0.75 per million input tokens and $3.75 per million output tokens.

    Compare that to Muse spark 1.3

    $1.25/M input, $4.25/M output (without data sharing) $0.10/M input, $0.20/M output (with data sharing)

    It is dirt cheap, but only if you are willing to share your data with meta and allow them to use it for improving their models and products.

  • With the contributor pricing being more than 10x cheaper than the standard, that would make it best and cheapest on the DeepSWE leaderboard! It feels fast in my experience too. LLMs keep improving at an insane pace.
  • when are we going to stop pretending these benchmarks have any meaning?

    anybody who's used these models knows that their real-world software engineering performance has no relation to the ranking on deepSWE.

  • muse-spark-1.3-contributor. Say what you want and Meta, changing the pricing to explicitly say 'we train on this and value it this much' is what every model provider should do. As a side note, it is now completely obvious how much stealing my tokens for training is worth to model providers. I avoid/pay extra/try my best to make sure I am not getting trained on but it seems like it keeps popping up that I missed a setting somewhere. This is the first quantifiable number I have seen out there from a model provider. Maybe it can help in lawsuits to quantify the damages for copyright/other things?
  • This is also really smart business wise imo. For hobby projects, toys, quick scripts you don't really mind if they train on it. It's a win-win. Once you get used to the tools and you want to do more serious business you are more likely to buy a more expensive sub from them.
  • We already had a good idea of how valuable it is from how much X.ai acquired Cursor for, and the near-immediate improvements to their coding scores.
    by cma
  • This has been my hunch for a while about all the discourse of "OpenAI/Anthropic subscription pricing is unsustainable!!"

    We understand theoretically they're taking our data, but yeah, that data is vital to the entire business plan of all these companies and WAY more valuable than people are giving credit for.

    I checked up on Mistral recently and saw their Claude-alike coding harness is using GLM now, whatever it takes to keep users on their platform and feeding them data.

  • A model that (at least in benchmarks) is getting closer to SOTA. A clear separation between what’s used to improve their products and what’s not (at least this is what they claim).

    Good job Meta! Seriously. This is almost making me forget about the 18B$ lawsuit for children social media addiction.

  • How is it not SOTA? It's beating 5.6 Sol.
    by dbbk
  • I started using Spark 1.2 for development because if you're willing to let Meta train on your data it was dirt cheap and was actually really pleasantly surprised with it. It's not a frontier model by any means, but for work that didn't require a top of the line model, I really enjoyed using it.

    I'm anthropomorphizing it a bit, but it felt like it knew its weaknesses and didn't try to impose it's opinions on me. What I mean by that is that it did what I told it and if there was something unexpected in the code that it put out it was often because I gave it ambiguous or conflicting instructions. It didn't try to go above and beyond and just acted like a tool, which is what I want from a coding agent 90%+ of the time. I also felt that it did a much better job of following established patterns in my code than many of the other current models do. I'm a huge fan of OpenAI's models and Spark 1.2 is what I expected 5.6 Luna to be.

    I'm curious and a little excited to use 1.3, but honestly a little worried that as Meta pushes for better benchmarks that Spark will start to fall into the trap of trying to be "helpful" in ways I don't want it to be.

    Tangential, but when I first started using Spark 1.2, it made me realize how much I miss 5.3 Codex. That model was the peak of coding models, IMO, in that it knew how to write good code, but didn't try to overstep or be "helpful" in unexpected ways. That got me thinking about how the major labs seem to be stepping away from coding focused models toward more general purpose ones and how I can't help but feel like that's a mistake.

  • If it's a mistake, it should course-correct.

    I agree that some of the smarter models are actually worse. I hope they take a model that's good enough--there are many--and just try to get it chatjimmy.ai speed.

    I have to think that's the future, somehow, and I'm really excited about it.

  • I believe that you can still use 5.3 Codex in the eponym CLI tool, the "Spark" fast version. I hope that it will lighten your day! :-)