Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • 80% less for Luna is absolutely crazy, in my opinion we may reach a point in the next year where powerful models on the API could potentially be cheaper than subscriptions. Compute just keeps decreasing in price.
  • API will never be cheaper than subs because theres a ton of value created for companies by locking people into subscriptions that tend to be sticky.
  • It's a clever strategic move: grab the market of cheap low end models within the product range. It's lower friction to switch a model than a provider.
  • It says Luna is fastest, but doesn't it take way more steps to get the same job done?

    https://deepswe.datacurve.ai/ - (See the Agent Steps view)

    Or is the output speed so much higher that it cancels out?

    I don't see a lot of benchmarks that record actual time. But on AA, Sol on Low beats Luna on High for Time Per Task.

  • The type of job matters. If it is too complex then yeah, you should be using a larger model.
  • What are your use case for these? I’m manly interested in coding where more capability is better - give me a 10x model at 10x the price and I’ll take it. A worse model at very low cost has no appeal to me. At-least not for coding. Translation maybe? OCR?
  • Summarization of datasets, categorization
  • Dinky stuff I want to be fast like “update the PR with this other JIRA ticket”
  • General light/defined tasks, e.g. go fetch a list of rows from <X> and then run script <Y> against each.

    Yes, I can just do it myself, but even at API prices, I'd rather have the LLM do it.

  • Translation, moderation, classification, guardrails, etc..
  • Agentic layer.

    Your support bot.

    Your research long running bot.

    Your SEO Optimizer bot.

    Your incident analyser bot.

    Your personal assistent bot.

  • 80% price cut for luna is a very aggressive pricing move

    makes it by far the best choice for most workloads that do not need bleeding edge intelligence (reminder: luna can be comparable to opus 5!)

    by tosh
  • Is it though? I saw a noticeable difference between Sol (xhigh) and Luna (max). Sol appears to understand better my prompts, you need to be more specific/clear with Luna.
  • Luna is meant to compete with Haiku. What tasks are you seeing it equal Opus on?
  • I generally just check the Price/Performance graph on Openrouter: https://openrouter.ai/rankings#performance#benchmarks. Activate the "Show Pareto" toggle on the right.

    I was still using GLM-5.2 in my personal projects, but this just made Luna a very easy choice.

  • The official doc says, Luna = Previous Nano models, kind of. Is it really good at coding?
  • Hasn't OpenRouter had Luna and Terra on 50% off sale since they launched? I wonder what will happen to that.
  • Didn't expect that. Luna pricing is crazy now. I don't think there is anything on the market that competes at this price-performance point.

    For our production app, OpenAI clearly is the best provider now. Their API is very reliable and has many nice features. The price-performance of the model lineup is incredible. We used open weights model via Fireworks for a long time (e.g. Kimi K2.5). Fireworks is a great provider but we still ran into issues here and there (Same with Anthropic and Google). OpenAI just works, is fast and in my view has a better price-performance ratio across almost all levels of intelligence.

  • OpenAI's APIs are extremely reliable for sure. I don't even remember when the last incident or downtime was.
  • This feels like the dialup->broadband transition to me.

    I was already a huge proponent of Luna for things like deep research. Being able to run 5x more for the same cost is simply bananas. We are already running 10 parallel agents for hypothesis generation. I cannot imagine 50. The statistics become much more interesting & powerful when you can run so many samples of the exact same prompt+model without breaking the bank.

  • Very interesting. Can you share more about your hypothesis/research pipeline? I have been using Sol for those types of task because I figured you'd need more reasoning for getting good ideas, but maybe quantity > quality at a certain point?
  • How do you run 'deep research'?
  • Do you have a sense of which tasks benefit from more agents and which don't?
  • > The kernel work helped reduce the end-to-end cost of serving the model by 20%, while its experiments increased token-generation efficiency by more than 15%.

    If the cost of serving GPT-5.6 just dropped by 20%, does that add up to literally billions of dollars in savings per month?

    We know Anthropic spend $1.25 billion renting inference capacity from SpaceX (in two Colossus datacenters) from the SpaceX IPO, but we don't know how much of Anthropic's inference capacity that is (presumably a small fraction, since they were operating on top of AWS and other providers before the SpaceX deal.)

    I've not seen any numbers that hint at OpenAI's per-month inference bill, but surely that has to be in the multiple billions of dollars as well.

    So 20% is a really, really big deal.

  • imagine writing that on your resume

    > reduced inference cost by 20 percent saving company x billion dollars per month

  • Some numbers: https://www.wheresyoured.at/exclusive-openai-financials/

    If those numbers are accurate, I don't think 20% is a really, really big deal. It's like saying "we're digging our grave 20% slower." Ok, but they're still digging!

    Or, different analogy, if I'm going broke because I lost my job due to executive AI psychosis, cancelling my netflix subscription doesn't really change the math of not being able to afford rent. It doesn't even really slow it. The amount of money that OpenAI is spending is so absurd that a minor cost saving is like, uh, some progress, but they'd need to do it a lot more to move the needle

  • ~2 years ago gemini2.5 helped write better kernes for itself and (only) reached 1% efficiency gains. Today we're at 20%.
  • Making Luna, which was already very cheap and extremely capable, 5x cheaper is crazy. I use Sol at work but Luna at home, and while there's definitely a difference, it doesn't feel like night-and-day. After a year of ever-increasing prices it suddenly feels (between this, Kimi K3, GLM 5.2) that prices are falling again.
  • depends on what you are doing. if you are doing verifiable tasks like fixing bugs then any model would do as long as you write the right verification.
  • Its just a matter of time at this point.These companies are working day and night to capture the market.