

Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- This is a joke, right?by GolfPopper
- Not necessarily. There were some tests last year-ish from hf that showed that simply alternating (randomly) between claude and gpt (whatever their versions were at the time) on a task produced better results than either of them individually. So during a task, the first call was sent to one, then the other and so on.
There's also the concept of "smart routing" requests based on some heuristics / embeddings. You'd get "simple" tasks handled by smaller (cheaper) models and use a bigger model to curate / sort / merge the results.
There's a lot of things to try here. I wouldn't personally pay for this service, but I don't think it's "a joke"...
- Reminds me of <https://github.com/irthomasthomas/llm-consortium>by eevmanu
- Fugu Ultra <https://console.sakana.ai/models#fugu-ultra> sounds similar to GPT-5.5 Pro or Gemini 3.1 Deep Think .
Is there any official source that could confirms if Fable (or Mythos) is parallelized test-time compute (like GPT 5.5 Pro) or sparse Mixture-of-Experts (MoE) transformer combined with a multi-agent, inference-time compute scaling architecture (Gemini 3.1 Deep Think)?
by eevmanu - ngl, I thought sakana.ai was doing cooler stuff than this. that said, the release of a product like this makes sense because it follows your natural intuition when using these models. The best way to use LLMs is to have at least two in your pocket, because the models do a good job at covering each others assets and filling in obvious model-specific blindspots.
it's interesting that they're offering in the form of fixed cost subscription plans too. My impression was that the first party providers can do this because they api inference margins to the tune of 80ish percent. Anyone else orchestrating on top of these models have to pass through these costs or eat it themselves.
by prodigycorp - Got myself the $20 subscription and tried it out. The 5-hour limit runs out surprisingly fast. Quality is okay but it feels slow, and even with my $20 Claude subscription on Fable, the credit usage ends up being lower. Fable usually catches issues in my Opus 4.8-generated code that I'd miss otherwise, but Fugu didn't. Makes me wonder if it's really at the Fable level. Hard to see the value here.by Lwrless
- > Frontier-level performance without single-vendor dependency. [...] Plug collective intelligence directly into your workflows today with a single API.
Does multiple vendors run this "single API" or how is this not replacing a single-vendor dependency for another single-vendor dependency?
- Beta user: they piloted OpenRouter fusion before it was seen as the viable step. Everyone's understood for months now that having different models check each other is the best path forward.
This gets you that in a nice neat package, without the underlying tinkering mechanics.
If (big iff) the usage mechanics work out, then this is actually a really good anti-big-model strategy.
They'll be incentivized for your success, not token-maximizing for their investors.
The team is super smart too. What's not to like?
Wishing them the best on launch.
by epsteingpt - if you've used codex or claude, how do the usage limits on fugu feel compared to the pro plans on either? honestly wouldn't mind subscribing to this if it's as generous as what codex is giving me monthly, which seems unrealistic.by prodigycorp
- FWIW, I've just published this interview with David Ha an hour ago.
We talk in detail about Fugu and why these kind of routing models are likely to win out over the big frontier models.
He makes some very good arguments for them.
https://www.disruptingjapan.com/the-future-of-ai-looks-very-...
by timoth3y - I tried running this for some market research for my startup and it did a pretty nice job. It didn't necessarily find any obscure data, and it seemed to rely on older data than what I could find myself. On top of this, it had the same sycophantic tendencies as most LLMs these days (explaining why your idea is great and riffing on that), which I find to be unnecessary use of resources.
All put together, paying ~$60 to get a hit-or-miss report seems a bit excessive, but obviously as the models they use under the hood get better it becomes more and more worth it, assuming they also improve their grounding/search capabilities.
I'm a big fan of Sakana though, and have followed David Ha / @hardmaru since the world models papers (with the racing car game and the Doom clone), which were incredible at the time.
by blixt - Looking at the technical report I'm a bit confused. The improvement from using their orchestrator models seems minimal (in some cases lower than just the model which I'm assuming is in the orchestrator's pool?). Maybe it's sort of acting as an additional reasoning step upfront? Sort of like how if you asked Claude to create a plan for how best to prompt itself, you would probably end up with a better result than just the base prompt.
Also, from the technical report, looks like they're training on the output of Claude Code, etc. I'm guessing this doesn't violate TOS because they're technically not a directly competing model. This brings me to what I see as the main risk with this service, which is that it seems like an easy thing for a frontier lab to make obsolete, either by models beginning to converge in terms of strengths or by improving their own harnesses to include more of this meta-reasoning.
- As a developer outside the US I think it's vital to have alternatives to OpenAI and Anthropic, but sadly this is not it. For $200/month you get < 3 hours of use per week, the API is extremely slow, and the output quality in my tests is nowhere near Fable. It's nowhere remotely near usable as a day-to-day workhorse. Very disappointing.by cortesi
- I'm glad eager people like you test for lazy people like meby NetOpWibby
- u seem to be the only one who used it here - how did it compare to opus and gpt5.5? in theory it should be at least on par if not better at times right.by itemize123
- There are so many derisive comments here.
David Ha, CEO and co-founder, was one of the youngest managing director at Goldman Sachs before doing ML at Google. His ML publications were considered top-notch almost a decade ago. I had high hopes for him when he raised money and founded Sakana.
I do agree with some comments here that perhaps this particular product is not well thought out. I also agree with the criticism that David calls Sakana a frontier AI lab while making money just selling AI B2B applications to Japanese businesses. I also agree with the assessment that Sakana has abrasive and antagonistic, sometimes openly hostile, recruiting tactics. I also agree that his then-impressive publications may have lost their luster in the age of LLMs.
However, the man is clearly driven; and he and his team may have more to offer in future. I admire the man for not taking the conventional AI-research career path.
by quanto - Kind of shocking - a model comes out that beats mythos and offers a reasonable price and it ... gets downvoted?
Probably taking hate from both sides - OpenAI / Claude fans who are undercutting its moat. Chinese open-model fans that want it to be cheaper.
But it's a genuine accomplishment to hit those benchmarks and offer a reasonable plan?
Bizarre reaction TBH.
by epsteingpt - so he's the quintessential brilliant jerk, okby hsaliak
- IQT does not invest in normal AI B2B company.by gooddata
- Indeed. The world models research many labs are now chasing was to some degree ignited by David Ha and Schmidhuber's 2018 paper.
More broadly, Sakana is pursing a refreshingly distinct research path, with their focus on evolutionary methods, biological intelligence (e.g. continuous thought machines) and open publication.
by ainch - You pay $200/month to Anthropic, $200/month to OpenAI, $200/month to Cursor, $200/month to $200/month to Google, and seeing that it didn't come to a nice round $1024/month, you pay $200/month to Sakana to coordinate it all, because why not.
While you're at it, feel free to send me $200 as well, I'll generate a crypto address ending with "AI".
by holistio - Happy user here, pairing it with Composer 2.5, with Fugu Ultra as advisor and Fugur as planner. For scope/architecture it’s on par with useful Fable-style orchestration than one chat thread.
I've been shipping production on archive.tw with Fugu Ultra in /advisor on oh-my-pi.
Advisor doesn’t slow the loop if the driver stays fast. Worth it if your harness can split advisor from worker.
by audreyt - Or use openrouter and switch to model you want to use..(i think so)by someone_1234
- Does it work? I’m less interested in economics than fit with an MVP.
- I wish I only paid $200/mo for Anthropic! Multiply that by 20x.by bicx
- Pay $0 to run a local model or even a cheap DeepSeek V4 model via their API which is close to free per million tokens.
These prices are just going to get raced to $0.
by rvz - at this point I might just try Neuralwatt and see how much request I can get with GLM5.2. I've read a lot of reviews that its very cheap to run using Neuralwatt cloudby robertwt7
- TIL: I just found out that base58 disallows I (capital i), l (lowercase L), O (capital o) and 0 (zero), so I could only generate GrxoJt4eNXE2QaQ55iPSa7hhiYdzCo8ZeAuokmh2Cai.
(don't send anything, sharing only because of the base58 fun fact I didn't know)
by holistio - My current setup:
Opus at low/medium effort generates plans. Then several coordinator/worker pairs are possible: DeepSeek v4 Pro + Minimax M3, Mimo v2.5 Pro + Mimo v2.5, Mimo + Minimax, Sonnet 4.6 + Haiku. I've been running hundreds of long multi-agent sessions, topped up extra credits here and theere, but haven't reached $200/month spend yet. Relying entirely on Claude/Codex feels like a waste of cash now.$20/month: Claude Code $10/month: Minimax $16/month: Xiaomi Mimo $10/month: Opencode Goby ricardobeat