

Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- I'd just like to point out that the largest model Cerebras has ever served is Kimi K2.6 which is 1T parameters, so that either means that theyve had a breakthrough on the hardware engineering side of things, or GPT-5.6 Sol is likely a lot smaller than people think.
If it truly is only ~1-2T parameters, then this kinda kills 2 narratives for me.
1. all the handwringing about open source catching up via Kimi K3 (3T params) is complete nonsense. All that matters imo for determining which labs are leading is intelligence per parameter. Anyone with a enough compute can train a giant model, but being able to squeeze capabilities into smaller models gives you a massive inference and training edge.
2. Inference margins are clearly insane, and this explains why OpenAI was able to lower the price of Luna by 80%. Id guess that thing is probably 120b params based on the TPS they are serving it at.
by anthonypasq - Isn’t the fact Fable is more expensive than Sol-Max by multiples already an indication that Sol is way smaller?by manmal
- Would be extremely interesting if some of the closed models would be that small. Means maybe in future they could run locally.by Gecko4072
- Don't frontier labs distill their own bigger models into smaller ones? Opus 5 was probably distilled from Fable/Mythos. Since Chinese labs now have competitive models, they can distill those into their smaller version in Kimi K3.1 or something and achieve better intelligence per parameter results.by literallywho
- Whoa. This looks both powerful and expensive.
My prediction is that, this time next year, top developers outside ai labs will be spending 50k USD+ on inference.
Within labs, I've heard spend is already far beyond this per developer.
by owentbrown - I mean I don't think $50k is the ceiling, unless you're talking about actual cash out. Claude code subscriptions right now can easily clear you $25-35k a year in nominal value for $2400 out of pocket cost.
Given sufficient budget and scope, I could certainly productively burn a half million dollars in tokens a year or more. I think that's where we're headed anyway, buying a 2nd or 5th claude max subscription feels slightly excessive for personal usage, but at a corporate level...
by jaggederest - not sure how "top developers" are defined here, but there is huge diminishing return curve starts kicking in after $200/month price point for typical eng work.by andriy_koval
- Isn't TCO lower with Cerebras chips compared to Nvidia? Theoretically, most developers should eventually be running on Ultrafast.
- 50k per month?!
If someone subsidize maybe, but if the companies need to pay no way, unless there is hard evidence of the return.
by vb-8448 - > top developers outside ai labs will be spending 50k USD+ on inference
I think it is more like top companies, not top developers, and the problem with developers in top companies was and is - absolute majority of them are not actually directly working on things that increase revenue, so companies can spend a ton of money and see barely if any changes in the product and the bottom line, so companies, at least legacy ones will be reluctant to sponsor that long term.
by maxnevermind - Good news for Intel and AMD.
Rught now on large scale codebases the bottleneck is both claude/codex inference, as well as time it takes to run tens of thousands of tests. We put those workloads on dedicated epyc 9005 build machines - but it still takes minutes per run. Those who can afford the fast tokens will be in the market for faster CPU that money can buy today.
by aenis - why intel and amd ? these are cerebras wafers?
i know people are joking about the sol ultrafast prices (its unlikely to be accessible for average joes) but this shows scaling wafer cores works for inference boost
which makes me very excited, sol ultrafast will be as slow as it will get if that makes sense. at these token speeds , we will see a much deeper economic impact.
by zuzululu - > GPT-5.6 Sol on Ultrafast mode, delivering up to 750 output tokens per second
> delivering 17k tokens per second per user on Llama 3.1 8B model.
Obviously this is a much smaller model, but I really can't wait for ASICs to take over the LLM space.
Imagine running a model like Sol/Fable (even half the size with 60-70% of it's intelligence) on your own ASIC hardware.
- I'd take qwen3.6 (3.8 as of tomorrow) 27B running at 17k per second first on the way to Sol/Fable! And then dsv4-flash-0731!by auspiv
- ASIC makes it sound like it's a single chip, but in reality serving trillion-param models on Cerebras requires a full cluster (as in multiple racks, MW of power).
Some interesting twitter analysis here:
by mNovak - This is really cool. Someone here commented about similarity between this and hardware advancements for AV encode/decode.
I think it's only a matter of time before miniaturization can have a thumbnail sized user-replaceable accessory that contains the LLM built onto the hardware. I admit I don't know how any of that works, but would be amazing to experience. Fully local, fully offline, ultra fast local inference better than any personal computing product.
by thraway3837 - In 5-10 years nice smartphones will be able to run ChatGPT (~gpt3-4) class models. A memory rich laptop (highend mac/framework) can run GPT-OSS:120b or full Gemma4 at very interactive speeds.
High end phones can already run the smaller models at enough speed to be probably useful, especially for background/overnight photo tagging and curation and things like that.
- https://chatjimmy.ai/ Is that. Company behind it just got acquired by AMDby christkv
- This does look pretty incredible, but don't forget that incredible token thoroughput can only necessarily solve certain bottlenecks. If your e2e tests take an hour, they'll still take an hour after Ultracode. If the agent runs a 10 minute typecheck after a change, that will still take 10 minutes. grep over a massive codebase is still just as slow, etc. I say this not to take away from this accomplishment but just to ensure everyone here keeps a clear head about what it means - 14x faster tokens does not mean it completes every task 14x faster.
I suspect Humanity's Last Exam is without tool-calls, making it kind of the perfect benchmark to highlight how fast Ultrafast is, but not really the same as the everyday work you or I do.
by johnfn - I've been measuring waiting for tool calls/waiting for model response in my OMP with Sol 5.6 and usually it's 85%-95% of time spent waiting for model to respond, so 14x speedup in model perf would still be very significant. YMMV but speeding up tests and improving DX is somewhat well understood.by fireant
- The omission of Mimo v2.5-Pro Ultraspeed, released in June, which can achieve 1000tok/s is an interesting flaw in the comparison graphs.
It is a bit outdated (scores ± 40% lower), but smart enough for a lot of coding tasks, and can cost under 1/10th of Sol.
by ricardobeat - > Compared with output speeds reported by Artificial Analysis GPT-5.6 Sol on Ultrafast mode runs 11x faster than Fable 5, and 5x faster than Opus 4.8 on Fast mode.
Awesome work. I'm personally very excited for faster models/inference.
I think speed is underrated to some degree in the current conversation. For a while, I was using Cursor's Composer quite a lot, even over frontier models, just because of how darn fast it was.
by wxw - I've been using DeepSeek flash a lot this week to try it out. Now, I deeply want the smart frontier models to be just as fast.by kilroy123
- What do you need speed for? That's a genuine question, I feel like the limiting factor already is my creativity, attention span and budget. And I'm not even yet optimizing cost by batching things like review to slow local models over night, or schedule tasks to take full advantage of my subscriptions.by arw0n
- Unless I have read over it, besides the animation in the intelligence vs speed graph which only mentions internal data and not whether they truly reran the AA suite, there is no actually solid statement on the important aspect of performance.
Neither the Cerebras or OpenAI post [0] outright state that this performs exactly the same as regular 5.6 Sol. I feel if this was 1:1 just Sol but much faster, they'd (rightfully) scream that off the rooftops. A line such as "this is the same performance, just faster, with no downsides" would go a long way in clarity and communication. Along with no pricing information, I'll hold out on further information.
by Topfi - I mean I still regularly use 5.3 spark (the cerebrus model) that comes with my sub to do rapid reviews of 5.6's work and it finds oodles of problems in about a minute.by cududa
- This is what Cerebras does- take other people's models and run them very very fast.by conception
- You (or anyone else) can just benchmark and compare. If they were serving a dumber model it would be trivially detectable.by beering
- "delivering up to 750 output tokens per second and without any quality compromise" seems pretty definitive.by Scaevolus
- The corresponding OpenAI post https://openai.com/index/previewing-ultrafast/
There is no pricing info, which could mean it's "if you have to ask..." territory or they are simply gauging interest before deciding