

Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- > May we scale smoothly, exponentially and uneventfully through A[SI]
That sentence sounds weird to me. I can't really put my finger on why, maybe the combination of adverbs, or just the fact of writing the desire of scaling as a company so directly. It feels (to me) like openly claiming their selfish goals. Or maybe I am just misinterpreting and they are referring to the whole humanity as "We" (but knowing Broadcom and in a lesser extent OpenAI doings, I am not convinced).
by dadoum - We’ve entered the “if you care about software, build hardware” phase of AIby digitaltrees
- What are the other phases. Or what are you referring to in general?by zwarag
- “People who are really serious about software should make their own hardware.” ― Alan Kayby wmf
- I have been eyeing what Taalas is doing [1] by making pure hardware models. The speed is absurd.by some-guy
- I had Opus 4.5 design an LLM inference engine in verilog, including firmware and automated verification a while ago: https://github.com/cpldcpu/smollm.c
It's of course far from optical. But lowering the implementation through the abstraction levels turned out to be extremely powerful.
by cpldcpu - Can you suggest some tutorials for Verilog and FPGAs in general?
I have a spare Tang Nano 9k but I don't feel confident about blindly asking Claude to vibecode me a solution and still would like to have at-least a basic level of understanding.
by smetannik - Microsoft, Google, and Amazon also do this, but they also have the hyperscaler datacenter infrastructure to host the chips. Designing and taping out the chip is one thing, packaging, cooling, deploying, powering, and managing the fleet is another stack entirely. Wonder where that will come from?
- Don't forget Stargate.
Update: Somebody on Twitter said it's going to be hosted 50/50 at Microsoft and Oracle.
by wmf - I haven't seen this discussed here:
So far, the accelerator is showing cost savings of roughly 50% compared with typical AI graphics processing units, Broadcom Chief Executive Officer Hock Tan said in an interview. - [0]
50% cost saving. The picture changes so quickly, there are still a lot of low hanging fruits, that I find any discussion about whether a vendor has moats, or if they can recoup investment, is moot and futile.
[0] - https://www.bloomberg.com/news/articles/2026-06-24/openai-an...
by signatoremo - "Typical" is doing a lot of work there. That could mean much older chips than Nvidia is currently selling.by Schiendelman
- If GPUs have 75% margin then 50% cheaper is no surprise.by wmf
- >designed for initial deployment by the end of 2026 and expanding in the years ahead,
So after the IPO and will be featured heavily in the IPO sales brochure as a future promise?
I'm sceptical over any pre-IPO announcements.
by v5v3 - Who's IPO? Broadcom and Google are already listed, obviously.by frandroid
- Yeah, the narrative feels like pre-IPO shenanigans, and it looks like the lid on my laundry basket. I wouldn’t be surprised if this is a con.by estetlinus
- With the pace of AI, and with AI helping to pave the way for faster/better AI, I keep wondering if hardware like this will become obsolete well before it has a meaningful ROI. Huge AI models can be run with less resources already through quantization and offloading, but that's just the beginning. One day, maybe not far from now, a breakthrough will allow huge LLMs (say 200B in size) to run well on an old 5 year old Dell desktop. Think that's crazy? Look at the size of the first hard drives. The IBM 350 was a disk with 50 platters, 24 inches in diameter, that held 3.5Mb, and was leased for today's equivalent of $35K.
https://www.computerhistory.org/storageengine/first-commerci...
Compare that to a multi-terabyte ssd. Now apply that improvement to how an LLM is architected and run now. With AI assisting, it won't be long before a leap occurs and these data centers with all their current ultra-cutting edge Nvidia cards are nearly obsolete overnight.
by deweywsu - True but as someone else pointed out; at that time we'd be interested in running 200T parameter model rather than 200B. Why, you might ask? Law of human laziness - a human will become as lazy as the technology allows it to. With the 200T or 20,000 T model - I'd be heavily incentivized to ask it to make the bread for me that I enjoy making now or create a movie for me (featuring myself) which will maximize the dopamine production in my brain.by dwa3592
- Looking at the development of memory bandwidth, capacity and prices over the last 10 years there is little indication that’s likely.by hyhatqtv
- > I keep wondering if hardware like this will become obsolete well before it has a meaningful ROI
it will build expertise/infra/know-how foundation for next generation of hardware
by andriy_koval - > One day, maybe not far from now, a breakthrough will allow huge LLMs (say 200B in size) to run well on an old 5 year old Dell desktop.
I think there will be specialized hardware (beside GPUs) that would be custom made for LLMs. Yes TPUs exist, but mainly for datacenter. GPUs exist, but they are adapted from mainly graphic application. Once all the demand from data center dries up, innovation will kick in.
by 3abiton - Usually breakthroughs in computing lead to more usage of computing, not less.by gdiamos
- I think Jevons Paradox and scaling laws will make this not the case. If bigger models are always better (which seems they are), then will always need high-end hardware.by LZ_Khan
- Interesting comment, but the comparison with hard disk drives is probably unfair.
The IBM 350 was commercialized 70 years ago; it took 70 years for someone like you to be able to compare that to a multi-TB SSD.
Furthermore, nothing says that Moore's Law will necessarily apply to LLMs, for decades to come.
- > One day, maybe not far from now, a breakthrough will allow huge LLMs (say 200B in size) to run well on an old 5 year old Dell desktop.
But if you have such a breakthrough could you not also apply it and run 200T models on todays datacenters?
by admax88qqq - Pretty huge move. Google and their TPUs are looking infinitely more prescient as I think they are on their 7th generation, along with the offshoots it inspired like the LPU and even others, perhaps like Cerebras and their Wafer Scale Engine.
However, based off first impressions, it seems like this is meant for inference side, and not training, which is also an interesting choice.
by maz1b - With Reinforcement Learning, inference is very present in post-training stages now tooby ggcr
- > early testing shows that Jalapeño will deliver performance per watt substantially better than current state-of-the-art
We're starting to see what really matters here, and though this is hand wavy the TPU makes similar claims.
I think googles memo about having no moat still stands (see: https://newsletter.semianalysis.com/p/google-we-have-no-moat... if you are unaware). It kind of makes sense that all of this is looking more like 60's to 90's IBM, DEC, Cray, Sun and the hardware race that happened then. History doesn't repeat but it often rhymes and I suspect that these efforts will follow the same trajectory.
by zer00eyz - Cerebras's Codex Spark 5.3 has been a huge flop. Small context window and old model. But hopefully they can improve so that we can benefit from 1000 tokens/second with GPT 5.5.
- Inference costs are higher than training now. I think.
Nvidia is king of general purpose training chips. But inferences can be specialized.
- Training is pretty much a 1x cost, and efficiency there is already on the way down with architectural improvements. Inference though is an ongoing cost which over time takes orders of magnitude more resources, so focusing on making that far more efficient means way greater gains over time.by skeledrew