Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- Some of the most insidious parts of AI infrastructure includes the embedding model. Corporations have already spent an outstanding amount of time and money creating embedding vectors that are closed source and not reproducible. This means that all their data is locked into whatever embedding model they chose initially.
I highly recommend utilizing an open sourced embedding model instead of paying for a closed source one. It's vastly more reasonable to run an open sourced embedding model as a first step. They're much, much smaller and, due to the overhead of network latency, and running it locally has almost the same speed as through an API even on slow computers.
I would even go so far as to say that closed source embedding models have a high risk of data hostage. If a team doesn't have access to the embedding model, the embeddings become useless. A corporation like OpenAI could, say, hike the prices to that model by 1000x and everyone would have to pay up or forfeit any utility of the data.
I envision a future where open source embedding models are shipped with relevant technologies and implemented by currently under-utilized chips like NPU's. A startup developing cheap microprocessors that can run them is an idea I would pay cash for. Or perhaps they will be bundled with security tokens.
While it might be impractical for all corporate teams to run language models, it is very realistic for everyone to operate an open sourced embedding model, at least in their private cloud. Better yet, utilize transfer learning on an open sourced one to train your own, that way the embedding vector is more secure against competitors and trade secrets.
by kittikitti - What are people even using embeddings for these days? It certainly seems like giving an agent grep covers most of the use cases. Dare I say: grep is all you need.by Zambyte
- The financials are incredible. If you spend one engineer's salary on hardware, you get a system running local AI that can multiply the efforts of an entire (small) team of engineers. It's a very "you can't afford not to" situation. Even considering the hardware prices today.by unrented7977
- I'm puzzled why so many comments here are about local AI? Article is clearly about open models on openrouter etc.by spopejoy
- They couldn't pick a more sinister headline for such an awesome technological development.by falaki
- Its almost as if NYT has had a hawkish agenda for the past few ...generations.by overfeed
- Ah, so you're one of those hippies hooked on free software too, yeah? What's that your smoking there? Emacs, huh? what's your OS? Linux? I knew it.
Corporal, put him away.
by golem14 - Easy prediction: LLMs will get shrunk down further and further until GenAI is just something that ships on a chip as part of your hardware. In the future it will seem quaint that we needed a network connection to talk to our LLM.
Adoption of open-source models to my mind is a similar step in that direction. In all cases, the goal is to become untethered from a mercurial vendor.
by _doctor_love - for every small GenAI model there will be larger model or cluster of models which are smarter than small modelby andriy_koval
- Yeah, I'm looking forward to this actually.
https://chatjimmy.ai/ blew my mind at how fast etched model weights can be.
For on-device LLMs, there's a point of diminishing returns, meaning you don't need to have the latest frontier model for most operations.
by wnmurphy - We're gonna start baking in models like TTS with thousands of voices available in any language as a chip on device. They just need to hit 99% accuracy and then it's a done deal.by hparadiz
- Like taalas.com (very recently acquired by AMD), or cerebras.ai (whole wafer is a chip)? As you said, I also think that is one of the main direction many companies (and academia) is moving to.by honr
- > AT&T turned to artificial intelligence models from Anthropic and OpenAI in recent years to help with customer service, call transcription and coding. [...]
> By May, open models accounted for 20 percent of AT&T’s A.I. use. That has since risen to 40 percent and may jump to 60 percent in the coming months, Mr. Markus said in an interview.
This is missing a crucial detail. We know they "help with customer service, call transcription and coding", but which of those have been upgrade to open models?
Call transcription is trivial to do with open models. I can run Whisper or Parakeet on a low-spec laptop.
"Customer service" could mean a lot of things, but it sounds feasible for open models too.
"Coding" - they might go to open models for that, but I expect the costs involved in paying for closed models for software developers within AT&T are a fraction of the costs involved in transcribing all of their calls or handling aspects of custom service for millions of customers.
From later in the story:
> AT&T researches Chinese models but is not using them, Mr. Markus said. Instead, it is working with popular alternatives made by American companies such as the Gemma A.I. model from Google and the Llama A.I. models from Meta.
Gemma 4 is great, but really, Llama, in 2026?
by simonw - > Gemma 4 is great, but really, Llama, in 2026?
I'd assume the author is just getting confused because of ollama and llama.cpp and all the other ecosystem "llama" that are still in use. Llama really did kick off the open models thing
by mewse-hn - Spent most of today reworking a rack and rig of gpus for all of our internal ai work… our big server is 8 rtx 6000 pro and 3 psu, I definitely feel I made a mistake not upgrading our wall power to 240v but so far we have multiple 30amp 120v and with 3 PSU uninterrupted power we have been very stable. My big upgrade will be moving to epyc motherboard from threadripper so we get gen5x8 with bifurcation instead of what we are stuck with today gen5x4 due to bios limitations . What has been so encouraging though is first deepseek v4 flash at 200+ t/s for single user and much more in aggregate- now on qwen 3.8 flash next for image support and eyes on glm5.3 flash for some testing … next is realistically considering co location and quiet a sizable loan to scale this to real hardware instead of miner rip vibesby taf2
- What kind of cost are you looking at and how many people could use it? I’d be interested to know what the payback period is like, because Claude code is getting ridiculously expensive.
- I'd love to! For real coding though, SOTA models barely get the job done. It wasn't until Opus 4.5 that you could really get decent results.
I'm sure this will change (and I can't wait for it!) but as of today, open models might be fine for summarizing and writing docs, but you need SOTA to work on code if you want to be competitive.
by slowin - I generally use sonnet 5 for most coding tasks, a lot of coding is really pretty straightforward.by transdev12
- I'm having no difficulty getting deepseek v4 to blast out good code. I guess devs want to outsource all of their thinking now? Yes open models maybe can't design the whole thing soup to nuts but why is that necessary?by spopejoy
- You're not corporate America (and trust me, I mostly mean that as a plus).
I also work in software, and while I vaguely disagree that open models can't be used (they absolutely fit into productive niches here, and holy hell are the last generation [ex laguna s1, kimi k3, glm 5.3, etc] actually decent) - I will agree that SOTA are a better fit for software development, especially when used in conjunction with an already very expensive employee who's driving them.
But for "Corporate America"... no. You absolutely don't need SOTA. They're doing things like transcription, summarization, customer interaction, minor technical tasks like form creation in existing tools, report generation (ex - powerpoint, pdf, docs, etc) and other general "white collar tasks". Think about roles in business that are in the 60-85k compensation range.
It's mostly busy work that keeps existing processes flowing and the business on the rails. Important, but not research/novel.
And cheap ai... is a wonderful fit for a lot of this. No one wants to replace an employee making 80k with a less reliable AI that costs 45k a year in tokens (SOTA). But they're absolutely willing to drop 2-3k/year on AI (~100/month - right in the open model cost range) for that employee if they can get a 10% bump in productivity or happiness.
by horsawlarway - I do think there are now open weight models that are on par with (or beating) Opus 4.5 by now (e.g. Kimi K3, GLM5.3). But yeah obviously the frontier closed source models seem to have pulled away once again, so open weight seems to be a few months behind right now (which might be too long to wait for a lot of people!).by kbwal7
- Every time I check in on this I hear a more recent model is the one where they started doing good work. I'm excited to here that Astra is where it got capable enough to work on code next yearby nemomarx
- > Some U.S. firms remain reluctant to use Chinese A.I. models because of concerns over regulation and data privacy. AT&T researches Chinese models but is not using them, Mr. Markus said. Instead, it is working with popular alternatives made by American companies such as the Gemma A.I. model from Google and the Llama A.I. models from Meta.
This makes sense since corporations require legal certainty, and using an open model from an American company (probably) provides them some level of indemnity, and also someone to sue.
by petcat - Read the license again.
- Exactly. See TikTok trouble as example and quite honestly, try a local open source LLM and ask it to use profanity, paint nudes - the LLM doesn’t answer the question of it is from OpenAI or Google.
The thing is that needs more attention is reverse engineered a LLM which is highly fascinating. I tried it, but it seems I am not there yet to put it mildly. It requires serious effort.
I am just speculating but can LLMs be sleepers? You write software and it seeds traces here and there under certain conditions that pose a serious security risk.
Or a kill switch?
I don’t know. I distrust Chinese LLMs but even more due to training data.
It is after all not a Western model. Different biases and the might be subtle but nevertheless substantial.
In short: no open source LLM may be usable without additional Finetuning for certain valid use cases.
The real value is versioning and autonomy as well as lot more stable answering despite model rot.
Also testing and the supporting systems are easier to maintain.
It is mainly an infrastructure challenge.
- The current U.S. regime is also replacing some amount of that legal certainty with regime fealty. Picking Chinese options over American ones probably runs a risk of upsetting their leader. I've got to imagine American companies are weighing this factor in their decisions.by Waterluvian
- I swear, Qwen 3.8 27B @ Q8 is smarter than Sonnet 5 most of the time. Why wouldn’t corporate America self host at this point, especially with better options like Deepseek Flash and GLM 5.3 flash that’s a middle ground between Sonnet and Opusby syntaxing
- I mean, why self-host? Deepseek v4 flash is dirt-cheap on openrouter. Inference is a race to the bottom at this point.by spopejoy
- Why not Luna?by skybrian
- is this actually the case? I haven't kept up with the small models
but if there's roughly Sonnet 4.6 level capable open small models, then I'd be impressed
by r_lee