Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- it's just a wall of blah blah until I see a gguf on hfby dangoljames
- Why is this a DOE thing?by hammock
- They have an insane amount of compute at their disposal.by nunez
- One of the few sectors of the american federal government still funded to do science after the Big Beautiful Bill scrapped everything else.by Schlagbohrer
- Contributing to a project like this seems like a great way to get yourself export controlledby victor9000
- dystopian sci fi movie narrator voice in 2020, she was branded an Essential Worker shows her in chains at a grocery store checkout counter. In 2026, she was... dun dun dun, EXPORT CONTROLLED shows her getting those words stamped on her neck Coming this summer, American Worker, a new sci fi horror movieby Schlagbohrer
- Does Europe have an equivalent program?by andsoitis
- As part of a much larger series of initiatives towards digital sovereignty, yes. [0]
[0] https://commission.europa.eu/news-and-media/news/strengtheni...
by shakna - Do all these models have any significant architectural differences or training data sources? What are the factors going into the diversity of their performance?by an0malous
- The article posted is basically entirely about that.by ux266478
- What would the selected participants get from this? Looks like there is no offer of funding?by Smith42
- by Aeroi
- This is not the same thing, right? IIUC, the awards for what you linked have already been given out. There aren't awards for the linked initiative - I think that's just Argonne National Lab asking for volunteers to make their (ANL's) award money stretch further, right?by edot
- If you’re in the weird position of knowing more about the national labs than the AI lab scene (like I am), link to a TechCrunch profile:
https://techcrunch.com/2026/01/28/tiny-startup-arcee-ai-buil...
My question is: why is this being run out of Argonne? Why not NERSC proper?
by nxobject - Your question piqued my curiosity.
I thought maybe NERSC doesn't accept jobs from private corporations but I checked and that's not true, as long as results are not held proprietary.
Perhaps ANL was used as they have a lot of compute and they lead and host the Genesis Open Models Initiative?
From your link, 2048 B300 GPUs were used for 6 months. If google search is right, NERSC has 7168 A100. B300's are way more capable than A100. To do this training in 6 months, "3 NERSCs" would be needed.
Between the political angle and the technical, I'd guess these two make up a big chunk of the answer.
by frumiousirc - It would be extremely interesting to me if the usgov produces a model which honors copyright and is also useful. This would give them extreme leverage over the labs, who may be violating copyright in significant and obvious ways.by sroerick
- Honestly the battle for copyright with models is lost. The takeaway is copyright applies to you as a small user and not to billion/trillion dollar companies. Same as any other US law, really.by appplication
- There's no mention of "LLM" nor "language". It does mention "foundation model" which includes LLMs but that also includes non-LLM architectures and non-text data. Many of the Genesis Initiative proposals answer "foundation model" call with non-LLM systems. All the FM's I know about currently in this sphere are non-LLMs. The "about gs1" page also does not mention "LLM" but does talk more about agentic harness and workflows. That description certainly sounds LLM'ish but describes a more rich system. I don't mean to suggest that LLMs will not be part of these "genesis open models" but as described, this will not result in a replacement for the "claude" or "codex" commands.by frumiousirc
- I'm interested to see where they want to land performance-wise (i.e. which point they choose on the scaling curve) and the niche they want to carve. They have a decent ways to scale beyond trinity large, in paticular on posttrain/RL before they are competitive with open-weights, especially internationally.
Deepseek is explicitly banned [1] at LLNL and I wouldn't be suprised if there's a blanket ban on all Chinese models. But nowadays models like tera/luna could fill this area of the pareto front, and LANL already runs openai models on their clusters [2]. Maybe it's in custom SFT/RL, for instrument control or sensitive topics? But you'll still have to compete with frontier models + a harness.
I would have also liked to see a carrot tied to their offer. It'll be hard to get teams to contribute RL gyms or curated text. But throw in a "we'll fund a postdoc/student to do that" and I think you'd have teams scrambling to apply.
[1] https://hpc.llnl.gov/about-livermore-computing/ai-ml-lc/lc-l...
[2] https://www.energy.gov/nnsa/articles/nnsas-los-alamos-nation...
by lithobraking - That's interesting a locally hosted LLM would be banned. I'm assuming locally hosted is included. Do they think it's been trained to sabotage equipment?
- I'd actually suggest a great starting point would be a local command reviewer LLM. Could ostensibly be a modern AV type thing. Particularly seeing this lately has driven the need home deeper to me: https://x.com/chrisbanes/status/2085341561609425230?s=20
An open weight tool call auto-reviewer, has all sorts of achievable scaling curve milestones.
by cududa - Just realized that there are basically no American open models right now ever since the Llama series was abandoned. Basically Gemma and GPT-OSS I guess?
Ah but Mira Murati's new Inkling is Apache 2.0
But it makes sense that if you're a university researcher you are thinking about what's a model that will be open weight and developed over the long term and doesn't raise 'Chyna' concerns in Washington DC
by firasd - Even if by “open models” you specifically restrict that to LLM or LLM-backbone models with different or additional modalities to text released by major American firms then there are still a lot of American open models being released. Many of them are small models and/or highly-specialized fine-tunes of other open models, but there are still a whole lot.by dragonwriter
- Laguna is the most recent and capable one that comes to mind. In its size class it is not as "smart" in my experience as qwen 3.5 122 or DeepSeek v4 flash 0731 (all at q8), but it's also not terrible.by walrus01
- Also Nemotron and Arcee.by wmf
- review of AllenAI Olmo research team and commitment to OSS -- AI2 complete transparency including training data, code, intermediate checkpoints, and detailed logs for reproducibility and scientific rigor.by mistrial9
- What makes more sense is to do something like deepseek at one of the major universities put those bright computer science young minds to work, in the good old days almost every major university would have done that has any major US university done that? Stanford Harvard Berkeley if the Chinese can put together a team like deep seek why can’t that be done at a major university in the United States?
https://www.interconnects.ai/p/the-american-deepseek-project
by Danox - There's a bunch of American open models. Inkling, Nemotron, Trinity come to mind, but I'm sure there's others.by ipsum2
- LiquidAI LFM models are amazing, but very situational. IBM Granite series are also unique and interesting for trying to reduce liability and extend local context size. Nvidia ships some and there was also that Inkling model recently. Poolside just released theirs.
Meta might release something this year. X AI's Grok is still due to release a model, if Elon keeps to his word even if they only release a distilled version. Reflection AI has been quiet, but their access to compute is ramping up. Microsoft's MAI is considering releasing some open weight models which would be great to see!
Ilya's SSI is unlikely to release an open model since he's aiming for radical safety. That bet could pay off if the existing approach produces so much chaos within the next 10-20 years that some global ban is achieved and a super safe model is promoted as the compliant route.
We don't get many huge model releases though. I think it's harder and more expensive to safety align them. Even if you do, people will work around the safety and abuse the models. Plus it makes it even easier for Chinese companies to distill things that aren't as easy over filtered APIs.
There is a lot of internet propaganda to the effect that the US is simply unable to release open weight models or that China has so many more AI companies that the US is drowning in Chinese open weight models, but it's more like we're being careful and China doesn't care. If you host a model in China, it has to be censored and downloading any models requires you to provide your identity. Huggingface is banned there. When they release their open models in the west, they don't have to care whether the models are aligned in any way.
by CMay - Not only are there many American open weight models as others have mentioned, but Americans are the only ones doing actual open source models [0]. Not just distributing binary blobs and calling them "open".by petcat