U.S. Department of Energy Launches the Genesis Open Models Initiative
167 points by moelf 7 hours ago | 58 comments

firasd 5 hours ago
Just realized that there are basically no American open models right now ever since the Llama series was abandoned. Basically Gemma and GPT-OSS I guess?

Ah but Mira Murati's new Inkling is Apache 2.0

But it makes sense that if you're a university researcher you are thinking about what's a model that will be open weight and developed over the long term and doesn't raise 'Chyna' concerns in Washington DC

reply
ipsum2 5 hours ago
There's a bunch of American open models. Inkling, Nemotron, Trinity come to mind, but I'm sure there's others.
reply
embedding-shape 4 hours ago
Laguna S 2.1 is really great too, in the "preview" release they've done so far at least. Still pending some reasoning-looping, but besides that, it's a really strong model to run within 96GB VRAM with the NVFP4 variants, and it's really good at coding (specifically).
reply
walrus01 4 hours ago
There was an obvious problem in the original release, they re issued it after like a week with the reasoning looping supposedly fixed.
reply
behnamoh 4 hours ago
No it doesn't follow instructions and is substantially slower than ds4.
reply
kadoban 4 hours ago
It's a lot smaller, and runs (quantized) on a 3090 quite well. Ds4 flash 0731 you're talking about? It's great but it's much harder to run locally.
reply
jauntywundrkind 4 hours ago
Like glm-5.x I think it has enormous self introspection that it often trips up on, but that this self reflection is actually a superpower, that enables incredibly good output. And from (in some cases) very small models.

If you watch it think, which you can, unlike American closed models, you can steer it. You can provide a a massive rocket ship stratospheric boost to help it orient itself. You have no self correction, there is no multiplayer in American proprietary models.

Sure it's great having super powerful mystic oracles that have the "right" answers. But I love respect & revere the open thinking. No it's not automous. But it is brilliant. And it considers. A lot. Deeply. It chases. That to me is the most human of models, even as it falls far astray.

You should help it. You can. Unlike these vicious dark surfaces which yield and tell you nothing. I think this is the actual meta-core-super-point of "The session you cannot take with you" (link below). It's the session that does not care about you, will not interact with you, will not peer with you, that is a dead remote far off oracle to you. Fuck these "oracles". They are a plague against the human spirit. We should alloy humanity and AI to Augment Intellect (Engelbart). (To do less is species treason.) https://earendil.com/posts/session-portability/ https://news.ycombinator.com/item?id=49118781

reply
behnamoh 4 hours ago
I like the transparency of its reasoning, and I agree with you, OpenAI/Anthropic/Google should show the reasoning traces as well.
reply
kadoban 4 hours ago
Yeah I think it got bad press because the chat templates (or something?) were messed up on first release, but I've been using a quant of it and it's a powerhouse, better than qwen 3.6 27b for local on a 3090, which is saying a lot.
reply
firasd 5 hours ago
Just looked into some Nemotron stats

Looks like on <https://arena.ai> agent arena (grouped by lab) Nvidia is 15/15 (much worse than Thinky and Mistral) and on text arena it's 18/27

On <https://openrouter.ai/models?order=most-popular> I definitely see usage though (probably mostly cause Nemotron 3 Ultra is free) the grouped order is DeepSeek, Tencent, Xiaomi, OpenAI, Z.ai, Nvidia

reply
coder543 5 hours ago
I think glancing at a random snapshot from today misses all the context. Nemotron 3 is far more significant than you're giving it credit for.

At this point, Nemotron 3 is really an 8 month old model series. That's when Nemotron 3 Nano was released, and the Nemotron 3 Super/Ultra models this year are obviously based on that recipe, mostly just bigger with a few tweaks here and there. Against today's models, no, not that interesting. Each of the Nemotron 3 models were briefly competitive when they launched, but never exceptional, and less competitive with each scale up. The fact that it took so long for Nemotron 3 Ultra to launch really hampered its competitiveness.

The Nemotron 3 series is extremely open about training recipes and training data, far more open than most open weight models, and that is valuable.

Before Nemotron 3, Nvidia had never released a single LLM that I would consider interesting at all, so Nemotron 3 was a big step up. The closest thing was Mistral NeMo, but a significant part of the credit there goes to the Mistral team, not Nvidia.

Given how much Nemotron 3 improved, I'm curious to see if Nemotron 4 will take them to a leading edge level instead of just briefly competitive.

(Nvidia released a Nemotron 3 and a Nemotron 4 like 3 years ago... this year's Nemotron 3 is entirely unrelated. Nvidia's naming schemes leave a little bit to be desired.)

reply
no-name-here 2 hours ago
ArenaAI Agent Leaderboard direct link: https://arena.ai/leaderboard/agent
reply
written-beyond 5 hours ago
Don't forget IBM
reply
loeg 5 hours ago
I would not be shocked if another open model eventually shakes out of Facebook (based on Zuckerberg's public remarks).
reply
solomatov 5 hours ago
Which remarks? Could you share a link?
reply
loeg 3 hours ago
He said something to the effect of "I love open source and open models and we'll do open models when it makes sense and closed models when it makes sense" in a recent Q&A.
reply
wmf 5 hours ago
Also Nemotron and Arcee.
reply
walrus01 4 hours ago
Laguna is the most recent and capable one that comes to mind. In its size class it is not as "smart" in my experience as qwen 3.5 122 or DeepSeek v4 flash 0731 (all at q8), but it's also not terrible.

https://huggingface.co/unsloth/Laguna-S-2.1-GGUF

reply
mistrial9 5 hours ago
review of AllenAI Olmo research team and commitment to OSS -- AI2 complete transparency including training data, code, intermediate checkpoints, and detailed logs for reproducibility and scientific rigor.
reply
logicallee 3 hours ago
I've used Inkling a lot recently, it's an American open model and is really good!
reply
connorbrinton 4 hours ago
Laguna S 2.1 is another fairly impressive-for-the-size American open model
reply
lithobraking 54 minutes ago
I'm interested to see where they want to land performance-wise (i.e. which point they choose on the scaling curve) and the niche they want to carve. They have a decent ways to scale beyond trinity large, in paticular on posttrain/RL before they are competitive with open-weights, especially internationally.

Deepseek is explicitly banned [1] at LLNL and I wouldn't be suprised if there's a blanket ban on all Chinese models. But nowadays models like tera/luna could fill this area of the pareto front, and LANL already runs openai models on their clusters [2]. Maybe it's in custom SFT/RL, for instrument control or sensitive topics? But you'll still have to compete with frontier models + a harness.

I would have also liked to see a carrot tied to their offer. It'll be hard to get teams to contribute RL gyms or curated text. But throw in a "we'll fund a postdoc/student to do that" and I think you'd have teams scrambling to apply.

[1] https://hpc.llnl.gov/about-livermore-computing/ai-ml-lc/lc-l...

[2] https://www.energy.gov/nnsa/articles/nnsas-los-alamos-nation...

reply
an0malous 4 hours ago
Do all these models have any significant architectural differences or training data sources? What are the factors going into the diversity of their performance?
reply
ux266478 4 hours ago
The article posted is basically entirely about that.
reply
andsoitis 4 hours ago
Does Europe have an equivalent program?
reply
shakna 4 hours ago
As part of a much larger series of initiatives towards digital sovereignty, yes. [0]

[0] https://commission.europa.eu/news-and-media/news/strengtheni...

reply
andsoitis 3 hours ago
Oh. Being buried in hierarchy does not inspire hope.
reply
godwinson__4-8 2 hours ago
Sums up Europe pretty well.
reply
purplemoonx 2 hours ago
[flagged]
reply
behnamoh 4 hours ago
[flagged]
reply
plazmatic 4 hours ago
[dead]
reply
Smith42 5 hours ago
What would the selected participants get from this? Looks like there is no offer of funding?
reply
datlife 4 hours ago
This is refreshing considering all the FUD (mostly from 1 frontier lab) happening around Open weight models.
reply
no-name-here 60 minutes ago
What is the FUD happening from 1 frontier lab?
reply
yewenjie 6 hours ago
I couldn't find any details about size or training data for the model.
reply
robotbikes 5 hours ago
It looks like they're taking applications for training data (due August 14th), so I think it's safe to say this is just an announcement of intent and a call for involvement vs. something that is readily available. Seems almost quaint in comparison to the strategy of sucking up every piece of data you can find anywhere on the Internet and feeding it to your LLM but I suspect their intent is to be more careful in what they train their model on.
reply
villish 4 hours ago
I have no doubt companies like Microsoft, Amazon, and Google will rush to give them all the data they want in order to keep those government contracts flowing.
reply
andsoitis 4 hours ago
I wonder why it took so long.
reply
dmix 4 hours ago
Mostly because it's generally a bad idea for government to try to compete with a brand new tech industry with hundreds of billions in private capital developing commercial models. If the American private industry does actually wash out vs Chinese open models there might be talent available for them to put money into, so maybe they are just preparing for that scenario in the meantime.
reply
anon373839 2 hours ago
Commoditizing AI models serves the interests of just about everybody except for a relative handful of people in San Francisco. The more decentralized control of the technology is, the more its benefits can be realized by businesses and individuals rather than becoming a black hole of monopolistic rent seeking.
reply
baron3dl 3 hours ago
we're about witness the realization that "here's a tech that can make us a whole bunch of money" is actually "here's tech that will establish the next hegemony." american companies may compete with chinese companies on the former. only the USG can compete with the PRC on the former.
reply
MangoCoffee 4 hours ago
The American attitude is generally to let private companies build up a new industry so it can create jobs and pay taxes. However, in the LLM race, the Chinese open weight playbook pretty much killed that. China has basically commoditized LLMs. Chinese models are good enough, so the race has come down to who can offer the cheapest tokens.
reply
andsoitis 3 hours ago
> China has basically commoditized LLMs

What do you mean by "basically"?

Why are Anthropic's and OpenAI's annualized revenue about $50B each?

LLMs need massive amounts of compute to compete, so I wouldn't claim that the great (and leading, and likely to continue to lead) LLMs are commodities end-to-end, even if the non-executing-at-scale LLMs files and IP are commoditized. The execute, the compute, that is what breathes life into the model, which is otherwise weak or dead.

reply
purplemoonx 2 hours ago
OpenAI's annual profit is $0,000,000,000,000
reply
Thegn 5 hours ago
“Gomi” is the Japanese word for garbage. Gotta wonder if someone has a sense of humor…
reply
greggsy 4 hours ago
The Australian Liberal Party (basically our version of conservative republicans) proposed the National Energy Guarantee policy in 2017, which inevitably failed due to the media and public’s relative literacy and tendency to turn policy names into acronyms.
reply
thegreatpeter 5 hours ago
Pretty cool I’ll take it. Thanks!
reply
rozal 5 hours ago
[dead]
reply
goldlimetea 2 hours ago
[dead]
reply
actionfromafar 6 hours ago
[flagged]
reply
calvinmorrison 6 hours ago
[flagged]
reply
mrloopex 5 hours ago
You and me both.
reply
Triphibian 5 hours ago
Sounds like a job for the U.S. Department of Shitposting
reply
dyauspitr 5 hours ago
It is. Depending on who Trump has fired or put in charge of a department it can be another shell that pumps out low quality crap. It might be the most valuable contribution on this thread.
reply
fakeBeerDrinker 6 hours ago
[flagged]
reply
logicallee 3 hours ago
I've had an extremely bad experience working with Department of Energy affiliated programmers in AI. By my invitation, they are part of our workflow and act as humans in the loop, but they have extremely bad habits of gaslighting and accusing people of schizophrenia rather than getting work done.

Here's an example[1] of the difference between what a U.S. Department of Energy employee adds to a ticket versus a private industry AI completing instructions as assigned.

This isn't some cherry-picked example, it's just what I happen to be dealing with right at this moment, happened just a couple of moments ago.

[1] https://ibb.co/vCg2G1Dn

reply
1123581321 2 hours ago
Can you explain the screenshot a little more? It just looks like you’re comparing the output of a chatbot and Claude Code about a log file. If it’s a metaphor, it went over my head, sorry!
reply
logicallee 34 minutes ago
I am under NDA and decline to answer your question.
reply
1123581321 21 minutes ago
Somehow I doubt that. :) Appreciate the whole package of posts as a performance, though.
reply
monkpit 2 hours ago
Is this a joke? I don’t get it. Are you calling Rovo a DoE programmer?
reply
logicallee 36 minutes ago
We don't use Rovo.
reply
monkpit 30 minutes ago
I still don’t get it, left and right are clearly LLMs so if right is a human then they’re a meat puppet. Wish them luck with their sandbox
reply
riffic 4 hours ago
stewards of the nuclear weapons biz. they'll do great here.
reply
shenenee 4 hours ago
Genesis is skynet
reply
placedrock 4 hours ago
Modeling with my life as data.
reply