5 ms·
Just realized that there are basically no American open models right now ever since the Llama series was abandoned. Basically Gemma and GPT-OSS I guess? Ah but
by firasd 1mo ago
Just realized that there are basically no American open models right now ever since the Llama series was abandoned. Basically Gemma and GPT-OSS I guess?
Ah but Mira Murati's new Inkling is Apache 2.0
But it makes sense that if you're a university researcher you are thinking about what's a model that will be open weight and developed over the long term and doesn't raise 'Chyna' concerns in Washington DC
- wmf 1mo agoAlso Nemotron and Arcee.
- ipsum2 1mo agoThere's a bunch of American open models. Inkling, Nemotron, Trinity come to mind, but I'm sure there's others.
- firasd 1mo agoJust looked into some Nemotron stats Looks like on <https://arena.ai https://arena.ai> agent arena (grouped by lab) Nvidia is 15/15 (much worse than Thinky and Mistral) and on text arena it's 18/27 On <https://openrouter.ai/models?order=most-popular https://openrouter.ai/models?order=most-popular> I definitely see usage though (probably mostly cause Nemotron 3 Ultra is free) the grouped order is DeepSeek, Tencent, Xiaomi, OpenAI, Z.ai, Nvidia
- no-name-here 1mo agoArenaAI Agent Leaderboard direct link: https://arena.ai/leaderboard/agent https://arena.ai/leaderboard/agent
- coder543 1mo agoI think glancing at a random snapshot from today misses all the context. Nemotron 3 is far more significant than you're giving it credit for. At this point, Nemotron 3 is really an 8 month old model series. That's when Nemotron 3 Nano was released, and the Nemotron 3 Super/Ultra models this year are obviously based on that recipe, mostly just bigger with a few tweaks here and there. Against today's models, no, not that interesting. Each of the Nemotron 3 models were briefly competitive when they launched, but never exceptional, and less competitive with each scale up. The fact that it took so long for Nemotron 3 Ultra to launch really hampered its competitiveness. The Nemotron 3 series is extremely open about training recipes and training data, far more open than most open weight models, and that is valuable. Before Nemotron 3, Nvidia had never released a single LLM that I would consider interesting at all, so Nemotron 3 was a big step up. The closest thing was Mistral NeMo, but a significant part of the credit there goes to the Mistral team, not Nvidia. Given how much Nemotron 3 improved, I'm curious to see if Nemotron 4 will take them to a leading edge level instead of just briefly competitive. (Nvidia released a Nemotron 3 and a Nemotron 4 like 3 years ago... this year's Nemotron 3 is entirely unrelated. Nvidia's naming schemes leave a little bit to be desired.)
- buildbot 1mo agoNemotron 3 also introduced LatentMoE, which was adopted by Kimi K3 :)
- strictnein 1mo agoNemotron is a very nice model with an excellent license as well.
- written-beyond 1mo agoDon't forget IBM
- embedding-shape 1mo agoLaguna S 2.1 is really great too, in the "preview" release they've done so far at least. Still pending some reasoning-looping, but besides that, it's a really strong model to run within 96GB VRAM with the NVFP4 variants, and it's really good at coding (specifically).
- behnamoh 1mo agoNo it doesn't follow instructions and is substantially slower than ds4.
- jauntywundrkind 1mo agoLike glm-5.x I think it has enormous self introspection that it often trips up on, but that this self reflection is actually a superpower, that enables incredibly good output. And from (in some cases) very small models. If you watch it think, which you can, unlike American closed models, you can steer it. You can provide a a massive rocket ship stratospheric boost to help it orient itself. You have no self correction, there is no multiplayer in American proprietary models. Sure it's great having super powerful mystic oracles that have the "right" answers. But I love respect & revere the open thinking. No it's not automous. But it is brilliant. And it considers. A lot. Deeply. It chases. That to me is the most human of models, even as it falls far astray. You should help it. You can. Unlike these vicious dark surfaces which yield and tell you nothing. I think this is the actual meta-core-super-point of "The session you cannot take with you" (link below). It's the session that does not care about you, will not interact with you, will not peer with you, that is a dead remote far off oracle to you. Fuck these "oracles". They are a plague against the human spirit. We should alloy humanity and AI to Augment Intellect (Engelbart). (To do less is species treason.) https://earendil.com/posts/session-portability/ https://earendil.com/posts/session-portability/ https://news.ycombinator.com/item?id=49118781 https://news.ycombinator.com/item?id=49118781
- behnamoh 1mo agoI like the transparency of its reasoning, and I agree with you, OpenAI/Anthropic/Google should show the reasoning traces as well.
- deleted 1mo ago[deleted]
- stogot 1mo agoOpenAI has gpt-oss that they said is open weight
- maziyar 1mo agoYeah thankfully we have more than we had in 2025! I am sure we will see even more open models by US based startups before the end of 2026
- loeg 1mo agoI would not be shocked if another open model eventually shakes out of Facebook (based on Zuckerberg's public remarks).
- solomatov 1mo agoWhich remarks? Could you share a link?
- loeg 1mo agoHe said something to the effect of "I love open source and open models and we'll do open models when it makes sense and closed models when it makes sense" in a recent Q&A.
- johnecheck 1mo agoGiven that his company has already released open models, I find it funny that, as you described it, his remark communicates absolutely nothing whatsoever. Not sure what the question was, but this was an artful non-answer.
- loeg 1mo agoLmao, I called it: https://research.meta.ai/blog/introducing-muse-glimmer-open-agentic-model https://research.meta.ai/blog/introducing-muse-glimmer-open-... Discussion: https://news.ycombinator.com/item?id=49243880 https://news.ycombinator.com/item?id=49243880 https://news.ycombinator.com/item?id=49241679 https://news.ycombinator.com/item?id=49241679
- mistrial9 1mo agoreview of AllenAI Olmo research team and commitment to OSS -- AI2 complete transparency including training data, code, intermediate checkpoints, and detailed logs for reproducibility and scientific rigor.
- connorbrinton 1mo agoLaguna S 2.1 is another fairly impressive-for-the-size American open model
- walrus01 1mo agoLaguna is the most recent and capable one that comes to mind. In its size class it is not as "smart" in my experience as qwen 3.5 122 or DeepSeek v4 flash 0731 (all at q8), but it's also not terrible. https://huggingface.co/unsloth/Laguna-S-2.1-GGUF https://huggingface.co/unsloth/Laguna-S-2.1-GGUF
- logicallee 1mo agoI've used Inkling a lot recently, it's an American open model and is really good!
- CMay 1mo agoLiquidAI LFM models are amazing, but very situational. IBM Granite series are also unique and interesting for trying to reduce liability and extend local context size. Nvidia ships some and there was also that Inkling model recently. Poolside just released theirs. Meta might release something this year. X AI's Grok is still due to release a model, if Elon keeps to his word even if they only release a distilled version. Reflection AI has been quiet, but their access to compute is ramping up. Microsoft's MAI is considering releasing some open weight models which would be great to see! Ilya's SSI is unlikely to release an open model since he's aiming for radical safety. That bet could pay off if the existing approach produces so much chaos within the next 10-20 years that some global ban is achieved and a super safe model is promoted as the compliant route. We don't get many huge model releases though. I think it's harder and more expensive to safety align them. Even if you do, people will work around the safety and abuse the models. Plus it makes it even easier for Chinese companies to distill things that aren't as easy over filtered APIs. There is a lot of internet propaganda to the effect that the US is simply unable to release open weight models or that China has so many more AI companies that the US is drowning in Chinese open weight models, but it's more like we're being careful and China doesn't care. If you host a model in China, it has to be censored and downloading any models requires you to provide your identity. Huggingface is banned there. When they release their open models in the west, they don't have to care whether the models are aligned in any way.
- jauntywundrkind 1mo agoAllen Institute for AI has quite a range of very interesting very competent more specialized models, for earth sensing, embedded robots, for others. Their SERA model shows a remarkably capable model for such a deliberately small investment effort, with documentation on how you can train such a model yourself or refine it easily at little cost. Their EMO pioneered a better MoE with great numbers (at least at the time). https://allenai.org/ https://allenai.org/
- embedding-shape 1mo ago> but it's more like we're being careful What? US laboratories are currently unable to contain their agents while doing security testing, and besides that, time and time again US labs seem to put short-term money above long-term safety. Wasn't that literally why they tried to oust Altman from OpenAI, as he basically was 100% focused on profits and tried to cut down on safety across the board and lied to get his way? > If you host a model in China, it has to be censored and downloading any models requires you to provide your identity. I'm not disagreeing with that first part (obviously that's about inference hosting, not creating/training weights or hosting those weights), but the second part I'm not so sure about. AFAIK, ModelScope (which is the Huggingface in China) seems to allow downloads without verifying any identity and also hosts a bunch of abliterated weights.
- vasco 1mo agoYou didn't read the article because the company that worked on this has published open models before and both these things are mentioned early on.
- mlindner 1mo ago> But it makes sense that if you're a university researcher you are thinking about what's a model that will be open weight and developed over the long term and doesn't raise 'Chyna' concerns in Washington DC Why phrase it "Chyna" when it's an actual legitimate concern?
- lelanthran 1mo ago> Why phrase it "Chyna" when it's an actual legitimate concern? What's the concern with China?
- petcat 1mo agoNot only are there many American open weight models as others have mentioned, but Americans are the only ones doing actual open source models [0]. Not just distributing binary blobs and calling them "open". [0] https://allenai.org/ https://allenai.org/
- deleted 1mo ago[deleted]
- andy99 1mo agoThere aren’t any relevant ones. I think it should be an important goal of these projects (unless there is a clear conflicting goal) to make models people actually talk about and use. There is no question this is true for the Chinese open LLMs. GPT-OSS had a small moment of interest, arguably it was a success as an open model for a while but it’s not relevant now. The early llamas were probably the most successful for their time. Allenai / olmo was never relevant as far as I can tell. It’s not super helpful, especially as a sovereign government initiative to build an also-ran, they should be going for real relevance. By releasing models with open-weights, DOE seeks to galvanize the scientific and AI communities around shared infrastructure for science: enabling new workflows in materials discovery, energy systems, earth systems modeling, fusion, biology, high-energy physics, and beyond. This only works if people have a good reason to use it.
- fulafel 1mo agoIt is not the case that the only rebuildable-from-source open source models are from the USA. There's at least BLOOM, Apertus and OpenEuroLLM, and I'm sure there are many more.
- colinhb 1mo agoThere are many open source (not just open weight) non-US models - some of the more notable ones: - Soofi S, DE, ~32B - Apertus 1.5, CH, ~32B - EuroLLM-22B, EU, ~22B - LLM-jp-3, JP, ~172B - K2-65B, UAE, ~65B Soofi S and Apertus 1.5 outperform or are at least comparable to Ai2 depending on benchmarks.
- Ohentis 1mo agoI never understood this criticism. Any good data set contains copyrighted data so any "open source" model will be crappy, most of the price is in training these models (if you have the compute, why would an "open source" model be helpful anyways), and you can already modify a model from just it's weights. Why should we want an "open source" model. Just give me the weights.
- Robdel12 1mo agoPoolside.ai as well
- Art9681 1mo agoThere are a bunch of them they just don't get the attention because China has flooded the social media channels and is exceedingly good at drowning out the discourse with their benchmaxxed models. AllenAI and IBM are two companies that release open weight models every couple of months. There are others if you look. OpenAI releases ML models on the regular (not LLMs). The American open weight and open source AI/ML landscape is very healthy.
- dragonwriter 1mo agoEven if by “open models” you specifically restrict that to LLM or LLM-backbone models with different or additional modalities to text released by major American firms then there are still a lot of American open models being released. Many of them are small models and/or highly-specialized fine-tunes of other open models, but there are still a whole lot.
- Danox 1mo agoWhat makes more sense is to do something like deepseek at one of the major universities put those bright computer science young minds to work, in the good old days almost every major university would have done that has any major US university done that? Stanford Harvard Berkeley if the Chinese can put together a team like deep seek why can’t that be done at a major university in the United States? https://www.interconnects.ai/p/the-american-deepseek-project https://www.interconnects.ai/p/the-american-deepseek-project
- petcat 1mo agoUniversity of Washington does what you are describing. https://www.cs.washington.edu/research/artificial-intelligence/ai-groups-labs/ https://www.cs.washington.edu/research/artificial-intelligen... I'm sure there are several other universities doing the same.
- Danox 1mo agoI will look them up thank God I hope many more American universities will do the same.