6 ms·
America needs its own DeepSeek or Z.ai, a lot of people (myself included) root for open chinese models to win because they have no other choice. Thinking Machi
by ls_stats 2mo ago
America needs its own DeepSeek or Z.ai, a lot of people (myself included) root for open chinese models to win because they have no other choice.
Thinking Machines might be it.
- soundworlds 2mo agoAllenAI is also one to keep your eye on. Founded by Paul Allen of Microsoft, they are one of the best teams working towards truly transparent / open AI (including training data)
- maxloh 2mo agoI love Allen AI. I find it wonderful that, as a non-profit, they are only one to two years behind SOTA models that cost billions of dollars to build, if not more.
- FrankBooth 2mo agoI love the tasteful thickness of Paul Allen’s model.
- nl 2mo agoAllenAI is great, but they don't have the budget or remit to build large models.
- timmg 2mo agoI wonder if the recent sale of the Seahawks will change that. IIRC, ~$10B and all is supposed to go to charity. Not sure how much of that will go to AllenAI, though. (If any.)
- nl 2mo agoNo reason to think it will. Paul Allen died after donating the money to create AllenAI and I don't think there are any links. Hopefully it somehow works out though!
- ReptileMan 2mo agoI will wait for Modernist AI by Myhrvold.
- icase 2mo agoi refuse to root for our enemies, but otherwise you are correct.
- mstank 2mo agoDo you think American companies will secretly distill frontier models to build open weight ones?
- upmind 2mo agounlikely I think, they're likely doing this to garner some interest in their company but they seem pretty interested in revenue (judging by the companies they're working with)
- xnx 2mo ago> root for open chinese models to win What does "winning" mean to you?
- verdverm 2mo agoIts not as good as GLM 5.2 for agentic workflows while also being bigger. Competition is going to be ruthless because the super low cost to switching. There is also AllenAi in the US, but they have yet to produce a model at this scale. Thankfully, new contenders can come out of nowhere and do well, as long as they can produce a competitive model.
- InsideOutSanta 2mo ago> Its not as good as GLM 5.2 for agentic workflows while also being bigger GLM 5.2 underwent extensive post-training and iteration since its original release to reach its current state. This seems like an extremely strong model for a first release, with a lot of potential for improvement, just like DS4. Sometimes I wish Meta had stuck with Llama 4 a bit longer to see how much further it could be pushed.
- hirako2000 2mo agoLlama 4 wasn't deemed a success, and Meta pivoted away as its now former head of AI couldn't demonstrate, nor even showed interest in, business profit. They overspent on llama 3 anyway so money ran dry, LeCun is good at running research, but budgets didn't stretch. Meta isn't investing in frontier big models anymore.
- nl 2mo ago> Meta isn't investing in frontier big models anymore Yes they are. Meta Muse is their attempt. It's below frontier performance at the moment but they are spending on getting there.
- nl 2mo agoLlama 4 was a bad architecture. Meta Spark is moderately promising but of course closed source.
- verdverm 2mo agoThis is a great point
- gkapur 2mo agoIt could be but there are a host of companies going after open weights models: Arcee, Reflection, Llama (TBD on Meta's focus on closed-source versus open-source), etc. That said, the fine-tuning API + open weight model at least is a semblance of a viable business that could work so I will be curious about it. I'm not sure the synergy is fully there (why is someone with an open weights model privelaged to fine-tune it better if it's just QLora or Lora) but let's see!
- paxys 2mo agoLlama is dead. Meta is now releasing proprietary models (Muse Spark).
- anon373839 2mo agoThey’ve made some wishy-washy statements about their intention to release a future version of Muse Spark as open weights. We’ll see.
- andriy_koval 2mo ago> It could be but there are a host of companies going after open weights models: Arcee, Reflection, Llama (TBD on Meta's focus on closed-source versus open-source), etc. my bet is that Chinese government fund Chinese models way more compared to what those companies receive (except llama, which is outdated but was strong foundation at its time)
- deleted 2mo ago[deleted]
- bostonvaulter2 2mo agoWhat is the business model for an open weight model?
- deleted 2mo ago[deleted]
- matsur 2mo agoThinky has a potential answer in Tinker — give away the weights and charge for the SFT (and maybe RL down the line) to make the model more capable for specific tasks.
- andriy_koval 2mo agoSFT/RL can be done without parent company.
- alightsoul 2mo agoBut not conveniently. This is why you outsource to vendors. Not everyone will do it
- ergocoder 2mo agoThe same business model that Deepseek is using. Open-source models + services. This is more attractive because it doesn't lock in the vendors. If I grow larger, I can decide to deploy the open-source models.
- tyre 2mo agoSo they're constantly hemorrhaging their most valuable clients? Tech history is littered with the corpses of "open source but we sell hosting" services. Models are so expensive to train, you can't be losing the big clients once they get super profitable.
- MikeTheGreat 2mo ago
- joshmarlow 2mo agoI don't hear about them a lot but it looks like arcee.ai is aiming to be just that. Here are some of their current open weight offerings: https://www.arcee.ai/open-source-catalog https://www.arcee.ai/open-source-catalog
- wgd 2mo agoYou don't hear about them much because their models aren't really competitive. I really wanted to try Trinity Large as a daily-driver in the MiniMax M2 sort of niche but I couldn't make it through a single day. The models need another couple point releases worth of post-training to make useful agents and if memory serves they weren't any less slopped in writing style and those are really the only two things people look for in models.
- tonic_note 2mo agoisn't that what Reflection is trying to be?
- UncleOxidant 2mo agoHopefully they'll release some smaller models (<100B) that we can run on home hardware at faster than 10tok/s.
- fastball 2mo agoWhat about Meta?
- insane_dreamer 2mo agoIt’s what Meta was supposed to do but Llama fell of the wagon. There’s also Prism
- jauntywundrkind 2mo agoAlso the fact that China is building solar power like crazy: that makes it fantastically more well spirited an endeavor to wish well.
- codemog 2mo agoI’m trying to be charitable but your comment reads as “China bad” propaganda to me. Who cares that DeepSeek and Z.ai are Chinese companies?
- xenospn 2mo agoTry selling SaaS for finance (think Private Equity/Wall St type customers) that is powered by a Chinese model. See how far you get.
- maxloh 2mo agoLLMs aren't just for coding and math. Many people understand the world through LLMs, even when it comes to philosophy and politics. If you understand the world through a Chinese LLM, you are seeing it through a biased lens stemming from biased training data. (Also, in that way, having all major LLMs developed by the US carries a risk too. We need more diversity than just the viewpoints of the US or China.)
- alightsoul 2mo agoAll 6 UN languages should have their own dedicated LLM at the very least: American for English, Chinese for Chinese, LATAMgpt for Spanish, Russia has their own, Mistral for French, the Arabs don't have one yet I think, then there's sarvam for India and South Korea's sovereign models like SK telecom's
- nimchimpsky 2mo ago[dead]
- citrus1330 2mo agoyou live under a rock or what?
- nodja 2mo agoIt doesn't matter until it does. If the chinese government decides that open weight model releases are no longer allowed, that's a lot of companies that can't release new models. Same with the US government, etc. Having diversity is important.