8 ms·
Apertus – Open Foundation Model for Sovereign AI
- yreg 3mo agoprevious thread: https://news.ycombinator.com/item?id=45108401 https://news.ycombinator.com/item?id=45108401
- trvz 3mo agoThe previous version of this model has been pretty bad, but claimed to adhere to copyright laws. However, based on my testing, that's not true either. So in my view this is completely useless.
- embedding-shape 3mo agoAs long as the following remains true, this release ends up a bigger contribution to science at large than most other models trained "behind closed doors": > Fully open model: open weights + open data + full training details including all data and training recipes
- coder543 3mo agoIs a recipe useful if no one likes it? There are equally open, much more useful models out there: https://artificialanalysis.ai/?models=nvidia-nemotron-3-ultra-550b-a55b%2Cnvidia-nemotron-3-super-120b-a12b%2Colmo-3-1-32b-instruct%2Colmo-3-1-32b-think%2Ck2-think-v2%2Capertus-8b-instruct%2Capertus-70b-instruct#intelligence https://artificialanalysis.ai/?models=nvidia-nemotron-3-ultr...
- khalic 3mo agoNemotron still has partial closed data. Having multiple models to chose from is a good thing
- simonw 3mo agoIt uses fineweb, which is derived from Common Crawl, which is an unlicensed scrape of web pages.
- reedciccio 3mo agoYou don't need a license to scrape the public web and analyze it, turn it into tokens and other transformations. Let's not expand copyright beyond the horrible monster it already is.
- simonw 3mo agoI think it's likely that US law will continue to find training on scraped, unlicensed data to be legal. That doesn't mean much to the many people I know of who refuse to use a technology that they see as being unethically created using the work of others without compensating them. I continue to hope that someone will train a "vegan" model on licensed or out-of-copyright data so those people can experience the benefits of this class of technology. (I compare them to vegans because, like vegans, I think their ethical position is credible and has merit even though I do not choose the same ethical framework for myself.)
- EnergyAmy 3mo agoThis is as ethical as it gets. They're getting compensated by being able to use the result of their work freely. This is the rising tide that lifts all boats.
- simonw 3mo agoGood luck convincing the training data licensing holdouts of that.
- mwcampbell 3mo agoWell, the frontier models aren't freely available in the same way that the training data (the public web) is; they're only available as a limited no-cost tier of a paid service. There are models where the weights are released and you can run them locally, but one counter-argument I've seen is that these aren't really the models that people are excited about and making extraordinary claims about.
- markhahn 3mo agoI'm curious how you test; could you explain? Do you have a set of factoids that should be subject to copyright, but are somehow literally (whole work) generated by the model in question?
- dofm 3mo agoSo far the smallest model I have actually seen behave in a way that feels consistent with the contemporary LLM chat experience is Gemma 4 12B. (The QAT build particularly). The E4B model is not bad — it has a good conversational flow, it responds well if nudged — but the 12B model feels capable. Nothing below that really seems to be good for anything other than training for specific tasks. I have not been impressed by the earlier Apertus 8B model, which doesn't feel like it really responds to nudges. I am a strong believer in smaller models, so I might try one of these out of curiosity to see if it might do useful things in limited contexts.
- throwaw12 3mo agoLooks like their instruct models are Llama3.1 fine tune from last year. Is there any progress on new models? My last hope for soverign AI is from Chinese open models
- kordlessagain 3mo agoSovereign AI is not about using just one model. It's about using the right model for the right job, and getting them to talk through the solution TOGETHER before presenting the answer. If you want to mix models like this, check out https://github.com/deepbluedynamics/nemesis8 https://github.com/deepbluedynamics/nemesis8
- atemerev 3mo agoI use it extensively. It is not ready for agentic use, but as a generic driving model for RAG use cases, it is pretty competent. You can build useful software with it.
- _pdp_ 3mo agoI want to believe.
- maxloh 3mo agoOther fully open LLMs include Allen AI's OLMo 3.1 and MBZUAI's K2 Think V2, both of which have released their full training pipelines and datasets. Nvidia Nemotron is also an open training source model, though a portion of its dataset remains proprietary. Quoting lambda's comment: > Note that the Nemotron models are generally stronger than Olmo and K2 Think V2 (according to Artificial Analysis benchmarks), and there is a lot of overlap in their datasets (lots of datasets are based on the same sources with different filtering, Olmo and K2 Think V2 both have used some Nemotron datasets). > But yeah, Nemotron is a modern and fairly capable LLM, even the 122b is more capable than Deepseek R1 (a 671b model) on most benchmarks, and there's also the recently released 550b Ultra now. https://news.ycombinator.com/item?id=48492439 https://news.ycombinator.com/item?id=48492439
- vcryan 3mo agoMaybe I'll give Nemotron another try. Yesterday I used the latest one on OpenRouter and it was bad - worse than StepFun
- soundworlds 3mo agoAllen AI do not get enough love. They are doing GenAI how it should have always been done. In fact, if the frontier companies had taken their approach, it would have started much slower, but I think we would be far more advanced by 2035. Instead we have a majority of society that wants to see AI fail.
- AndrewKemendo 3mo agoFully agree with this and they were leading robotic learning as well even back to 2019. IsaacSim was (and might still be) the best robotic learning sim and I ran MLAgents.
- hit8run 3mo agoCare to elaborate on this?
- sawjet 3mo agoIs there any evidence that "a majority of society wants AI to fail" Or is it just vibes?
- pferde 3mo agoFor a model that claims to focus on many languages, it's quite unreliable when it comes to simple questions like "how to say X in language Y" or "how to conjugate verb X in language Y". It keeps hallucinating words that do not exist, and when corrected, it only hallucinates a new lie.
- 8note 3mo agoit probably doesnt know what language each set of words is referencing. i doubt they are including a lot of training data labeled with the language. "how to say X in language Y" is a different task from saying X in language Y
- einpoklum 3mo agoActually, it isn't all that different. There are only two words separating "how to say X in language Y" from "say X in language Y". And this "vulgar" metric is actually quite relevant for an LLM, which answers based on conversational context.
- SwellJoe 3mo agoI like the idea, and it has become more pressing that everyone outside the US think about tech sovereignty because the US has become an unsafe place to keep your data, but the impression I get from Apertus is that it moves at the speed of a committee. I have no expectation they'll deliver a competitive model. At least, not competitive with current models. Maybe competitive with models a year ago (though they haven't even done that yet, right?).
- nezuzen 3mo ago"the US has become an unsafe place to keep your data" I empathize with this but curious what would make any other country a better safehaven for your data? I personally like the EU's approach to data safeguards, but are there other locales/data protections you have in mind that would keep your data "safe".
- digitaltrees 3mo agoThe rule of law exists in other countries in a way it does not in the US right now.
- SubiculumCode 3mo agoCan you give examples?
- sscaryterry 3mo agoWow, this is a bit obtuse. It is a commonly accepted "fact" right now, outside the US, that the US is not to be trusted (right now), due to some orange guy, and his mates, manipulating markets, running their mouths, doing all kinds of criminal and/or infantile shit. I'd say there is quite a bit of evidence for this all around.
- SubiculumCode 3mo agoHardly obtuse. It's good to be specific when making broad claims. The graft of Trump is a big problem, IMO, but the claim was larger than that, as being something about America's system of Law and Justice, and I don't see these as being completely busted (yet) by the Orange Man
- maxloh 3mo agoGreat to see more fully open LLMs. I think a problem with open-weight models is that while you can improve them, you are not going to create the next generation of LLMs by fine-tuning. We are at the mercy of frontier labs for access to SOTA LLMs. For example, Anthropic recently started requiring identity verification for Claude [0], same for OpenAI [1]. If one day China's distillation labs stop releasing their LLMs as open-weight, I doubt American labs will continue to release free LLM weights without that competition. That's where fully open pipelines shine: they enable the community to create the next generation of SOTA LLMs. That is the only way LLMs truly become sovereign. [0]: https://news.ycombinator.com/item?id=48618455 https://news.ycombinator.com/item?id=48618455 [1]: https://news.ycombinator.com/item?id=48618606 https://news.ycombinator.com/item?id=48618606
- dofm 3mo ago> We are at the mercy of frontier labs for access to SOTA LLMs I disagree with this use of SOTA, and this topic is why. Anthropic and OpenAI have “cutting-edge” models. These are beyond the state of the art but they are closed, secretive, hard to quantify. The “state of the art” is open source, open weights models that can be inspected, studied, shared and critiqued, because that is what is meant by “the art” —- it is the knowledge and principles and evidence and materials available to all. The “state of the art” is the highest point of that. I wish we could make this distinction and stop blessing two secretive, unverifiable loss-making companies with so much power. (Putting that aside, I suspect — without evidence, mind you - that the endless march to solving models by making them bigger is not the solution anyway.)
- sockaddr 3mo agoSorry but I think you’re requirement that something only be “the art” if any arbitrary person can critique it is off. The frontier labs are working on the state of the art but it’s just art that you aren’t allowed to see. Unfortunately.
- dofm 3mo agoIt is work using the principles of the art, obviously. But "state of the art" implies the highest state of general availability, not just in terms of access to some product, but of use of the ideas, concepts, methodologies etc. Anthropic and OpenAI have "cutting edge" models; the state of the art is behind the cutting edge. The state of the art is the best open source, open weights model available. More or less by definition. I am probably tilting at windmills here.
- dTal 3mo agoIt's good that there is a movement for open LLMs, but it's not where the battleground is right now. The battleground is local vs service LLMs, and we are losing that battle badly despite all the software being here now and viable, entirely because UX sucks. How many normal people do you know who use "ChatGPT"? A lot, probably. How many even know what "Gemma" is, let alone have downloaded llama.cpp, a GGUF file from Hugginface, and run "llama-server" from a text console with all the correct command arguments? How many are thinking about this use case when speccing out their next computer? Where is the breathless marketing copy boasting x tok/s? We are sleepwalking into slavery.
- azinman2 3mo ago> We are sleepwalking into slavery. That’s a bit hyperbolic…
- MrDrMcCoy 3mo agoSome hyperbole is useful. The problem is real and serious, though short of the specific verbiage.
- double0jimb0 3mo agoYea, anyone who understands what makes products actually usable is opting to get paid for said skill.
- idiotsecant 3mo agoBetter UX does not buy you a datacenter farm to train state of the art cutting edge models. Right now the only people who can do that are the technobility class.
- dTal 3mo agoIt does not, but it might encourage more people to care. Worrying about training is a luxury when you are starting from a baseline of "OpenAI spies upon me and controls my access". Let's focus on getting every Tom, Dick and Harry 1) on board with LLMs, because they're happening, 2) habitually using local software.
- mrshu 3mo agoBy far the most impactful product of the Apretus project are the people. To quote a memorable line from Dominique Paul (https://www.thisiscrispin.com/ https://www.thisiscrispin.com/): > What most people miss IMO is that this is not a team who is doing this for the fourth time like virtually any other LLM provider and who could learn from its own past experiences. I bet if the team would do another model training they could get way better results at one fourth of the costs.
- reconnecting 3mo agoA chat interface where you can try Apertus: https://chat.publicai.co https://chat.publicai.co
- einpoklum 3mo agoYou will need to register with an email and password though, i.e. your sessions will be recorded and identified. Also even after you do that, and start a chat, you currently get: "JSON.parse: unexpected character at line 1 column 1 of the JSON data" so it's not quite there yet.
- jawns 3mo agoI am curious about how opt-outs and PII removal work. Who confirms those requests are legit?
- dangoodmanUT 3mo agoHow are they going to be competitive with top models at 70B size?
- kennywinker 3mo agoQwen et al shows size isn’t actually the only useful metric for an llm.
- Ainaguade 3mo ago[dead]
- neom 3mo agoI'm curious to know what stuff like this means for cohere? Their whole value prop is Sovereign AI. It seems they spent a lot of money developing models but own none of their own infra, what is the point of a country spending a lot of money on coheres solutions when stuff like this is becoming increasingly available and usable? Feels like I must be missing something here??
- markab21 3mo agoI'm mildly surprised that more people aren't using Nemo models for this reason. We've moved most of our processing to a combination of Nemo Ultra and Super, with some support for multi-model-specific tasks on Omni. The setup is working REALLY well for us, and I'm comfortable with the more measured pace of improvements. We work with many long-context problems, and the ecosystem is great. There were a number of use cases where we needed to use Gemini (audio modality), and Ultra has been a VERY cost-effective alternative once we got through the nuances.
- khalic 3mo ago[dead]
- focusgroup0 3mo ago[dead]
- holistio 3mo agoKnowledge cutoff is March 2024. Incredible.
- uberex 3mo agoDoes anyone care about this anymore with context windows and tool harnesses.
- nisten 3mo agoAs an opesource AI researcher with a lot of models and datasets on huggingface I am very appreciative of these types of project but we are ignoring the elephant in the room here ( or lack of ) the swiss have no gpus
- kennywinker 3mo agoHow is this a real problem? Genuine question, because i don’t really understand the urgency of everyone buying up ram and gpus as prices for those skyrocket. I can run the 8B version of this swiss-ai model on a ten year old GPU. For the larger one, $2000 consumer hardware can run it fine. Beyond that, there are plenty of places where time on a GPU can be rented, and if the model is good, there will be hardware to run it.
- pu_pe 3mo agoYou can run it, but you can't train it. While this type of toy model could actually be trained in Swiss equipment, a state-of-the-art LLM probably could not. My charitable reading of GP's point is that the bottleneck for true compute sovereignty is the chips, not the models.
- markhahn 3mo agowhy do you say the Swiss have no gpus?
- T-A 3mo agothe Apertus model was trained on the Alps supercomputer, operational at CSCS since September 2024, a data center of over 10'000 top-of-the-line NVIDIA Grace-Hopper chips https://log.alets.ch/110/ https://log.alets.ch/110/
- khalic 3mo agoDo some research before posting that kind of stuff
- david_shi 3mo agoThese models don't seem very competitive, who's their target audience?
- poplarsol 3mo agoEuropeans who fetishize "compliance".
- 3997531578 3mo ago[dead]
- markhahn 3mo agoresidents of the universe who recognize the US as a supply-chain risk. no, actually, from the docs it sounds mainly motivated by the country's unique linguistic requirements.
- zitterbewegung 3mo agoSort of interesting license not sure if anyone will do it long term. The training data and the Apertus LLM may contain or generate information that directly or indirectly refers to an identifiable individual (Personal Data). You process Personal Data as independent controller in accordance with applicable data protection law. SNAI will regularly provide a file with hash values for download which you can apply as an output filter to your use of our Apertus LLM. The file reflects data protection deletion requests which have been addressed to SNAI as the developer of the Apertus LLM. It allows you to remove Personal Data contained in the model output. We strongly advise downloading and applying this output filter from SNAI every six months following the release of the model.
- JSR_FDED 3mo agoFrom a sovereign AI perspective, how does this compare to Mistral?
- pizlonator 3mo ago> compliant at scale The jokes write themselves.
- uberex 3mo agoBeing childish I https://oss.zuericitygpt.ch/?q=hello+talk+like+a+pirate https://oss.zuericitygpt.ch/?q=hello+talk+like+a+pirate
- yashthakker 3mo ago[flagged]
- wg0 3mo agoYou might dismiss it as nothing but the Linux analogy does not work here either. It is more than that and direct threat to commercial AI labs and their business model. These labs are milking bunch of foundational papers for years now and the end is near. Going forward would be such open source, open data and open recipe models possibly someday even with the training being crowd sourced if not inference like the BitTorrent model. Lastly, even Chinese models (GLM, Deepseek, MiMax) work really really good and any user would testify that they do not miss OpenAI/Anthropic/Gemini at all if they're using those Chinese models which is argument enough that with such models, no one is going to miss Chinese models as well.
- naklitechie 3mo agoWhat's the community's take on Sovereign AI being funded by states around the world? Why the emphasis on sovereign? Open is good enough. No?
- luplex 3mo agoSovereignty is a political buzzword. From the political point of view, you want your country to be as independent as possible. This means you need the capabilities to build and deploy good AI models. Initiatives like this are more about capability-building and less about LLM-building. Why do we need capabilities in Europe? Because Trump and Xi can't be trusted to keep providing us with new frontier models in the next years.
- khalic 3mo agoIt was in reaction to the possible threat of main actors restricting use. The latest US gov stunt with Fable just made it concrete and pressing.
- andrewshadura 3mo agoNot to be confused with Apertium and Apertis.
- runnig 3mo ago[dead]
- jocelyner 3mo ago[flagged]
- firstrowraver 3mo agoapertvs.ai? seriously?
- iamyemeth 3mo ago> Conclusion There are 2 r's in the word "strawberry". Not looking good so far
- sigmoid10 3mo agoI guess they still use a tokenizer? Why would this kind of issue be solved? The model fundamentally can't see the word character by character like you do. For o200k tokenizers for example, what the model sees are 3 tokens: [302, 1618, 19772]. These are shown to you as ["st", "raw", "berry"] in the UI. The only way any model can infer individual characters is by using external tools or implicit knowledge picked up during training or (what many of the big labs apparently do) special training for these edge cases that fail once the next special case comes along.
- Bobaso 3mo agoApertus V1 performance were sub-par. The Team is working on v2 ATM. Looking forward to testing it.
- khalic 3mo agoI don't know, I'm implementing a translation system right now, and Apertus is very good for the model size. I wished they added some chain of thought training to increase precision and context understanding.
- flixspiek 3mo ago[flagged]
- rcdwealth 3mo ago[dead]
- chris_explicare 3mo ago[flagged]