8 ms·
Rio de Janeiro's city government model Rio3.5 beats Qwen3.7 in recent benchmarks
- deleted 3mo ago[deleted]
- deleted 3mo ago[deleted]
- VoidWhisperer 3mo agohttps://github.com/nex-agi/Nex-N2/issues/4 https://github.com/nex-agi/Nex-N2/issues/4 Seems that they didn't make/train a new novel model, they did a mix of two existing models and then gave it an instruction to say it was 'Rio, trained by Rio AI Labs'
- w4yai 3mo ago> The model is built via a merge of https://huggingface.co/nex-agi/Nex-N2-Pro https://huggingface.co/nex-agi/Nex-N2-Pro and https://huggingface.co/Qwen/Qwen3.5-397B-A17B https://huggingface.co/Qwen/Qwen3.5-397B-A17B, proceeded by On-Policy Distillation from a stronger model. We detected an incorrect upload in the previous version, where the base merged version was upload instead of the final distilled model. We are sorry for the confusion and apologize profusely. https://huggingface.co/prefeitura-rio/Rio-3.5-Open-397B/commit/a778c1ec4e21180ee55c3ea016a348e549e75f09 https://huggingface.co/prefeitura-rio/Rio-3.5-Open-397B/comm...
- daquisu 3mo agoIt was a recent edit though. Yesterday snapshot: https://web.archive.org/web/20260613072958/https://huggingface.co/prefeitura-rio/Rio-3.5-Open-397B https://web.archive.org/web/20260613072958/https://huggingfa...
- giancarlostoro 3mo agoHow does that contradict that they uploaded the wrong model?
- danieldrehmer 3mo agocan you offer a 4-bit quantized version and name it Zé Pequeno, pretty please?
- scotty79 3mo agoI'd love to see people figuring out how to build models from several smaller ones. We could then train small specialized models and deploy setups more optimized for any given task. Modular LLMs should be a thing.
- giancarlostoro 3mo agoThis is something I've been trying to figure out for a bit, some models are really good at instructions, but their context window is too small, I do wonder if having a cluster of smaller models would be feasible. Been building a custom coding harness so once its nice and polished I might experiment with this more.
- urbnspacecowboy 3mo agoSee discussion: https://news.ycombinator.com/item?id=48528371 https://news.ycombinator.com/item?id=48528371
- pixel_popping 3mo agoTo be fair, I still find it to be a great initiative.
- xbar 3mo agoSexy.
- mrandish 3mo ago> Rio de Janeiro's city government model... Because... lack of a good open weight LLM is a pressing need high on the municipal priorities list for Rio de Janeiro citizens?
- true_religion 3mo agoShould governments not take actions that later benefit the academic, scientific, and economic welfare of their constituents? Or is it that it’s a city doing this? Now Brazil does know how to boondoggle its finances for a prestigious cause with little return (e.g. the Olympics games) but this is far smaller a cost, more akin to a city setting up a tech accelerator or making a media campaign about how important STEM is.
- senorrib 3mo agoIt's the municipal IT company, and the dude that did this is a volunteer.
- pelasaco 3mo agoThe Taubaté LLM Hoax https://en.wikipedia.org/wiki/Taubat%C3%A9_pregnancy_hoax https://en.wikipedia.org/wiki/Taubat%C3%A9_pregnancy_hoax
- betimsl 3mo agoThe problem with these is the tool calling. From my experiments qwen agent almost always fails with tool calling and porting the correct config is quite tedious. Rio3.5 with Qwen compatible tool calling, we need that :)
- dizhn 3mo agoMr Erdoğan launched and initiative yesterday to become the leader in the AI space. As absurd of a claim as his 2023 (hard) landing on the moon.
- mettamage 3mo agohttps://xcancel.com/ZenMagnets/status/2065796012820848699 https://xcancel.com/ZenMagnets/status/2065796012820848699 Correct me if I'm wrong but reading through the comments of the thread this seems to be post training/fine tuning.
- oceansky 3mo agoYes. It's post training in qwen using the novel SwiReasoning framework.
- hedgehog 3mo agoI hadn't seen SwiReasoning (https://swireasoning.github.io https://swireasoning.github.io, paper and code), it looks like that works at generation time without any requirements on the model. It increases token-efficiency and accuracy, but at first skim it seems like this would be incompatible with multi-token prediction. For large reductions in token budget it could be worth it.
- rafaquintanilha 3mo agoDoesn't look like it's incompatible. Someone already released a quantization using MTP: https://huggingface.co/foxipanda/Rio-3.5-Open-397B-GGUF https://huggingface.co/foxipanda/Rio-3.5-Open-397B-GGUF
- hedgehog 3mo agoAs I understand it the basic premise of all the speculative decoding schemes is that the logits on the draft don't need to be exact so long as you mostly sample the same tokens, and because each position is fed by the embedding associated with the previous position's token you sort of "round away" error. With SwiReasoning I think you skip the sampling/rounding part and do something continuous using the whole distribution, so it would seem to rely on the accuracy of those values. MTP still makes sense outside the latent reasoning chunks though.
- 3mo ago
- cuzezzzbbfofai 3mo ago[flagged]
- blahblaher 3mo agoyes, let's instead trust a bunch of billionaires, that "for sure" have your and all of our interests at heart. And no, the "invisible hand" does not exist, it's the Epstein class hand, you just don't see it
- atoav 3mo agoA government ideally is a representation of the democratically chosen will of the people. If it is not, work towards making it so. IMO wherever someone says "the government" we should mentally substitute "we all, collectively". But a specific type of person appears to labour under the illusion that somehow we can get by without we all collectively steering our direction and choosing people who do what needs to be done without commercial interest. Their idea is that instead of choosing people who do it, we just make them compete for who can squeeze the most profit out of dealing with a problem and "somehow" that leads to a better result. When you press them for the details on that part of the mechanism, you will usually get crickets.
- latency-guy2 3mo ago"we all" is wrong, always. You do not agree with me. You can't claim to have my interests or my will if you are against it.
- atoav 3mo agoYes? With sufficient pedantic spirit anything can be argued against. This is what you're doing. So to give a counter-example: You drive with three friends in a car. You ask them: "Do we all want to go to MC Donald's?" Explain how it is wrong and why it would be. If it is always wrong it follows it has to be wrong here too. The answer is that the meaning of "we all" is context dependent and that friend of yours that argues that we all somehow includes people in the whole city is an oddball that doesn't pick up the context within the words have been said. We can all go around and make each others day worse with deliberate pedantry by ignoring the context of words, but that is basically just a waste of human energy. If you disagree with the fundamental point I made, argue against it based on the merits of the idea instead of arguing semantics.
- HeliumHydride 3mo agohttps://www.reddit.com/r/LocalLLaMA/comments/1u4fzg1/new_model_on_huggingface/ https://www.reddit.com/r/LocalLLaMA/comments/1u4fzg1/new_mod... https://x.com/SemiAnalysis_/status/2065894494935933191 https://x.com/SemiAnalysis_/status/2065894494935933191
- ramon156 3mo agoEvery day I'm reminded why I don't spend time on twitter. What use does it have to claim "X is better than Y in benchmark Z, disagreeing with that means disagreeing with me" Information is power, dick measurements are not.
- reed1234 3mo agoNo, I love twitter— and you are wrong.
- itsthecourier 3mo agomy length is a valid data point for the sake of science
- adrian_b 3mo ago> Post-trained from Qwen 3.5 397B Model Card: https://huggingface.co/prefeitura-rio/Rio-3.5-Open-397B https://huggingface.co/prefeitura-rio/Rio-3.5-Open-397B
- hmokiguess 3mo agoNever let them know your next move
- deleted 3mo ago[deleted]
- Aurornis 3mo agoA city government funding a fine-tune of a model is interesting. As for the benchmarks: If you spend any time playing with fine tunes of published models you know that benchmarks are gamed so much that they're a useless indicator of performance for models from small teams. It's too easy to fine tune a model to perform well on the benchmarks, release it, put a line on your resume saying you released a model that beat the major labs on benchmarks, and then try to use that to jump into a new job. The temptation is high. There are a lot of fringe models and fine tunes that claim to have better performance on some benchmark. Then you try to use them and find they're often worse at general tasks than the base model. I would wait and see if these results hold across other benchmarks. It's cool that the city is doing something with AI, but this is something where extraordinary claims require extraordinary evidence. I doubt a small, previously unknown team has unlocked something secret that the team who made Qwen couldn't figure out. It's more likely it was fine tuned for a specific outcome (possibly these benchmarks) and performance in other areas was reduced as a consequence.
- embedding-shape 3mo agoIndeed, this is all very true, I'd say it's true for the larger teams too, the entire ecosystem is so gamed by now that if you don't have your own private benchmarks with private test cases you haven't shared publicly, it's almost impossible to get a fair picture how well a model works, unless you actually sit down and use it.
- marcosdumay 3mo ago> A city government funding a fine-tune of a model is interesting. Looks like it's an IT services government-owned company. Most likely, they saw some business opportunity on selling it around for cities.
- arjie 3mo agoBenchmaxxing is the new “have a crypto trading strategy”. No one is impressed by it except non practitioners.
- kruxigt 3mo ago[dead]