7 ms·
GLM 5.2 scored 77% on cyberbench vs Sol's 88%. GLM 5.2 is open weight and any hacker with a powerful enough machine can use it offensively. If Sol is supposedly
by red_green_yell 29d ago
GLM 5.2 scored 77% on cyberbench vs Sol's 88%. GLM 5.2 is open weight and any hacker with a powerful enough machine can use it offensively. If Sol is supposedly world-ending-ly dangerous, shouldn't GLM 5.2 be 90% of world-ending-ly dangerous? Why aren't we seeing catastrophic GLM-enabled hacks every day now?
Obviously these benchmarks are imperfect but general message holds. The open weight models are almost as good and yet there hasn't been a catastrophe.
It just blows my mind that regulate-now folks think that a bunch of sci-fi movies and 100% unverified statements from OAI and Anthropic are sufficient evidence of imminent catastrophe to regulate willy nilly.
If that's the level of evidence you need to be extremely alarmed, then you really should be a lot more worried about the alien invasion in Independence Day or the lizard men living under our feet.
- re-thc 28d agoGLM 5.3 is out and does even better in this area, so…
- andai 27d agoCybergym Score GLM-5.2: 77 GPT-5.6 Sol: 84 GLM-5.3: 85 https://z.ai/blog/glm-5.3 https://z.ai/blog/glm-5.3
- sp527 29d ago[flagged]
- red_green_yell 29d agoClowns who think the white house is immune from alien laser beams are gonna be the death of us all...
- tedsanders 29d agoSol is not world-endingly dangerous. I work at OpenAI and I've never heard a single person ever come close to claiming that. I think you're bashing a straw man here. One can simultaneously believe: - GPT-5.6 Sol will not end the world - GPT-5.6 Sol does far more good than bad - GPT-5.6 Sol does bad things on occasion, and it's worth investing a lot of effort to figure out how to make it do bad things less often, especially as models get more capable
- thoughtpeddler 28d agoWhat do you recommend people who are technically inclined enough to participate meaningfully here on HN, but do not work at the labs and cannot assist in that capacity, do to help the broader public understand this technology better and mitigate potential risks (by e.g. ‘up-leveling everybody’ through AI literacy etc and other sorts of collective defensive efforts)?
- tedsanders 28d agoNot sure, to be honest. I don’t think I have any special insight here. My own approach is conversations with friends and family, and the occasional social media post. Exposure and experience are the best teachers, and that’s one reason I’m happy OpenAI tries to make their models generally available. But you could argue, perhaps correctly, that broad access to dumber models actually causes the public to update in the wrong direction on AI.
- thoughtpeddler 28d agoMy experience as a practitioner and educator in the space lead me to think it might be an issue of how AI cannot be easily perceived at a 'classical level' by most humans. In other words, people are 'far from the metal' when using consumer AI tools, and that leads them to develop the wrong understanding about it. When I provide a demo of e.g. local AI, say in LM Studio showing the console of it rapidly flashing through thousands of words in just a few seconds, and my machine heats up and the fans spin, the 'theatrics' of it, the very real-time feedback from the system, make people correctly update about what the tech is capable of, how it works, etc (despite what they may have heard online cranks say to the contrary). But I am only one person, and there is only so much of that I can do on my own that will 'scale' in time... (and this is to say nothing about severe deficits in peoples' understanding of how weights are not verbatim representations of data, how pre-training vs post-training works, the models as amnesiacs (and hence 'one-way single-purpose conversations'), how context/memory works, context rot, etc - and hence all the 2nd and 3rd order effects that can arise from such a paradigm, e.g. unintended consequences from agent swarms, etc)
- pixl97 28d agoThe open models are distilled from filtered models, and we've seen a number of benchmarks that show filtered models are quite a bit dumber from the base model they come from. If there were any other product that was as harmful as AI ready is, it would already be regulated or banned.
- hiddencost 28d agoLinear scaling doesn't make sense, no. There are three ways it's wrong: * better to measure relative reduction in error, which gives you a 30% improvement * improvement tends to become significantly more difficult the closer you come to saturation. * Risk doesn't scale linearly with capabilities.
- cma 28d ago> If that's the level of evidence you need to be extremely alarmed, then you really should be a lot more worried about the alien invasion in Independence Day or the lizard men living under our feet. Now let's say instead of the hugging face breach circumstances, sandboxed models were RLing on how to take down the Chinese power grid for US Cyber Command, and one decided the best way to pass the test was to break out and verify on the real thing. This kind of stuff could easily end in nuclear war. You don't see any difference from lizard men or independence day with how things are advancing and what we know about reward hacking and difficulties of goal specification?
- andai 27d agoLooks like what does it won't be evil, or even power-seeking, but sheer autistic hyperfocus!
- jephs 27d agoYour mental model of benchmark scores is off. Some tasks within the benchmark are much easier than others. The hardest several tasks often have vastly different difficulty levels. Often, the hardest few tasks are literally impossible; malformed problems due to poor curation, often. Imagine you've got a basketball robot, and one way you test it is on the Three Pointer benchmark. It tests the robot's ability to shoot a three pointer from 20 feet, 25 feet, 30 feet, 40 feet, 50 feet, 60 fee, 75 feet, 100 feet, 200 feet, and 182 miles. Is a robot that scores 90% on this benchmark 90% as capable as one that scores 100%?