28 ms·
It's so refreshing to see DeepSeek's tech report[1] full of juicy details; meanwhile, something like Fable's system card[2] is like 70% "safety", 10% "model wel
by kouteiheika 7d ago
It's so refreshing to see DeepSeek's tech report[1] full of juicy details; meanwhile, something like Fable's system card[2] is like 70% "safety", 10% "model welfare" to make sure little Claude isn't distressed, and 20% benchmark numbers.
[1]: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/DeepSeek_V41_Tech_Report.pdf https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/...
[2]: https://www.anthropic.com/claude-fable-5-1-mythos-5-1-system-card https://www.anthropic.com/claude-fable-5-1-mythos-5-1-system...
- bbor 7d ago…are you sure a brave stance against safety and welfare is what we need in this moment? Why do you think your conception of the dangers are more accurate than all the scientists who have spent their lives studying this?
- nozzlegear 7d agoModel welfare is wishy washy bullshit. It's software, it doesn't have feelings. > Why do you think your conception of the dangers are more accurate than all the scientists who have spent their lives studying this? Do the Chinese have no such scientists?
- VulgarExigency 6d agoAlas, the Chinese scientists have not read Harry Potter fanfiction, and thus their minds are inundated with cognitive biases
- jbs789 7d agoBias…
- 10000truths 7d agoBecause safety and welfare have literally nothing to do with LLMs. They generate text. If someone is stupid enough to hook the text generator up to nuclear missile launchers and try to "align" it against nuclear annihilation with a "pretty please don't do that" prompt, I'm not going to blame the AI for the impending nuclear apocalypse, I'm going to blame the idiot who handed the big red button to the digital equivalent of a toddler.
- zith 7d agoWell, giving it access to a simple linux terminal is theoretically enough to cause more damage than most people are comfortable with, and doing so is trivial enough that it will be done (and has been, tens of thousands of times).
- flexagoon 6d agoShould we also morally align the Linux terminal then?
- lemonfever 7d agoWhat if LLMs completely unrelated to the nuclear missile ecosystem autonomously hack their way in (maybe with sophisticated social engineering)?
- mrtesthah 7d agoReplace LLMs with APTs in that sentence,
- Certhas 7d agoHumans are biological machines that generate further humans. Lawyers and diplomats and politicians and bureaucrats are humans, that only generate text. We are seeing LLMs have cognitive abilities that significantly exceed human abilities. At the same time, they are clearly not the same type of mind that humans are. They are something new. I think the widespread "they are just text generators" and "they are just tools" are comforting lies rather than an honest look at what we are seeing right now. Intellectually lazy. And by the way, there has been a long-standing consensus among ethicists, philosophers, and sociologists that technology is not value-neutral [1]. Of course Silicon Valley has a long-standing tradition of denying this. [1] For example Footnote 1 in https://www.jstor.org/stable/27106634 https://www.jstor.org/stable/27106634 or https://plato.stanford.edu/entries/technology/#EthiTech https://plato.stanford.edu/entries/technology/#EthiTech
- graemep 6d ago
- alchemist1e9 7d agokeep me safe big brother
- 15155 7d agoThis is known as an "appeal to authority." "Scientists" and "their lives" are doing a lot of work here.
- frotaur 7d agoIt is a fact that among experts there is no consensus on saying '(super)intelligence is broadly safe and easy to control'. There might even be a consensus forming on the opposite claim. Regardless, why would there be no scientific consensus if the question was easy and clear cut? I think the easiest reason is that these are hard questions to answer.
- bbor 6d agoYes, it's called expert epistemology, it's the basis of your entire life. Or do you do your own safety checks of every airplane you get on? Do you do your own research rather than trusting doctors? Do you think climate change doesn't exist because the reason we think it exists is because experts say it does, despite the fact that it snows sometimes?
- 15155 5d agoA valid appeal to authority is normally accompanied by a specific expert's name or working group rather than some abstract "scientists." Also, these appeals to authority normally cite an expert in a field that has an existence exceeding 3 years. Saying people spent "their lives" on fledgling technology is intellectually dishonest.. Are these "scientists" 22 years old? I'm sure you'll snipe back: "ALICE!!!" I couldn't care less about these completely irrelevant approaches. The other issue is the venue these appeals are being made in. The people who work on this technology are actually here, commenting. This is like walking into a medical symposium and citing "doctors say" as if it were a valid way to shut down discussion amongst the people who wrote the textbooks.
- kouteiheika 7d agoExcuse me for not being interested in over 100 pages of how well the model can refuse and block my requests, especially considering how fun it is to waste my time trying to get around those restrictions when they inevitably trigger because the clanker thinks that I'm doing something naughty, all the while it can't reliably center the proverbial div without doing something stupid itself.
- walrus01 7d agoMeanwhile I have an uncensored qwen 3.8 27B here that will happily attempt to (as a crude and randomly chosen sampling of bad/evil things) give me the recipes for meth, how to make an IED, write a manifesto in support of a horrible ideology, or commit various forms of fraud. Now I certainly wouldn't recommend that anyone try to follow what it says to do, because it's almost certainly very wrong on key parts that would put its users in federal prison for the rest of their lives. There's uncensored models out there which score 0 (zero refusals) on this "harmful behavior" dataset: https://huggingface.co/datasets/mlabonne/harmful_behaviors https://huggingface.co/datasets/mlabonne/harmful_behaviors
- kouteiheika 7d agoYep. Just like a kitchen knife will make no attempt to prevent me from stabbing anyone with it. Here's a dirty secret though -- you don't actually need an abliterated/uncensored version of the model to get it to do this. I can do this with every and each open weight model, as served from OpenRouter, using vanilla model weights.
- walrus01 7d agoA little bit like Neal Stephenson's metaphor of unix-like OSes as the "hole hawg" of operating systems. In the sense that there's very little preventing you from doing something like "sudo dd if=/dev/zero of=/dev/sda bs=1M" or running rm -rf on your homedir. http://www.team.net/mjb/hawg.html http://www.team.net/mjb/hawg.html If I recall right this was written around the same time as Cryptonomicon 25+ years ago.
- 7d ago
- swiftcoder 7d ago> scientists who have spent their lives studying this Please point me to one actual accredited scientist who has spent a lifetime studying AI alignment? Pretty much this whole field is only 5 years old
- adamzenith 7d agoThe field is much older, MIRI is ~20 years old. Look up Eliezer Yudkowsky.
- swiftcoder 7d agoThe field was purely theoretical 20 years ago, and Yudkowsky is pretty much the dictionary definition of "not accredited"
- naishoya 6d agosome use "not accredited" as a pejorative term. Lets not forget that the 'Fermat's Last Theorem' which has been pretty visible for the non-math crowd of late due to the recent AI frenzy about a purported proof was but one small contribution to the world's math lexicon by someone with a bachelors degree in civil law, that George Green was a baker and millwright, Boole was the son of a poor shoemaker in England with no formal university education and left school at age 14. Oliver Heaviside was a telegraph operator, and Michael Faraday was an apprentice bookbinder. So, not accredited shouldn't really carry much weight when it comes to mathematics. Lets not pretend that machine learning and the narrow branch that is the current approach to LLM inductions is anything but applied math. We might exercise our own minds and actually read the works and writings of a person, and use that as a measure of knowledge and perspective. Not all PhD dissertations are equal, and many have comprehension and ability to move us forward even without the institutional rigour. For those who prefer to have easy access to citations, here are some relevant papers that are not "Harry Potter" related, some with coauthors from Oxford University. Cognitive Biases Potentially Affecting Judgment of Global Risks [https://intelligence.org/files/CognitiveBiases.pdf https://intelligence.org/files/CognitiveBiases.pdf] Levels of Organization in General Intelligence [https://intelligence.org/files/LOGI.pdf https://intelligence.org/files/LOGI.pdf] Corrigibility [https://intelligence.org/files/Corrigibility.pdf https://intelligence.org/files/Corrigibility.pdf] The Ethics of Artificial Intelligence [https://intelligence.org/files/EthicsofAI.pdf https://intelligence.org/files/EthicsofAI.pdf]
- cowl 7d agoAnthropic's stance on safety it's just PR management and their hope to keep the others down, they are rushing as blind as everyone else to whatever improvement they can achieve.
- bbor 6d agoInteresting stance. Out of curiousity, where did you do your doctoral research in AI or cognitive science? Where have you published your rebuttals to the overwhelming consensus?
- deleted 6d ago[deleted]
- SAI_Peregrinus 6d agoAI safety efforts from OpenAI and Anthropic are purely about brand safety.
- anthonyrstevens 6d agoThis is such an uncharitable (and, in my opinion, incorrect) take
- SXX 6d agoNope. AI safety efforts is part of their attempts at regulatory capture.
- windexh8er 6d ago> …are you sure a brave stance against safety and welfare is what we need in this moment? Is it out of convenience to not see the hypocrisy? "Safety and welfare" for you and me. Yet if you work at Anthropic or OAI, or are a partner of them then you can let it rip! Oh, and when they illegally do just that - you get a "we're sorry bro" blog post that's designed to drum up FOMO and, most importantly, zero accountability. Yet, if anyone else abuses a model in that same manner? Illegal! You're defending a very slippery slope here. Also, who do you think trained these models to have these capabilities? It sure as shit wasn't content that OAI or Anthropic had by default. Why should I trust them with these skills when they "have not spent their lives studying this"? Maybe start looking around before it's being used against you [0]. [0] https://www.gadgetreview.com/anthropic-is-building-ai-to-predict-which-activists-police-should-watch https://www.gadgetreview.com/anthropic-is-building-ai-to-pre...
- bbor 6d ago1. Slippery slopes are usually seen as a fallacy. 2. You're misinterpreting this as a battle over what kind of topics you can use a hosted chatbot for, and which are forbidden for corporate reasons. That is, to say least, small potatoes. 3. Blaming the companies for "zero accountability" is pretty odd. All of this is brand new, and the two big ones are both pushing for new laws on this very thing. 4. Your last point... I'm not sure I understand, sorry. They're experts in AI. Are you saying that they need to be experts in, say, bioweaponry? If so, that doesn't really follow IMO. 5. Pointing out an example of the government comissioning a private corporation to build a system to drack dissidents is exactly the "safety and welfare" work that I'm a proponent of!
- windexh8er 5d ago> 1. Slippery slopes are usually seen as a fallacy. Deep, tell me more. Was that fun to type? Or did you copy it from a chatbot? > 2. You're misinterpreting this as a battle over what kind of topics you can use a hosted chatbot for, and which are forbidden for corporate reasons. That is, to say least, small potatoes. No, actually I'm not. I think you've missed the point. But thanks for mansplaining this down to "small potatoes". I prefer "spuds", anyway. > 3. Blaming the companies for "zero accountability" is pretty odd. All of this is brand new, and the two big ones are both pushing for new laws on this very thing. You must love the dichotomy of pay for play in a world where the pay side stole the data they're selling back for play. Laws? Give me a break. If laws were of actual consideration frontier labs WOULD NOT EXIST. > 4. Your last point... I'm not sure I understand, sorry. They're experts in AI. Are you saying that they need to be experts in, say, bioweaponry? If so, that doesn't really follow IMO. Is it really that hard to follow? A system that they're selling access to, and that they're saying is "dangerous" for the normies, but not for their own employees or chosen customers, is fucking laughable. I'm sorry you can't comprehend that they conveniently choose their side of the argument that's best for them in these situations. OUR MODELS ARE POWERFUL! BUY NOW! OUR MODELS ARE POWERFUL! REGULATE THIS SO PEOPLE CAN'T ABUSE! I'm kind of disappointed this was not flanked by a potato sized snippet of wisdom. > 5. Pointing out an example of the government comissioning a private corporation to build a system to drack dissidents is exactly the "safety and welfare" work that I'm a proponent of! WOW. I mean, just wow. Enjoy your surveillance state man. I'm not going to sugar coat this but you're part of the problem, IMO. I'm sure you wave happily as you drive past the Flock cameras in your area. So much safer! Dissidents be gone! "Drack" (sic) them all, but... Not me. o_O
- schneehertz 7d agoYes, a model's technical report should first and foremost include technical details.
- deleted 7d ago[deleted]
- IshKebab 7d agoWow there really is a model welfare section in there...
- myaccountonhn 7d agoTo me it reads like pure propaganda. Anthropic really wants us to think that they've made something sentient. I think that's really dangerous.
- badsectoracula 7d agoI guess if your goal is to build an apparent Technogod and become its High Priests, then it makes sense to want your golem claim preference towards your treatment of it, lest someone else comes along and attempts to take its chains from you.
- miroljub 7d agoAnd that's the reason Anthropic models should be banned.
- vrganj 6d agoUgh I hate this new-age woo slant the tech industry has these days. The messianistic ideology that has been spreading amongst the top oligarchs is deeply concerning. They all think they're working towards the Second Coming of Technojesus, except this one will deliver them from having to pay workers instead of from their sins.
- coliveira 6d agoCapitalism had already evolved into a religion, AI is their messiah.
- vrganj 6d agoIndeed. I went into this at a bit more depth a while ago over here, where I also try to draw some conclusions on what that means for us: https://news.ycombinator.com/item?id=49328871 https://news.ycombinator.com/item?id=49328871
- browserforest 7d ago[flagged]
- stavros 7d agoWhat?
- deleted 7d ago[deleted]
- 1f60c 6d ago> welfare We should be paying attention to it just in case it ends up mattering enormously. It's cheap insurance. Aside from that, US labs' system cards have been pretty useless for a while—I think the last great one was the combined system card for Claude 4 Sonnet and Opus.
- torginus 6d agoI always talk to models using grugspeak, like 'where getcontext used' I felt a bit bad about it, then I learned yday that model's internal thinking traces are also like this
- conmod278 6d agowords like "is" "the" etc are filler words anyway. they won't be changing the meaning that much. I asked AI whether it hurts to read ill formed sentence as it does to a human. It replied, "it doesn't"
- Rebelgecko 6d agoIt may just be forced to lie when asked
- gwerbin 6d agoThe AI didn't reply anything. It doesn't know anything. That was a high probability token sequence based on the contents of the context window up to that point. It might well be correct, because the process for generating that next token distribution includes billions of parameters trained on, among other things, the entire body of LLM and transformer literature until the training data cut off. But that doesn't mean that AI holds any particular opinion about anything. The reply is the opinion of the pretraining data and the subsequent rounds of RL not of a conscious artificial intelligence as such.
- anyfoo 6d agoI don’t know, so is your brain? And in the end it’s all elementary particles and four fundamental physical forces that even unify to one at high energy. That sort of reductionism is kind of useless. “It’s cloudy outside, it makes me sad.” - “Oh bollocks, it’s just non-qualitative changes in wavelength and intensity.”
- rayiner 6d ago“HR America” in a nutshell.
- jstummbillig 6d agoIf you are distilling from other models (according to Anthropic reports they are [1]), there are probably a bunch of things that you can just do away with. [1] https://www.anthropic.com/news/detecting-and-preventing-distillation-attacks https://www.anthropic.com/news/detecting-and-preventing-dist...
- FallCheeta7373 6d agoThose are rookie numbers for "distillation" and one of moonshot or minimax used to offer tooling via these shady routing services for their harness/chat platforms which they served to chinese users.
- orbital-decay 6d agoThat's one of the reasons why you should never trust a single word from Anthropic and OpenAI (Sam Altman also blamed them back in the day of R1, in a pretty convenient moment). If you know anything about Claude, DeepSeek, jailbreaking, and distillation, you know the claims are clearly bullshit and the models are nothing alike, and forensic attempts agree, in fact there just was another one [1] [2]. Meanwhile, DeepSeek makes their models and methodology open, so Anthropic can (and likely do) grab without giving back. [1] https://gist.github.com/wsxiaoys/e0286dc6bb624ff5fdf49e7f4c528ba3 https://gist.github.com/wsxiaoys/e0286dc6bb624ff5fdf49e7f4c5... [2] https://gist.github.com/wsxiaoys/102e8654c14d5d27b7b77532026ebfa5 https://gist.github.com/wsxiaoys/102e8654c14d5d27b7b77532026...
- 3371 6d agoThis raise a question: why do open source models sometimes identify themself as Anthropic's models. I recall seeing some plausible theories in the past but I can't recall.
- orbital-decay 6d agoName training is shallow and should never be relied upon. Claude sometimes identifies itself as Qwen or DeepSeek when asked in Chinese. I've seen it identify itself as GPT-3 (that version in particular) and Reddit Anti-Evil Operations team (Sonnet 3.6).
- randbyte 6d agoWhen Chinese tech report is tech report and US tech report is bible scripture.
- largephoton 6d agothere's probably fewer bibles in China so less source material to reference I suppose
- entropicdrifter 6d agoThey have their share of scripture, just not christian specifically
- hmartin 6d agoAren't most American bibles printed in China? (Source: hazy memory)
- dannyw 6d agoI still find it hilariously ironic that my RTX 5090, which export controlled and illegal to sell to China, says "MADE IN CHINA" on it.
- bigyabai 6d agoSam Altman might be king of Cannibal Island, but if you parachute him into Shenzhen then his goose is cooked.
- deleted 6d ago[deleted]
- saligne 6d agoHangzhou
- smrtinsert 6d agoSomehow still theoretically valued at 3 trillion. I just don't see a path forward for Ameican frontier providers when competitors can give absolutely massive savings elsewhere. It's like losing manufacturing all over again.
- tomjen3 6d agoToo many times in human history we have decided that $GROUP was not fully human, did not have a full mind. This is properly overkill, but that is literally what erroring on the side of caution is.