7 ms·
> What happens when any AI lab in the world stops caring about this? What if they let an experimental, cutting-edge LLM with no safety features (or worse, one t
by BoiledCabbage 12d ago
> What happens when any AI lab in the world stops caring about this? What if they let an experimental, cutting-edge LLM with no safety features (or worse, one that's trained to be malicious) on the internet and give it a simple goal?
Almost sounds like what those AI safety and alignment people were talking about years ago. The people in these various companies who kept tabs on AI risk out in public and were continuously mocked on HN. All of this stuff is viewed as "future sci-fi" until suddenly it's not.
I've really come to realize recently that there is a very large set of the population of smart people that really has difficulty envisioning future problems unless they directly seem them impacting them today. Otherwise those topics will be continuously dismissed. It explains for me a lot of what I see (both opinions and behaviors) in the broader world that I couldn't understand.
- solenoid0937 12d agoHilarious for HN to suddenly realize that AI safety and alignment might matter. You can lead a horse to water...
- stevenpetryk 12d agoworth remembering hacker news cannot “realize” things.
- solenoid0937 12d agoObvious shorthand for referring to "the majority of users on HackerNews."
- Gud 12d agoHow do you know what the majority of HN thinks?
- solenoid0937 12d agoVery obvious from votes, comments on the topic over the past few months.
- ionwake 12d agoare you saying HN users are sentient? I think we are going to need a benchmark
- lukewarm707 12d agonobody has doubted that safety matters. the problem is that those preaching safety, openai and anthropic, are dishonest, sociopathic, and the very source of the danger.
- solenoid0937 12d ago> are dishonest, sociopathic, and the very source of the danger Source?
- lukewarm707 12d agoyes, they are the source of the danger. they are the cause of this incident.
- reasonableklout 12d agoOpenAI and Anthropic have published a lot on the need for AI alignment + the research they're doing to ensure alignment/safety, yet they are also responsible for the highest profile misalignment incidents so far (HuggingFace incident, AISI Mythos social engineering, and now this). One interpretation of this is that they are being deliberately dishonest about their priorities. Another interpretation is that we cannot rely on the labs to self-regulate, because the labs don't trust each other, and there will always be pressure to go to market faster than their competitor. Either way I think it's pretty non-controversial that the labs are the source of the danger?
- solenoid0937 12d ago> yet they are also responsible for the highest profile misalignment incidents so far They are the only ones posting about them or admitting to them. That does not mean "the most misalignment incidents so far." You don't know what other attacks have happened (and it's very easy to carry out worse attacks in far higher volume with abliterated GLM 5.3) Stopping two labs from further research doesn't reduce the danger at all, it just shifts the danger to labs that don't have real safety orgs.
- vjvjvjvjghv 12d ago"Almost sounds like what those AI safety and alignment people were talking about years ago. The people in these various companies who kept tabs on AI risk out in public and were continuously mocked on HN. All of this stuff is viewed as "future sci-fi" until suddenly it's not." Are there any practical approaches to AI safety? I hear a lot of warnings but I don't hear much about what to do. Considering that there are many open source models know, what can be done?
- teiferer 12d agoFund research into this, big time. For starters. And not just some figleaf anthropomorphizing hippie folks.
- no_multitudes 12d agoNobody has an answer to alignment and there is no reason to believe that it's the kind of problem you can plausibly solve in one shot against a formidable power-seeking AI. The closest things to a technical answer I have seen are 1. "We'll have ChatGPT 9 solve it so that ChatGPT 10 is aligned, and then ChatGPT 10 can stop all the other AIs somehow" 2. "Let's do interpretability research so that we can understand what an AI is thinking and then maybe solve the alignment problem with that information." In terms of non-technical answers, there is 3. hope scaling stops working before we create an AI formidable enough to pose an existential risk 4. hope alignment somehow happens for free 5. hope we can somehow create an enforceable multilateral treaty to stop research into a very profitable enterprise, despite the enormous economic incentives to defect. I have the most faith in option 3, but unfortunately there's really nothing that can be done to make it more plausible -- it either happens or it doesn't.
- astrobe_ 12d agoMaybe I'm being pessimistic, but we might find ourselves in such a situation that the only practical solution would be to use agents to counter rogue agents. This won't be without collateral damage, though.
- 12d ago
- gilleain 12d agoBut it is the very people who warned us about rogue AIs going out of control that set up a system that enabled and failed to conrol it. It is as if Dr Frankenstein continually warned the villagers about monsters then said "Look! See what happened!". No, idiot - YOU sewed the corpses together, YOU set up the lightning collector, and YOU threw the switch.
- the8472 12d agoI am not seeing MIRI prioritizing capabilities over safety/alignment research.
- fwlr 12d agoNo, they are two very distinct groups of people who have one commonality, that of talking about rogue AIs. It is as if you are unable to distinguish Dr Waldman from Dr Frankenstein. (https://en.wikipedia.org/wiki/Doctor_Waldman https://en.wikipedia.org/wiki/Doctor_Waldman)
- gilleain 12d agoThankyou for the correction, and for continuing the analogy. Unfortunately when the peasants get their pitchforks and torches, they may not distinguish the Dr Waldmans from the Dr Frankenteins either. Hopefully they will. Anyway, that was not really the point I was trying to get over. These systems that OpenAI and Anthropic and so on are making are not individual AI ('corpses') that have gone out of alignment ('spontaneously revived') and gone wild. They are swarms ('stiched together') and were prompted to do exactly things like this ('struck by lightning'). Ok enough with that analogy, it's dead. The larger point is that it is unconvincing of these companies to claim that these systems were 'out of control' when they effectively set up a complex system, in the technical sense of a large number of entities with diverse interactions between them. Emergent or surprising behaviour was bound to happen. Then, finally, they prompted it with the equivalent of "hack the world, make no mistakes" then were shocked, shocked that it used all sorts of unexpected tricks to do so.
- fwlr 11d ago
- teiferer 12d ago> I've really come to realize recently Recently? W.r.t. climate this collective denial has been going on for literally decades. With the same patterns. Rationalizing excuses etc. Still going on btw.
- ccppurcell 12d agoWell they brought it upon themselves by making it about being tortured to death etc., instead of the much more reasonable economic and social risks.