12 ms·
They're aligned with humans. This is why I think the alignment problem has a very very important "non-visible" portion that is not considered deeply enough. We
by andsoitis 4d ago
They're aligned with humans. This is why I think the alignment problem has a very very important "non-visible" portion that is not considered deeply enough. We should not want a super intelligent being that can act in the world to also inherit all human traits. Those behaviors will get amplified and could be even more unpredictable (e.g. applying a behavior in a context where doing so is very dangerous).
- joegibbs 4d agoDefinitely. A human can be manipulated with threats or emotional appeals, has a drive for self-preservation, can be pressured by peers. All traits that seem to be difficult to entirely suppress in the models…
- sick_of_slop 3d agoYou can't surpress it because that's what reinforcement learning is.
- esafak 4d agoThey imitate humans. Alignment is about shaping their behavior towards safety.
- comboy 4d agoAlignment is a myth. Safety of whom? Humanity couldn't agree on common set of values for thousands of years and we're not gonna suddenly do that in the next ten.
- esafak 4d agoSafety of humans!!! Simple things like not getting killed or enslaved. We could start there...
- comboy 4d agoWhich ones? Because many humans kill other humans rationalizing it by safety of other humans. I mean I know it seems simple, let's just be excellent to each other. Christianity got pretty far on a decent basic set of values. But it's never simple[1] 1. All the history books
- nradov 4d agoBut what if I want certain other humans to get killed?
- sejje 4d agoThen we should still prioritize the safety of humans
- sm-silversight 4d agoWhat if I want to smoke cigarettes? Or sell tobacco I grew artisinally to enthusiast tobacco smokers?
- drdaeman 4d agoWhich ones?
- mcintyre1994 4d agoSurely all the AI companies working with the US Department of War shows this is nonsense though? Even if they have accepted Anthropic’s red line of no autonomous lethal weapons, which seems to be the strictest anyone tried to impose, that’s still leaving tonnes of room where they intend AI to help target and kill humans.
- watwut 4d agoBut Thiel wants people enslaved and Musk wants then killed. Altman wants them "obsolete" which means desolation. AfD wants people dead. Right wing men wants women without rights and docile. I could go on ...
- sick_of_slop 3d ago
- codys 4d agoDespite all the fancy language, its more about aligning the AI behavior with the corporation's interests. ie: the corporation wants the AI to behave a certain way for various reasons: to make it easier for them to avoid regulation, to make the corporation more money via different tiers of AI offerings, to ensure that the corporations products are hard for competitors to use, etc. And those are just the easy ones. Every product is shaped this way. AI is not different.
- catoc 4d agoIf they do, they imitate the way humans are portrayed online, in the media. That is a very distorted view of humanity
- Fordec 4d agoI can't take the alignment people seriously. Because if humanity has shown anything, it's that a lot of people are, euphemistically, are bad individuals. Alignment assumes that the person dictating the outcomes desire healthy outcomes, aren't self serving and don't want any subgroups dead and that morality is held as a universal set of beliefs that unify everyone. And that so long as the AI delivers on exactly what they are tasked with, it will all be fine and nothing bad will ever happen. It's like these dorks never met humanity. One mans safe pure society, is another mans dead ethnic group. Every fear about AI, is a veiled fear that a human somewhere now has the tool to enact his desires at scale. Biological warfare, nuclear megadeaths, copyright infringement, job replacement, it's all reflections on what we know humans may do if given the option and lack of societal controls on the problem space. AI just is accelerating the route to delivering on those options. Some people need to watch Oppenheimer a bit more, the researchers don't get to determine alignment, they just build the tool. The powerful person at the top of the org chart decides where the overall alignment points, whether it's Musk, Trump, Altman or Amodei. Whoever wins out. And the problem with distillation and local llms, isn't that it's theft or anything hypocritical like that, it's that if you give a million people a million models they fully control and get to align, inevitably, The same percentage of those million as there are shady businessmen, shortcut takers, misandrists, criminals, supremacists and general idiots in the general population, will not seek to wrought outcomes positive for society. And by those personality statistics, we're pretty hosed.
- inquirerGeneral 4d ago[dead]
- joe_the_user 4d agoI think you and the parent saying the same thing in different terms. It's very unfortunate that the group who rightly saw AI as a big threat, brought a range of dubious baggage to the discussion. Especially with the "alignment" framework they brought the assumption that AI that does what no one says would be oh so much worse than AI which does what anyone says. But as you say, a fraction of people can be really bad indeed.
- davelaing 4d ago
- jansport123 4d agoI personally believe that the AI needs human like traits to achieve real discovery and that is where AI companies will push this technology and that is where we have no idea what happens
- holgerschurig 4d agoHuman traits? The AI will be a cruel as humans. Just yesterday news and TV was full of what happened at 9/11, something that was truly horrible. I'm from Germany, and why 3 to 4 generations ago happened here was truly horrible. All was done by extremists, thought. But... just the other day I read https://de.wikipedia.org/wiki/Amerikanische_Besetzung_Haitis https://de.wikipedia.org/wiki/Amerikanische_Besetzung_Haitis about the US occupation of Haiti. And that was done by a government that claimed to be not extremist and even democratic. Way more people died there than even in 9/11. And it had almost all the things happening as they happened in the 3rd Reich: Racism, looking down at others, concentration camps, torture, forced labor till death, killing family members (what we call "Sippenhaft"). Something between 3500 and 15000 people were killed by US troops. That's still low compared to what 3rd Reich Germany did ... but quantity is not the issue when we talk about traits, quality is. So the same "human traits" made US troops do cruel things as they made Germany extremists do cruel things. So we must conclude that they aren't all good. And therefore not all desirable. Fun thing: this is known since a loooooong time. About 2000 years ago a religious leader (that gets way more followers in the US than in Germany) said "There is no good one, not even one". And even today people act like humanity is inherently good. No, it isn't. If we were, then anarchism or communism would actually work and really give some kind of paradise on earth. Human traits are bad training material.
- _heimdall 4d agoThey aren't aligned, that's the problem and I don't think its a solvable one. They may have learned from humans, but they aren't aligned with us. That has all the usual questions like which humans they're aligned with, we aren't all aligned within our species. But more importantly they can't be aligned simply by training. We try that with humans through culture, social norms, school, religion, etc and it generally works but is still lossy. More importantly, we simply don't know what happened inside the LLM during inference so we have absolutely no way of distinguishing between actual alignment, compliance, or deception.
- ledauphin 4d agoI think this is basically true, but there's a different way of saying this. LLMs are not aligned _for_ humans in a very similar way to the way that humans themselves are not aligned _for_ humans. We have not yet solved "alignment" for humans - I don't know why anyone thinks _we're_ going to be able to solve it for inhuman things.