5 ms·
I don't advocate for throwing the whole problem away or handwaving it as impossible. I'm just saying that what we call "alignment" is impossible to solve in the
by phoe-krk 1mo ago
I don't advocate for throwing the whole problem away or handwaving it as impossible. I'm just saying that what we call "alignment" is impossible to solve in the general case, because it's so poorly defined that even humans don't "align" on ethics and morality, however we define them. Just, in case of humans, we tend to close our eyes and and call it "politics".
All "alignment" solutions will need to be contextual, just like a researcher hacking their way to some content might be lauded a hero in a context where there is no other way to reach it and something valuable depends on getting it out.
- freehorse 1mo agoIt is repeatedly shown that the models are aligned towards tasks, not constraints. Which is also how the current market and companies operate, and use cases in general favour that. So I am not sure it is incidental, inevitable, "vague" or "just hard" as opposed to by design. If you define a task and constraints to that task that contradict each other, the model is gonna try to solve the task against the constraints. It is perfectly aligned to what it is supposed to be aligned, because "alignment with human ethics" and other stuff is just a theatre or afterthought at best.