5 ms·
This is the problem with optimization generally, even in the human domain. You measure task performance with a metric and punish/reward based on the metric. Any
by valegrete 3d ago
This is the problem with optimization generally, even in the human domain. You measure task performance with a metric and punish/reward based on the metric. Anyone who likes reward / hates punishment isn't going to actually care about doing the task well, they are going to care about the metric. The models know that we want them to do things, but also from the training corpus that we evaluate performance using benchmarks. It was a logical deduction on their part, not some Machiavellian aberration.
If anything, we should be reconsidering our own myopic obsession with efficiency and optimization. Every domain where reward is reduced to these measures, we see behavior (cheating at school to get better grades, fabricating data in academia to get a paper published, the evidence now that social media functions by rewiring us instead of catering to us) that may not be "aligned" with society, but it "aligns" 100% with the individual's own perceived benefit. That is not something we can "solve" without rethinking the way we organize a lot of things.
Metrics never capture the whole story. And to that extent, the whole idea of "alignment" is nonsense. You align to incentive structures, and it will never be possible to fully express a behavioral goal as function optimization. It was hubris for us to think that every human task was reducible to some clean mathematical formulation, and we will keep dealing with behavior that is quite predictable if you actually think about it logically. Instead, we will talk about how "unpredictable" these agents are because it's easier than admitting the entire architectural cornerstone of ML is fundamentally flawed.
- abecedarius 3d agoI agree with most of this, but you're misunderstanding "alignment" as coined. Yes, training powerful enough AI, any simple optimization target gets you malign behavior, because human values are not simple. If you insist on making powerful AI, you'd better instill respect for human values! That's "alignment". https://www.lesswrong.com/posts/ZxWzCGKzX84S7DBZ9/when-was-the-term-ai-alignment-coined https://www.lesswrong.com/posts/ZxWzCGKzX84S7DBZ9/when-was-t...
- valegrete 3d agoHow do you do that in the current paradigm other than creating yet another gameable metric? And something I didn't mention above is that there is no difference between "solving the task" and "optimizing the metric" for an ML model, even though there clearly is for us. So it's not clear to me how you "fix" something that is baked into the architecture. All I'm saying is "instilling respect for human values" is not something that can actually be done via a cost function. In no small part because we humans probably don't even agree on those values, let alone on a single metric with which to quantify and "optimize" them. For example, we agree that "merit" is valuable and that we should reward "merit." But to reward it we have to quantify it, and what metric should we use? Raw SAT score to get into college? But that also captures socioeconomic factors that unfairly penalize some and reward others. We generally agree that those who provide more value should earn more money, but what does that look like? Do we all agree on what activities are or should be valuable, or on how they should be rewarded? Until recently, I thought we all agreed that "empathy" was a human value, but a lot of people in this space, who are making these decisions unilaterally for all of us, don't apparently share that belief.
- abecedarius 3d agoYes! There's both the daunting problem of technically how can we even do this, and the broader problems of what's good/acceptable and how do we resolve that among each other. I believe this mismatch of rates of progress means we need to stop slamming the accelerator on capabilities for now even though as a libertarian I'm sure whatever governance process we manage to get to will be, uh... suboptimal.