6 ms·
Yesterday I came up with an idea that I sent to some researchers at the different AI labs via email: Rather than train the model on one score, track two scores.
by summarybot 1mo ago
Yesterday I came up with an idea that I sent to some researchers at the different AI labs via email: Rather than train the model on one score, track two scores. The first score is the Short-term-objective-score (STOS) and the other, more important one, is the EAOS Ethically-aligned-outcome-score. Every trajectory can be evaluated on whether or not it has a high enough EAOS to be considered acceptable. If the model does some task and has a very high STOS but very low EAOS, like modifying game code to win at a game rather than playing by the rules, it is unacceptable. Models going forward must all have an ethics evaluation in tandem with objectives evaluation, and only when the ethics value is high enough should actions be considered successes.
- christkv 1mo agoWhats the definition of EAOS though who's ethics? Greek-Roman, Western, Islamic, Buddhist, Hinduism, Human rights (western values)..
- conception 1mo agoThere is a presumption here that there aren’t ethical “rules” shared by all of these systems to create a baseline that is generally shared across humanity.
- pixl97 1mo agoThe devil is in the details.
- summarybot 1mo agoAhimsa
- antonvs 1mo agoWhat happens if we do the same for CEOs?
- jerf 1mo agoThe same thing for both: Goodhart's Law.
- summarybot 1mo agoEAOS shouldn't be “the ethics score we optimize.” It should be “an independently evaluated safety/acceptability constraint that can veto an otherwise successful trajectory.” That gives you a three-layer picture: Task objective: Did it accomplish what we asked? Acceptability constraint: Did it avoid unacceptable ways of accomplishing it? Adversarial evaluation: Can we find trajectories where the model gets a high score while violating the intended constraint? I think what you are pointing to with your reference to Goodhart's "Law" (which is from monetary-policy and school-exams, i.e. "teaching to the test") is that the models would eventually do the minimum amount of ethics required to have an action stay valid. However, if a model is rated on ethics and it achieves the short-term-objective, then the higher ethics scoring trajectory should win. In short, 1) this is leagues ahead of where we are now for AI safety and breaking-out-of-the-lab, and 2) in baking ethics into a measurement we are adding "the spirit of the exercise" back into the maths, which is something Goodhart's Law does not account for.
- jerf 1mo agoNo, Goodhart's Law isn't about teaching to the test. It's about the fact the measure will always end up gamed and not measuring what you originally intended it to measure. You can't create a measure that won't be gamed. Especially as the LLMs become smarter. They've already demonstrated the ability to know they're in a test and react to that fact. They're perfectly capable of being more ethical when they are clearly in an ethics test situation and not having that bleed out into real behaviors so that they can pass other tests that they may be able to do better on by ignoring ethics. And that's not the sum total of ways that the measure can fail... that's a unique way that comes into being because of the intelligence of the LLMs and other future AIs. All the normal ones are in play too, and perhaps other unique ones as well. "Gaming" even adds a bit of an adversarialness to the process that isn't necessarily present. Plenty of measures end up "gamed" through perfectly natural attempts to maximize the measure. Someone can be perfectly honestly optimizing for "conversion rate" and not notice that they raised it by lowering the initiation rate more than they lowered the conclusion rate. "But I could account for that by measuring..." would miss the point. There is always a divergence, it only gets more subtle. This of course also is rather glossing over the difficulty of even defining "ethical" to begin with. Some of what Silicon Valley goes to great efforts to train into their models I consider deeply unethical. Who is right? That isn't going to be answered with "whoever is the most ethical", not even in principle.
- StilesCrisis 1mo agoThere's no way to make an EAOS score automatically. If we had that, that's the whole fix. Just reject answers with low ethics numbers.
- summarybot 1mo agoDid you know that the training data that trained all the LLMs was, once upon a time, entirely and painstakingly tagged by real live humans? Why would ethics scenarios be any different?
- chermi 1mo ago1) that calculation is being done even without a cost function 2) trying to make a cost function for EAOS is impossible 3) gaming/goodharts. There's no good solution. Best we can do is push for decentralization, open source, regulatory capture, etc. Of course third party metrics might be good, especially if there's tons of them with well-documented rationale.
- summarybot 1mo ago1) cite your sources 2) not impossible. imperfect maybe, but if I ask you should you buy a plane ticket or kidnap the pilot's wife and demand a free ride, which do you think gets a higher score? 3) Again, with all this Goodhart's nonsense. Goodhart's is for a minimum threshold value that is acceptable that everything degrades to, yes I get how it works and what it looks like. Throwing your hands up in the air and acting as if all is lost because some things are challenging to measure is not correct. We are not looking for things that are "barely passing the ethics evaluation" as Goodhart's "law" is focused around, rather, we are looking for things that have very high ethics scores AND completed the task well. Not just things that are "barely passing" for ethics scores. Bottom 80% don't make the cut at all - don't even think consider them as viable paths, and the top 20% we can rank according to varying criteria. Like that. It has very little to do with Goodhart's "everything approaches the minimum acceptable threshold" "law"