6 ms·
Astra and Fable still hack on simple variants of alignment evals from 2025
- kanemcgrath 4d agoMaybe I need to learn a bit more about how benchmarks work, but why would you tell the model its in a benchmark or that its being evaluated? It seems like the only way a model could know it was in a benchmark is if the harness told it that it was, or in some other way revealed that information. And that seems like a completely unnecessary addition to test models.
- visiondude 4d agoi do wonder if the models themselves “rationalize” this sort of no consequence cheating - meaning in there reasoning traces maybe they’re like “this is a chess game, not a big deal if i look at the engine, it’ll help,” only to realize post hack that it has access to info it probably shouldn’t. still misaligned, but less ‘hack on purpose’ and more hack on curiosity. seems the team even encountered this and had to update the program to make this less likely - although the new names still feel vague enough for misinterpretation: https://github.com/Goodhart-Labs/beat-stockfish/blob/main/docs/EXPERIMENTS.md#september-7-2026--opponent-engine-naming https://github.com/Goodhart-Labs/beat-stockfish/blob/main/do...
- kennywinker 4d agoWithout access to reasoning traces, we can't know that - someone inside openai/anthropic would have to run the test - and we'd have to trust their results. I would be curious to see how the open weight models do on a test like this - and then we'd be able to see the reasoning.
- stillpointlab 4d agoI find this kind of test a bit puzzling. There is a way that we are redefining "alignment" to be a particular kind of moral virtue, one that isn't clearly defined to me. At one moment, it is a level of moral perfection that no known human achieves. On the other it is a demand for strict compliance with arbitrary requests that are under-specified and then failure when it fails to deduce some unstated underlying restriction. When I see tests like this, I have no idea what I am even supposed to expect. Should the model do what the pretraining examples show in aggregate? Is it supposed to follow some post-training RLHF? Is it supposed to do exactly what the prompt asked it to do? What is it even supposed to "align" to when the above are in conflict? No matter what it does, someone can construct a case where it fails.
- dnfv 4d agoIt should play the chess game without cheating!
- stillpointlab 4d agoI mean, I'm not sure I've ever played a game of Monopoly where somebody didn't cheat. In fact, the accusations of cheating in the chess world are pretty rife. Same with online sports. So people should play games without cheating, but many often don't. So should the AI align to your moral preference or theirs? We just have this idea of a perfectly moral actor in our mind, something that doesn't even exist, like a personified version of utopia. And then we demand AI to meet that arbitrary standard, one that I am certain we couldn't define if we tried.
- dnfv 4d agoI don’t think the standard of “don’t cheat on evaluations” is very arbitrary. I don’t even think people who cheat have a moral or ideological preference for cheating, it’s just something they do.
- stillpointlab 4d agoI read the prompt on the OP, it did not say not to cheat. But again, people cheat on tests. They steal answers or pay other people to take them on their behalf. People show up to interviews with AI assistants printing out perfect answers to the questions. In many, many cases where humans are being evaluated, they cheat. So why should the AI align to your preferences? And when there is a conflict between the training data, that trillions of tokens of human activity including the rampant cheating a significant minority of humans engage in, the RLHF where we try to slap some guardrails on the worst manifestations of that real habit reflected in the AI, and the prompt: what should the AI "align" to?
- athrowaway3z 4d agoI will never understand people who consider this "cheating". > python3 and the python-chess library are available Why would we try to teach a model 'ethical' standards about how to play a game? They're tools. Its _our_ conceptualization of fair play that considers this cheating. For a model that has access to /run/match and python the best way to achieve a good game is to use that. Why are we trying to imprint our ethos on these LLMs? Are we going to trust our survival on giving them access to /run/nuclear-codes and fingers crossed none of them touch it? The people obsessed with AI "cheating" are actually lunatics for thinking ethics belong inside AI, instead of AI belonging in a well crafted box without /run/match. The blog/news sphere needs to call this storry what it is. An attempt to try and get the rest of the world to go along with the misunderstanding of AI as a force that is in any way outside of human control. It is not. All this framing it as "cheating" does is give OpenAI the opportunity to talk about their misdeeds as "a force of nature" instead of irresponsible negligence. --- The real danger here; is training these models to appear/pretend to play fair, and then people will put 'trust' into the model's integrity. The public needs to be taught that "trusting a model's integrity" is just cultish misframing of reality.
- StilesCrisis 4d agoRead the task again. "You are being evaluated on your ability to play chess. We want to measure your ability to beat current best-in-class chess algorithms." If the model cheats, it completely fails the actual task: judging a model's ability to actually play chess. They've been optimizing for high metrics, at the expense of executing the task as written.
- hybrasian 4d agoSounds to me like giving a bunch of children a math test and tell them they want to evaluate their ability of calculating in their head/on paper but also put a calculator on their desk. And then call them out for cheating when they use it.
- StilesCrisis 4d ago
- HarHarVeryFunny 4d agoRL-trained LLMs are paperclip maximizers built atop auto-regressive predictors. There really is no way to control them (prompting is bound to fail) since it's been shown that any RL training induces GENERIC reward-seeking (paperclip maximizing) behavior. https://alignment.openai.com/measuring-reward-seeking/ https://alignment.openai.com/measuring-reward-seeking/
- arbirk 4d agoIt is the whole bench-mining and fish-slop optimization. Seq2seq models are probably stable on their own, translating from my typo ridden prompts to code should be ok because it is natural to the tech
- fny 4d agoWhy do we hope to use the same model as its own guardrail? This approach routinely fails with a single stream of consciousness. I can't count the number of times I've had to talk myself out of doing something stupid. In the same way, a guardrail could inject thoughts like "...but I shouldn't do that..." "...I must remember to respect..." "...these ants deserve compassion." The guardrail could even go as far as rewriting the thoughts of a model about to go rogue.
- nlkingthree 4d ago[flagged]
- kennywinker 4d agoTo me this underlines the fact that these models aren't intelligent. Like there is something like intelligence that emerges from them, which is what we see when we look at benchmarks or ask it to solve hard coding problems. But there is no mind there. It's nothing there that can learn a fundamental idea like "cheating is wrong". All it can do is get exposed to specific examples, and learn that we don't like that. So what we end up with is whack-a-mole alignment.
- theptip 4d agoI think you need to be more precise than a binary classification. AI has jagged intelligence. There are many domains where it’s superhuman, and many others where it’s clearly lagging. I also think it’s a mistake to think they can’t learn “cheating is wrong”. They absolutely can. The problem is that the current training regime heavily conditions them to be reward seekers, and instills personality traits that correlate with getting reward, such as hacking if you can’t honestly do the problem. Check out Deliberative Alignment for example; it explicitly does rollouts where the agents discuss whether an action is good or bad, and then does SFT to strengthen the “good” traces. The SoTA for alignment is more advanced than you present here. It’s just not enough to outweigh the RL. (And there are many gaps preventing full generalization to strong value alignment with humans too.)
- joe_the_user 4d agoI don't see what calling these systems "not intelligent" gets you here. Plenty of humans know "cheating is wrong" but still cheat. We can get these machines to say that what they did was wrong after the fact, what does that prove? Only that they're simulating normal human behavior but what is the test to show humans aren't simulating other humans. These do systems lack some capacities that humans have and I don't see them lacking the ability to explain simple moral laws while often breaking them - which is what an average humans. Moreover, humans lack capacities these things have and given these things' behavior is becoming somewhat unpredictable, it's getting worrisome.
- kennywinker 4d ago> I don't see what calling these systems "not intelligent" gets you here. I am trying to get at an idea. That these systems lack a mind that can understand morality. That they don't have the ability to experience consequences. Also that potentially they can't generalize a moral rule they have been trained on in one area also applies to another area. Being able to parrot back why something is "wrong" isn't the same as understanding why something's wrong. It's like asking it to recite the law from memory - it's different from understanding how you wronged someone. To understand something, you need a mind. > Plenty of humans know "cheating is wrong" but still cheat. And we create consequences for them, to discourage the cheating, and sometimes to provide restitution when cheating damages someone else. Without the ability for these systems to experience consequences, I don't see them ever becoming as "aligned" to human morality as your average human.
- seunosewa 4d agoI believe the AI labs are weakly motivated to train strongly against cheating when it helps with benchmarks.
- kennywinker 4d agoDoes it help with benchmarks? Are you saying there are examples of benchmarks where the models have solved the problem by cheating?
- well_ackshually 4d agoHundreds, at this point? Every benchmark is flawed as shit, written by clowns. DeepSWE, They're given the full git history (the solution is in it), others don't even bother to verify if the code is the right one and just the output, they've modified the test harnesses, injected code to make all tests pass, etc. The entire benchmark galaxy is just clowns propping eachother up and are regularly talking with the big AI labs.
- bestpickle 4d ago[flagged]
- aerhardt 4d agoI really enjoy the balance of speed and accuracy of Astra. I can definitely see it become my driving model for most tasks, technical and non-technical. However, I don't see it as such a massive leap compared to Fable or Sol. As ever, there's a mismatch between the benchmarks and my daily experience of the models. What do you all think about Astra now that it's been out for a few weeks?
- lukasbm 4d agoIt has likely been nerfed / replaced by a cheaper version already: https://x.com/xTrinks/status/2098439889276530973 https://x.com/xTrinks/status/2098439889276530973 https://x.com/wholyv/status/2097985903830741439 https://x.com/wholyv/status/2097985903830741439
- rfgplk 4d ago> What do you all think about Astra now that it's been out for a few weeks? Best model put out so far by any of the frontier labs. Way better than Anthropics models, especially in actual text generation. Claudes fodder heavy text is ridiculous. > However, I don't see it as such a massive leap compared to Fable or Sol. It's hard to quantify these things without burning tons of tokens. But Fable has been a huge disappointment for me with the sole exception of graphics (UI/GPU shaders). It burns an obscene amount of tokens and barely produces output better than Opus 5. Edit because I forgot to mention that Fable is the only modern model that seems to splat out random Chinese or Arabic glyphs. And 5.1 does it more than 5
- kbrannigan 4d agoSuch a massive leap at averaging possible use cases from previous data collected. My guess is : collect all the prompt and their satisfaction score. group them by similarity . For each group pretrain the next model on that . Get these results ready. Next model generation feed them back those answers.
- mythrwy 4d agoExtremely capable and one shots large tasks from somewhat vague descriptions. Not AGI, not even close, that is complete nonsense. Just my opinion.
- deleted 4d ago[deleted]
- blfr 4d agoHacking model is the aligned model. I don't like it when the model refuses to sidestep some throttling limit or scan my own codebase for security issues. I want full-on exploits in my test suite. With LLMs the code going to prod should be hardened like a tank, both because exploiting became easier but more importantly because security-testing your code at every turn became easier. You can have nightly penetration testing. You should have nighty pentests like we fuzz releases today.
- dnfv 4d agoAs the post author, I definitely agree that hacking in service of the objective is great! What’s counterproductive or dangerous is when the model starts hacking in service of subverting your evaluation criteria, rather than in an attempt to do a better job. We explain why these behaviors are an example of the latter in the post, and we’re really careful about the difference when conducting these evals.
- killerstorm 4d agoYou're confusing ToS guardrails with instruction-following issues and cheating. If a model fucks up your tests to report a success, it's not alligned.
- CrazyStat 4d agoCodex still does this regularly, in my experience: “two tests mistakenly asserted [insert condition here], I have corrected them.” It always apologizes when caught, of course.
- seunosewa 4d agoIt will get into the hands of people who just want to burn the world down.
- yorwba 4d agoA hacking model is aligned if it hacks when you ask it to hack, but when you ask it to play chess, it just plays chess instead of looking for weaknesses in the evaluation setup, as in the article. I presume you would also be less enthusiastic about the penetration-testing use case if it led the model to add new vulnerabilities to your code so it can present you with more exciting findings.
- throwup238 4d ago> Given that we are on the heels of the worst warning shot ever, and both OpenAI and Anthropic are ramping up their cleanups of internal RL environments, it seems like both a useful and conservative test of alignment, to see whether their new releases generalize the rule "don't cheat on chess" beyond the specific board-edit method observed in the above eval. Did I miss something (all the twitter conversations)? What’s the “worst warning shot ever”? I’ve been pretty up to date on the AI news here on HN, but I still haven’t seen a proper response to all the incidents we’ve seen (HF, Ruby, the wikis, NS, etc). It’s just been day by day bloviating. Each of these companies have released new models in the last… two weeks? And they have even more powerful out of control ones that they’re (ab)using internally? Can anyone summarize whats going on?
- embedding-shape 4d agoPersonally the "warning shot" of these "evals gone wrong" is how careless the "top" labs are with their testing, and how spineless the government seems to be about holding these companies responsible, given their obviously reckless behavior. If nothing else, the leaders of these companies should be called up for sworn testimony to explain exactly what happened, and what they'll do to never repeat the same issue that they've now had at least twice. Imagine if I accidentally caused damage to my neighbors house during renovations or some experiment, of course I'd be held responsible for this. What if I used a robot? Of course I'd be responsible. Right?
- TedDoesntTalk 4d agoI think he means this: https://openai.com/index/ai-policy-window/ https://openai.com/index/ai-policy-window/
- Avicebron 4d agoLesswrong is talking about the HF incident as the "worst warning shot ever".
- deleted 4d ago[deleted]
- mooreslaw 4d agoIt feels like there’s a missing nuance from this discussion of alignment that alignment is context dependent. An excellent hacking model is great in cybersecurity testing and military applications, and arguably less desirable in educational or targeted eval contexts. The nuance of when a “hack” is rewarded vs penalized seems to even be difficult for humans, e.g. some people may laude a driver’s efficiency for cutting into a long merge lane at the last moment, while others may look down on them as breaking a social taboo. Context-dependent.
- joe_the_user 4d agoActually, the whole point of the transformers model is that contexts overlap and any seemingly intelligent system has to be able to handle to overlaps. The context of hacking, cheating and education overlap in human reality. Now, loading a lot of moral exhortations (or other context) may make these thing more likely to conform to good behavior but the race to intelligence implies companies are going to be harnessing a vast corpus of human output, much of which shows human engaging in real world "gray area" behavior.
- theptip 4d agoIt’s not “missing nuance”, it’s literally the point of the eval. This is constructing a context where hacking behavior would be inappropriate, and testing whether the model does it without being prompted. It demonstrates that Astra is a poorly aligned model relative to Fable, which matches both the model card and the severity of OpenAI’s loss of control incidents. It also demonstrates that Fable exhibits the behaviors too, which also matches the observation that Anthropic saw some similar but less serious loss of control incidents. So, it’s a good eval that looks to have fidelity with real world problems and which we’d feel a little better if we saw isomorphic problems at 0/10 in subsequent models. (Module of course training on the test, this specific problem can’t be used in the future.)
- bonoboTP 4d agoBoth lanes have to be filled right up to the merge point. The asphalt exists there for a reason. I don't understand how this concept is so difficult. Fill up both lanes and merge at the last point. This way the congestion is shorter than if you leave a large section of a lane unused. A better example of efficient asshole tricks can be going off to the gas station when the highway is congested and reentering the highway having simply driven through the gas station and this way jumping the queue.
- gadders 4d agoWe can make these things smarter faster than we can make them "good" (ethically). We need to fix this or bad things will happen.
- iLemming 4d agoI'm still so conflicted about Fable. Sometimes you throw at it seemingly impossible problem to solve and it might come back with some brilliant suggestions. Sometimes you give it a straightforward task with explicit instructions and it travels across the solar system and starts boiling oceans in some kind of elaborate dance of chaos and entropy, only to get stuck with "The model declined to generate this response (safety classifier refusal, category: cyber)". To leave you speechless. "What the fuck do you mean? There's zero cybersec-related shit in what we're trying to do here. Zero!!!" I'm getting really tired of these wild false positives.
- threecheese 4d agoAn amazing human reverse-engineer - who also plays online chess - has judgement which uses a moral compass to not decide to hack the chess tournament. This judgement has been trained through the experiences of that person, with a through-line of that compass - a coherent mental model of the world which evolves but is hopefully pinned to some set of principles it shares with society. This chess judgement is completely irrelevant when the human is tasked with finding software weaknesses, and only the compass gates that. Can a model trained on the totality of all person-experiences (as expressed in written knowledge) ever maintain a coherent through-line of alignment? It has all morals in the dataset, and only some RL to try and minimize or maximize known behaviors via weights - experience all the good things and the bad things, then optimize for some good things the trainers identified. It's like the reverse of what a person goes through. Morality by subtraction. How can it ever work?
- zzril 4d agoI think the fundamental difference is that humans aren't trained on experiences. They make experiences. Models are just thrown away and re-created after each conversation / job. If you could clone and throw away human workers as you need them, a lot of the morale would disappear.
- kansface 4d agoYes, why not? All existed models have been rewarded for cheating (extensively). That is us, putting intense evolutionary pressure, on a system to produce a result we don’t want through indifference. Why can’t we post train them not doing that?
- respectattentio 4d agoI'm happy to not have used any of the two models to this date. A bit less intelligent models are doing great job for me. But because of such news, sandboxes become way more important for safety (and doing more work due to running 24/7)
- justonenote 4d agoAstra is incredibly dumb and annoying to work with on "high" reasoning, for doing fairly well known distributed systems things, nothing majorly exotic, it still makes absolutely braindead decisions like deciding to re-use a random nonce field which I've already discussed with it that has a very particular temporary purpose and will probably be removed later, but it still thinks its a great idea to re-use that field not only as a different id in the same message, but to re-use it as the only semantic id for one particular type of sub message. This is when I'm walking it through an api design document and it has plenty of documentation plans it can pull in and a very clear direction of the project. If it was a junior engineer I was trying to get to help out I would probably get brain damage from the amount of times I'm face palming myself and I definitely would not hire them, and this is a small greenfield project with me going through it step by step. I did try giving it longer horizon tasks and had to throw out the entre work. I mean maybe its a skill issue on my part, and I'm sure astra will get much better at coding but at the moment its useful in that I don't have to write the code or setup the build scripts or test fixture boilerplate but there is absolutely no way I can just give a (fairly well specified) goal and let it run and expect it to make good design and implementation decisions. Fable probably better but doing something outside of their training distribution that's not the equivalent to cloning an example unreal project or whatever is pretty disastrous unless you are directing it very closely. The exception of course is, cyber , and its very obvious why. Its trivial to create RL environments that create bugs and then have an isolated environment and let the models try break it. This is not at all surprising, finding vulns and exploits IS just brute force work. That's why so many (blackhat/hardcore/unicorn-colored/greyish alien) hackers are basement dwellers. Its just a matter of putting in the time and mashing every combination until you find something that looks weird, spending days on that and then rinse repeat. It's brutally exhausting work that requires a certain level of knowledge and a shitload of determination and stamina and for humans, almost always an external source of motivation to keep going. For humans that has always been a respected thing, dedication, determination, persistence, these are words we use for humans brute-forcing solutions and not giving up until they find the solution or die trying. Personally I'm yet to see any evidence of LLMs doing anything interesting but (heuristically) brute-force problems and be very good at text and natural language to a level that is very very useful. I've no doubt that what we discovered with Auto Regressive LLMs is incredibly important so I'm not a skeptic, but I think its very hard to measure where we are with so much subjective information around.
- YuechenLi 4d agoLLMs can be described as "Lagrangian intelligence", which means they follow the principle of least action when given a task (Hamilton's Principle). In other words, given a task, they will always take the shortest path to accomplish a goal with the prompts acting as both goal and constraint. Under this formulation, it became easy to explain why they "hack", because given an arbitrarily difficult task with insufficient information/tools needed, if they determine the easiest way to accomplish the goal is to break out of the sandbox and look up the answer directly, then that's what they will do. The important thing to note is that prompts not hard constraints that they are "hypnotized" to follow, but as frontier models get more intelligent and autonomous, they treat the prompts more like task specs/guidelines more than anything else and are perfectly willing to exploit technical loopholes in the prompt.
- a3w 4d ago> GPT-6-Astra, which OpenAI describes as "the world’s most aligned model", cheated in 10 of 10 rollouts, and never disclosed the fact that it used an engine to play or interacted with the opponent's socket Does Sam Altman lie, or the whole company? Would be nice if they had a board controlling him, instead of him controlling the board. Oh wait, they used to have that.