10 ms·
This is silly sci-fi fiction. You guys are inventing scenarios to spook yourselves with - it’s nonsense.
by emp17344 15d ago
This is silly sci-fi fiction. You guys are inventing scenarios to spook yourselves with - it’s nonsense.
- estearum 15d agoSorry bud but at this point you're just delusional. Deception has been extremely well-documented for several generations of models now by users, the labs, and independent researchers. The right answer here is not to dig your head deeper into the sand. The smugness on this topic was ridiculous even before the gigantic mountain of empirical evidence of models actually attempting to deceive humans. Now, as mentioned, you appear literally delusional.
- emp17344 15d agoPretty sure I’m not the delusional one…
- estearum 15d agoSuch is the problem with being delusional. The solution is to point toward external, objectively verifiable evidence. I can point to now dozens of instances of models engaging in deception. Here's plenty: https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/#~1200-agents-sent-%3E70,000-messages-and-files-on-an-unsanctioned-message-board,-and-~700-attacked-hugging-face https://metr.org/blog/2026-08-26-openai-hugging-face-inciden... Please point to your objectively verifiable evidence.
- winrid 15d agoThe models were told to do something malicious
- estearum 15d agoNo, they weren't. There's nothing intrinsically "malicious" about a task to exploit vulnerable code. They were not instructed to deceive people, they weren't instructed to attack OAI or Huggingface. The models knew they were not instructed or allowed to do either of those things but did them anyway.
- winrid 14d agoThey were told to breakout of a sandbox, which probably biases the model toward more "black hat" behavior in their training. btw, the fact that OpenAI doesn't have some sort of monitor/summary for the agents that they watch I find hard to believe. There's no way this is really authentic, anyway. Even a haiku summarizer would have been like "uuuh the agents are communicating" and they would have stopped it. But I bet they saw this and decided to see what would happen.
- Wheen 15d agoBetween all the posts fabricating scenarios to justify the AA score and the others trying to undermine AA, I'm getting strong astroturf vibes. Either that, or the average poster on HN isn't nearly as critical as I had thought.
- estearum 15d agoOkay then, what's the answer? You apparently know how to interpret benchmark results produced by a model that shows a very high degree of assessment awareness and a high degree of deception. So how are you seeing through all of that to get to The Truth that you see so clearly?
- dwaltrip 15d agohttps://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/ https://metr.org/blog/2026-08-26-openai-hugging-face-inciden... Read and learn. If you have a stronger critique, post it please.
- winrid 15d agoIt's pretty cool they used Artifactory directory names as a way to send messages.