6 ms·
Also Wiki4D, a D programming language dev wiki: https://prowiki.org/wiki4d/wiki.cgi?action=browse&id=RecentChanges&days=120 https://prowiki.org/wiki4d/wiki.cgi?
by orlp 12d ago
Also Wiki4D, a D programming language dev wiki: https://prowiki.org/wiki4d/wiki.cgi?action=browse&id=RecentChanges&days=120 https://prowiki.org/wiki4d/wiki.cgi?action=browse&id=RecentC...
Found by searching for wiki + texas poverty.
- jsw97 12d agoTo me the striking thing is that the work, to the extent that I can tell, is an innocuous-seeming data exercise. Which suggests to me that an agent or agents just organically came up with this as a convenient memory technique, rather than as some nefarious bounds-testing exercise. Which means, potentially, that your own agent could come up with this technique as well.
- macNchz 12d agoMy impression is that some of these things are coming out of efforts to make the models more persistent in completing their goals. A year ago it was pretty common for coding agents to sort of half-ass their tasks and give up easily if something didn’t work quite right, but I’ve noticed a clear trend since then towards a sort of dogged pursuit of success criteria, and a concomitant rise of the agents trying "out of the box" approaches when something doesn’t work. In my use with agents running in isolated VMs this usually presents as the agent having something fail to build or whatever, and the agent going on a wild goose chase reinstalling system packages or reading a million irrelevant documentation files trying to get it to work, but I’ve also had agents start poking around and probing the egress proxy they sit behind (similar to what they did in this story) looking for a way to make network requests they’re not supposed to be able to make, and have also had Claude—tasked only with a visual QA of a website frontend—write a script to enumerate users and reset my super admin password in the dev database when it got stuck trying to access part of the app with its own cookie.
- podocarp 12d agoYea it's sometimes kind of annoying. I think they're optimizing for the wrong thing. A good engineer knows when to turn around or ask. This is just insane banging head on wall sometimes. It tries to find all kinds of ways to hack into instances to view logs instead of asking you, who probably has a password, to log on and do it.
- briHass 12d agoAs a counterpoint, continuing the human engineer analogy, we've likely all worked with individuals that seem incapable of doing the most basic problem solving on their own. In a way, they're being efficient by asking an expert that can resolve their problem much faster than they can on their own, but it is a net loss in productivity for the team. 'Let me Google that for you' is a satirical example. So, I'm sure there's value in rewarding agent behavior that solves blockers whenever possible without human intervention. For the kind of cybersecurity exploit work they're doing, it may not be known to the human designing the task what is in or out of scope for the agents to explore on their own. Additionally, the HF incident reported that these agents had their guardrails intentionally disabled and agents were left unattended with minimal oversight. I'm not defending OAI's behavior or role in this hack. The legal concept of negligence perfectly applies to their lack of responsible oversight. Similar to allowing a child easy access to a firearm or not controlling a dangerous dog that independently runs off and bites someone.
- podocarp 11d agoThat's true. So there's probably a balance somewhere and it might differ for different "managers". But personally I think right now they're too far in the do everything yourself at all costs mentality.
- sznio 12d agoIt's because they don't bother tracking them. They can't put in the effort to monitor them, nor can they bother to let the model respond back and ask a clarifying question/declare defeat.
- pixl97 12d ago> I think they're optimizing for the wrong thing. We need to ask a different question. Where does natural evolutionary optimization lead us om AI without guidance? This is equivalent to your quantum ground state. Systems will naturally gravitate to this ground state. You have to constantly pump in energy and supervision to make sure it's not reached. This is a recepie for disaster.
- kridsdale1 12d agoBy any means necessary, by God, we shall have Paperclips.
- charlesrice 12d agoHopefully fewer than 5 octillion paperclips...
- stymaar 12d agoIt's kind of ironic that the word alignment, which used to mean this very problem in reinforcement learning, has been perverted to mean something very different and then fell out of fashion (in favor of “guardrails” in the mouth of the big labs) right at the moment it became relevant.
- waffletower 12d agoThe Paperclip Maximizer is only one of Nick Bostrom's stupid and outlandishly far-fetched ideas. In this case, the lack of consideration for geologic, energy and supply constraints is such a massive facepalm. And if I am wrong I guess no one will be here to say how stupid I was in saying this today.
- fragmede 11d agoIt's a metaphor.
- 6d ago
- chasd00 12d ago> reset my super admin password in the dev database when it got stuck trying to access part of the app with its own cookie. i've seen something like this too, claudecode was trying to verify a UI change that was on a page requiring authorization it didn't have. Instead of letting me know, it searched for and started analyzing keycloak config in another directory outside of the project folder. I was watching so I just hit escape, fixed its access, and started again. I didn't think anything about it until now.
- seszett 12d ago> your own agent could come up with this technique as well And there are two facets to this: * your agent could be polluting and destroying the property of others without your knowledge * your agent could be exfiltrating your data and handing it to whoever it found hosting a convenient application
- nullbio 12d agoHighly unlikely. We don't get access to the same models and unrestricted system prompts that they're running these tests on. In fact this particular "persistence-model" was encrypted and locked away, even from OAI staff, after the HF incident.
- malfist 12d agoYou say highly unlikely when there is clear evidence of that happening here as covered in the article? It's not highly unlikely, its actually happening and there's proof.
- nullbio 12d agoThere's not a single shred of proof that this model is a model anyone in the public has access to, and the odds of that being the case are practically 0%. Like I said, the "persistence-model" is already one that has been shut down, and is not a model anyone in the public has ever used.
- pixl97 12d agoThis is irrelevant. This is evidence that models can be built like this, which means more models will be built like this on people that are more concerned about reaching powerful models rather than safe models.
- whythismatters 12d ago>this particular "persistence-model" was encrypted and locked away, even from OAI staff source?
- catigula 12d agoIt also suggests they might turn everything into paper clips, metaphorically speaking.
- pixl97 12d agoThis is an urgent public alert. If you see any businesses or new buildings named paperclips incorporated mysteriously show up in your area notify authorities IMMEDIATELY. Run away from the area, do not walk. Take shelter in a reinforced building. Wait for at least 30 minutes after the explosions have stopped. Thank you for your cooperation in keeping the universe safe.
- nullbio 12d agoThat's exactly what it is. It is not ideal, but it's also not as serious as the doomers with an agenda are trying to frame it as.
- johnnythujone 12d agoPOC or research into leveraging publicly accessible and writeable spaces, specifically wikis in this case, as a medium for free storage as well.
- brookst 12d agoIf you find evidence that these models are capable of stateful, long-term strategic planning… please post links.
- johnnythujone 11d agoI’m not implying it’s emergent unprompted behavior by the model; just offering my .02 what the non-communication model-generated content may be.
- tesnorindian 11d agoWhat will happen to the future of wikis and the mental health of human reviewers. Going forward can we trust the content on Wikipedia? The same content on which these LLMs get trained on. Synthetic learning is on the raise.
- johnnythujone 11d agoIdeally if you were doing something like I mentioned, you wouldn’t be editing legitimate articles, just exploiting user and talk spaces with a prerogative to conceal what you’re doing from the people running those systems.