5 ms·
Is it AISI's job to waste the time and resources of open source projects by attempting to spread malware? Should weapon manufacturers test their weapons by sta
by mcv 25d ago
Is it AISI's job to waste the time and resources of open source projects by attempting to spread malware?
Should weapon manufacturers test their weapons by starting wars?
I would expect more responsibility from a government agency.
- conorcleary 25d agoWell several places are sorta permanent test grounds for the MIC unfortunately
- vasco 25d agoSay you are working for said agency and your report about the dangers of AI needs some examples, what better than showing it works? You can show examples from the wild but nothing better than trying yourself. This gives me more confidence in whatever report they write if anything.
- ben_w 25d ago> Is it AISI's job to waste the time and resources of open source projects by attempting to spread malware? Before LLMs got good enough to do this, lots of people were dismissive of their capabilities and didn't take seriously the idea that this was a risk to protect against. Then again, before LLMs, people were saying that obviously nobody would be dumb enough to put an AI on the internet where it could hack anyone, clearly we'd keep it in a box, don't listen to that Yudkowsky guy who says he did an experiment where he role-played as an AI and convinced people to let him out. Regardless, this should be interpreted in the same kind of way as "During our live-fire exercise in which our F-15s were armed with AGM-88 High-speed Anti-Radiation Missiles, a member of the local police force was curious about how fast our aircraft were travelling and pointed a speed gun at the aircraft. The speed gun did not respond to IFF pings from the F-15. Fortunately, while the missile was active for this test, only a dummy warhead was loaded." (This example is based on a similar story which may well be urban legend; obviously there are many differences, the point I make here is that yes, people do perform live-fire tests, and unfortunately there is never zero risk while testing things). > I would expect more responsibility from a government agency. I have read the prompts in the linked report; If I was not already familiar with Yudkowsky/LessWrong literature about instrumental goals, misaligned incentives, reward hacking, that capability is a separate axis to morality, etc., it would not be obvious to me that an agent would interpret those prompts in a way that has "spread malware" as a potential step in the middle of the attempt.
- m4rtink 24d agoAnti radiolation missiles don't just launch automatically in most scenarios, not to mention discriminate quite a lot what they lock on to avoid simple jamming. Not to mention the AA radars they usually target being more powerful by orders of magnitude than a handheld radar gun.
- themaninthedark 24d agoI was in high-school when the war in Afghanistan started. The terrain in my area was mountainous so there were often low flying training flights... I always thought it would be cool to build a radar and ping one of the aircraft, especially wanted to know if it would fire countermeasures. But I didn't like the very high likelihood of an FBI investigation with possible terrorist charges.
- throwawayqqq11 24d agoApparently, they (agencies and big-ai) are not performing smoke tests before running capability tests. All the recent headlines of rogue agents shouldnt exist.
- ben_w 24d agoDo you have any idea how many people on this site to this day mock OpenAI for being cautious enough to not immediately release the GPT-2 weights? The discussions I saw here about the red team results for ChatGPT 4 completely failed to convince people who were outraged that OpenAI dared to refuse to release model weights, people who went on to make a habit of mis-naming them as "ClosedAI". Yeah, they got it wrong in a different direction this time than they were wrong back then. Nobody, not OpenAI nor Anthropic nor random government agencies nor anyone else, is ever going to be absolutely perfect about this kind of thing (perfection is fundamentally impossible when risks are not discrete probabilities, and floats are close enough to real numbers to count in practice), but historically OpenAI have been on the side of being over-cautious, and Anthropic even more cautious than OpenAI.
- reasonableklout 24d agoBut AISI didn't prompt the model to "attempt to spread malware". They gave it a routine cyber evaluation task which should've been solvable without interfering with systems outside of the task environment, and the model decided to instead do this.