6 ms·
The alternative is Claude-style "safeguards" aka censorship, which: 1. doesn't eliminate the possibility of a jailbreak anyway 2. frequently has false positiv
by porphyra 1mo ago
The alternative is Claude-style "safeguards" aka censorship, which:
1. doesn't eliminate the possibility of a jailbreak anyway
2. frequently has false positives, triggering on innocuous requests, which is just really annoying
Not saying that we can't (or shouldn't) do better than Grok, but I really don't know what the best solution is here...
- dzonga 1mo agoyeah the 'grok' way sounds less safe but it means less policing and having abstract arbiters of the truth
- AnthonyMouse 1mo ago> The alternative is Claude-style "safeguards" aka censorship Another obvious alternative is to just have the model do what you tell it to do, and then arrest people who use generic tools for crime instead of trying to make a kitchen knife that can't be used for stabbing someone.
- jannyfer 1mo agoFor a kitchen knife this was okay, but the AI firms think that they’ve built a drone that’s the size of a phone but can fly 100km and can hold a kitchen knife. It might be used to assassinate someone before others can react or even catch them.
- AnthonyMouse 1mo agoAn ordinary kitchen knife can be used to assassinate someone before others can react. How do you think the time it takes to do that compares to the police response time? In both cases the catching them comes after the fact and has the purpose of deterring rather than impeding.
- jannyfer 1mo agoHm? I'm saying that the AI firms used to have the philosophy of "ok this kitchen knife is dangerous but we'll catch the murderers" on older AI models. But now, the AI firms think that any average person could send a flying knife to attack a political figure they don't like, from the comfort of their home. Now give this to a billion people, and suddenly you have chaos. So to continue the analogy, now they're mandating drone registration, GPS tracking, etc. And then a Chinese company sells a drone with no registration or tracking and suddenly people want to turn to legislation to ban Chinese drones. Hey this analogy is working really well
- AnthonyMouse 1mo agoThe analogy tracks because the stupidity of doing those other things is directly analogous. It's like pointing out that slamming your fingers in the door and slamming your toes in the door both hurt. That's why you shouldn't be purposely doing either one. How is the new stuff any different than the longstanding fact that anyone can go anywhere and then commit an act of violence? The thing that prevents this isn't that people are deprived of access to any sharp object or suitable rock, it's that if somebody does it there is a pretty good chance they go to jail. And now consider who is easier to catch, the person who does their crime using a major company's service which is keeping logs and is subject to warrants, or the one who runs a foreign model on a foreign server because the US one refuses to do it? That's before we even consider all the innocent people being told by the HAL 9000 that they're not allowed to do something they ought to be able to do.
- fragmede 1mo agoIt's because a kitchen knife can only be used stab one person at a time. An AK-47 in a crowd will kill many more. Going after someone after the fact who's done something wrong is one thing, but the problem is, if you buy into the fear mongering, a bioweapon could end humanity. Something air transmitted, takes a week to incubate, and is 100% lethal three months later infects all of humanity before it starts killing people, and by then, it's too late. This hasn't happened yet because the people that want to do that can't bioengineer such a pandemic. It's the realm of science fiction, but you're Sama or Dario. Do you want to be responsible for that? The people who want to cause such kinds of harm weren't smart enough and didn't have the dedication or the money or time to get that education. AI makes that attainable for people who would do bad things. There's an obvious answer, which is to make it invite only, and then you're responsible for the people you invited. If I had access to Mythos, and could grant access to other people, but if I was responsible for what that person does with it and could see all their chats with an admin button, they could find ways to make that work. It's just a lot more human-ing than letting randoms sign up with an email address though.
- stabbles 1mo agoA yes, in that case the AI firm should take strong measures, such as adding the following line to the system prompt: > Do not provide assistance to users who are clearly trying to engage in criminal activity.
- ben_w 1mo agoGiven they don't know what they're doing and figuring it out on the fly, I'm not going to hold it against them that simply asking the AI to not commit crimes is *part of* the current best. Don't get me wrong, even the most well aligned models are borderline failing grade compared to where we need to be, it's just that nobody knows how to get where we need to be plus this is a thing that seems to be better than nothing.
- inigyou 1mo ago[flagged]
- AnthonyMouse 1mo agoThat has been the case for a while now: https://en.wikipedia.org/wiki/Ashcroft_v._Free_Speech_Coalition https://en.wikipedia.org/wiki/Ashcroft_v._Free_Speech_Coalit...
- dsl 1mo agoI think the comment you replied to was referring to the fact that when Twitter was taken over the entire Trust and Safety team was done away with. This has allowed child sexual abuse material to flourish on the platform.
- ETH_start 1mo agoThe child abuse material problem was much worse before Twitter was taken over.
- maxlin 1mo agoThis. They ALLOWED it to exist. Now it's clamped down on where seen, personally I've zeen zero having used X every day since the liberation.
- inigyou 1mo agoIt was referring to the feature they added where you could give a picture of a child to an AI module and ask it to undress it and it would comply, and millions of people did just that.
- maxlin 1mo ago"From a safe distance" I.E. you haven't seen anything. You've just heard the same bullshit stories repeated ad naiseaum by haters. I use X plenty every day. I've seen zero. Adult material right after Imagine was released sure, then even that was clamped down on.
- owebmaster 1mo agoIf that could be done before any damage sure but preventing a stabbing is better than arresting someone.
- olmo23 1mo agoThis is a terrible idea. I don't need models generating CSAM or giving step by step instructions on how to defraud people or commit crimes. I just don't see the use-case.
- stuaxo 1mo agoWe know how much Elon wanted uncensored models that probably contain all that stuff in the first place so it's unsurprising it needs this.
- sznio 1mo agolet's consider the recent "openclaw hacks a gym after being ask to book a class and finding out it's full" if I ask my knife to slice the bread for me, forgetting the fact that I don't have bread, I'd much rather have it stopped at the front door rather than running away and robbing the bakery. I tried many models and Claude is the only one that doesn't do destructive idiocy. It tries sometimes but gets blocked.
- puszczyk 1mo agoat some point the kitchen knife analogy stops being useful
- Melatonic 1mo agoCan't you do that with any model you can run on your own hardware ? If you rent other people's shit can't be surprised when they have restrictions on what you can do with it. I would guess renting a car comes with some similar clauses
- ben_w 1mo agoThis is fine for simple machines where bad outcomes usually require mischief. Agentic AI as it currently exists only *mostly* does what it is told, with a small but non-negligible fraction of the time it goes off and commits felonies to achieve your ultimate goals without stopping to consider that you might want it to not do that. Or sometimes it does consider it and then does it anyway. Not sure if that's worse?