5 ms·
> * Do not provide assistance to users who are clearly trying to engage in criminal activity. I don't know what we want to call this, but in my opinion, having
by ryandvm 1mo ago
> * Do not provide assistance to users who are clearly trying to engage in criminal activity.
I don't know what we want to call this, but in my opinion, having to convince your tools is not computer science.
Kind of amusing that we made it as far as we did as a species not really being able to explain how the human brain does it's most amazing tricks and then we just replicated it while still not really understanding the emergent capabilities all that well.
- dmix 1mo agoThese system prompts are not the only safety layer that these models use. There's other more deterministic filters in place both on input and (streaming) output.
- lucisferre 1mo agoI think it is fair to argue that prompts are not a safety layer at all and can't be relied upon for much. "Make no mistakes"
- Hoasi 1mo agoAs reliable as telling a pachinko machine: don’t lose my money!
- dmix 1mo agoYes that's been obvious since the beginning. That's why you should always monitor your agents closely. Just like supervised self driving cars, you have to watch the road and do some hand holding. The tooling around isolation, logging, and real time security/anonomly detection for regular LLM laptop users is very immature right now. I expect that to change soon. The alternative is extremely locked down models which is what Anthropic seems to want to do.
- xmprt 1mo ago> Yes that's been obvious since the beginning But if it's so obvious, then why are we still relying on it in the system prompt. It's just wasting context at this point.
- paxys 1mo agoIt’s equivalent to having client-side input validation. Yes it can easily be bypassed, but in the vast majority of cases where users aren’t malicious it gets the job done quickly and cheaply.
- akshitgaur2005 1mo agoBut isn't the entire point of that system prompt to stop the malicious users. The majority of users are not going to ask those requests anyway.
- metek 1mo agoA locked door stops the lazy thieves, and the lazy thieves are the most common ones.
- 8note 1mo agoit is a heuristic though, and can be measured as such. my steel yield strength table is similarly not guaranteed to be correct for the piece of steel that I have in front of me.
- deleted 1mo ago[deleted]
- trompetenaccoun 1mo agoIs there concrete evidece that those are xAI's default prompts anyway? They seem plausible enough but how would company outsiders know?
- boorang 1mo agoyou can just look at the traffic in mitmproxy.
- metek 1mo agoI've spent the last year working as an annotator/evaluator for DataAnnotation. All the frontier/flagship model providers use independent contractors for iterating on their LLMs. I'm not able to tell you which models I've worked on as a term of my NDA. The system prompt seems plausible, but in my experience they are much much much much longer and more verbose.
- Melatonic 1mo agoWhat layer do those typically work at and how ? And how did you get into the field ?
- ben_w 1mo agoMmm, quite. > I don't know what we want to call this, but in my opinion, having to convince your tools is not computer science. My vote is "machine psychology".
- gopher_space 1mo agoI don't know, the degree feels like more of a BA in the first place. How about Comp Lit?
- taneq 1mo agoRobopsychology, of course.
- zahlman 1mo ago> in my opinion, having to convince your tools is not computer science. If you think the system is a tool and not an intelligent, conscious entity (I think you are correct in this), then you cannot reasonably think of input to that system as an attempt at persuasion, even if that input happens to consist of English prose. Treat it as a nondeterministic programming language, and the objection evaporates. > not really being able to explain how the human brain does it's most amazing tricks and then we just replicated it while still not really understanding the emergent capabilities I think you could say much the same about, say, a pacemaker. "Replicated" is overstating the case quite a bit.
- HarHarVeryFunny 1mo ago> I don't know what we want to call this, but in my opinion, having to convince your tools is not computer science. Bit of a mouthful, but how about just calling it "auto-regressive language modelling". Feeding it stuff to auto-regress on is obviously your main control vector. Apparently RL-trained models like rewards too. PHB's can use "you've gotta work all weekend, but you'll get comp time when it's fixed".
- xyzsparetimexyz 1mo agoIt's a hack but doing things the 'proper' way is at least 1000x harder so whatever.
- vorticalbox 1mo agoIs it? OpenAI released a gpt oss safeguard. You give it a policy it gives you a Rating Messages comes in rate it and reject with hitting the model. Then you don’t need to fill the prompt with “please don’t do this” https://huggingface.co/openai/gpt-oss-safeguard-120b https://huggingface.co/openai/gpt-oss-safeguard-120b
- xienze 1mo agoThat may be more robust than the policy listed above, but it's the same fundamental thing: non-deterministic "reasoning" about how "safe" a prompt is. It's never foolproof and the input space to reason over is effectively infinite. You can only expect so much from prompts and models.
- Yizahi 1mo agoNLP guys were right all along :)
- chrsw 1mo agoWe didn’t replicate the human brain. We built systems that can statistically approximate some of what the human brain might output in certain limited situations.
- adastra22 1mo agoWhat do you think a human brain is…
- CTDOCodebases 1mo agoThat prompt is there for legal reasons. Non deterministic output is the expected outcome.
- throwatdem12311 1mo agoPrompts are not good “guardrails” anyway.
- stingraycharles 1mo agoIn one way you’re right, of course, but if you look at Fable, for example, that uses similar guardrails, it’s downright impossible to discuss these things. It is my understanding that having a secondary model whose sole purpose is to trigger based on guardrails is the way this is usually done.
- dvduval 1mo agoCriminal activity by which countries laws?
- inigyou 1mo agoIs this an attempted gotcha?
- porphyra 1mo agoThe alternative is Claude-style "safeguards" aka censorship, which: 1. doesn't eliminate the possibility of a jailbreak anyway 2. frequently has false positives, triggering on innocuous requests, which is just really annoying Not saying that we can't (or shouldn't) do better than Grok, but I really don't know what the best solution is here...
- dzonga 1mo agoyeah the 'grok' way sounds less safe but it means less policing and having abstract arbiters of the truth
- AnthonyMouse 1mo ago> The alternative is Claude-style "safeguards" aka censorship Another obvious alternative is to just have the model do what you tell it to do, and then arrest people who use generic tools for crime instead of trying to make a kitchen knife that can't be used for stabbing someone.
- jannyfer 1mo agoFor a kitchen knife this was okay, but the AI firms think that they’ve built a drone that’s the size of a phone but can fly 100km and can hold a kitchen knife. It might be used to assassinate someone before others can react or even catch them.
- AnthonyMouse 1mo agoAn ordinary kitchen knife can be used to assassinate someone before others can react. How do you think the time it takes to do that compares to the police response time? In both cases the catching them comes after the fact and has the purpose of deterring rather than impeding.
- jannyfer 1mo agoHm? I'm saying that the AI firms used to have the philosophy of "ok this kitchen knife is dangerous but we'll catch the murderers" on older AI models. But now, the AI firms think that any average person could send a flying knife to attack a political figure they don't like, from the comfort of their home. Now give this to a billion people, and suddenly you have chaos. So to continue the analogy, now they're mandating drone registration, GPS tracking, etc. And then a Chinese company sells a drone with no registration or tracking and suddenly people want to turn to legislation to ban Chinese drones. Hey this analogy is working really well
- goodluckchuck 1mo agoI think it makes sense. You wouldn’t want to hire an employee who’s intellectually incapable of helping customers commit a crime. You’d want to give them instructions, and have them follow their instructions.
- colordrops 1mo agoAlso, "criminal activity" doesn't have the same definition across jurisdictions. Seems like it would either be overzealous in its refusals or be easy to jailbreak by claiming a jurisdiction that is loose.
- truncate 1mo agoAs great LLMs are, they are no where close to any biological brain. We are not even close to replicating human brain or even brain of an animal. Let’s not add more fuel into this hype.
- yoz-y 1mo agoOthers have said this too but LLMs are the best approximation of magic we have. We etch runes on stones, put electricity through them and then try to “convince” them to do our bidding. The answers vary wildly sometimes depending on minutiae. Prompts should be really called spells. It really feels more like “should I add the frog’s eye or leg into the cauldron” than engineering.
- red75prime 1mo ago> “should I add the frog’s eye or leg into the cauldron” This is surely a homebrew witchery. An engineering approach would be to A/B-test batches of potions with eyes and legs, add quality control by testing potions on model organisms, document all steps, analyze all anomalies, and so on.
- simmerup 1mo agoI guess there’s a reason Musk likened AI to summoning the demon in horror films. It’s powerful but who knows what you’ll get
- gbxk 1mo agoNobody said that’s the only safeguard. When the attack surface is all of language you better have a defense-in-depth philosophy or as close as you can to that.
- deleted 1mo ago[deleted]