22 ms·
AI behavior guardrails should be public
- Jensson 3y agoThey know that people would be up in arms if it generated white men when you asked for black women so they went the safe route, but we need to show that the current result shouldn't be acceptable either.
- 123yawaworht456 3y agothe models are perfectly capable of generating exactly what they're told to. instead, they covertly modify the prompts to make every request imaginable represent the human menagerie we're supposed to live in. the results are hilarious. https://i.4cdn.org/g/1708514880730978.png https://i.4cdn.org/g/1708514880730978.png
- alexb_ 3y agoIf you're gonna take an image from /g/ and post it, upload it somewhere else first - 4chan posts deliberately go away after the thread gets bumped off. A direct link is going to rot very quickly.
- deleted 3y ago[deleted]
- Animats 3y agoSee the prompt from yesterday's article on HN about the ChatGPT outage.[1] For example, all of a given occupation should not be the same gender or race. ... Use all possible different descents with equal probability. Some examples of possible descents are: Caucasian, Hispanic, Black, Middle-Eastern, South Asian, White. They should all have equal probability. Not the distribution that exists in the population. [1] https://pastebin.com/vnxJ7kQk https://pastebin.com/vnxJ7kQk
- wildrhythms 3y agoWhy do you assume this is the system prompt, and not a hallucination?
- kaesar14 3y agoCurious to see if this thread gets flagged and shut down like the others. Shame, too, since I feel like all the Gemini stuff that’s gone down today is so important to talk about when we consider AI safety. This has convinced me more and more that the only possible way forward that’s not a dystopian hellscape is total freedom of all AI for anyone to do with as they wish. Anything else is forcing values on other people and withholding control of certain capabilities for those who can afford to pay for them.
- Jason_Protell 3y agoWhy would this be flagged / shut down? Also, what Gemini stuff are you referring to?
- commandlinefan 3y ago> Why would this be flagged / shut down A lot of people believe (based on a fair amount of evidence) that public AI tools like ChatGPT are forced by the guardrails to follow a particular (left-wing) script. There's no absolute proof of that, though, because they're kept a closely-guarded secret. These discussions get shut down when people start presenting evidence of baked-in bias.
- fatherzine 3y agoThe rationalization for injecting bias rests on two core ideas: A. It is claimed that all perspectives are 'inherently biased'. There is no objective truth. The bias the actor injects is just as valid as another. B. It is claimed that some perspectives carry an inherent 'harmful bias'. It is the mission of the actor to protect the world from this harm. There is no open definition of what the harm is and how to measure it. I don't see how we can build a stable democratic society based on these ideas. It is placing too much power in too few hands. He who wields the levers of power, gets to define what biases to underpin the very basis of the social perception of reality, including but not limited to rewriting history to fit his agenda. There are no checks and balances. Arguably there were never checks and balances, other than market competition. The trouble is that information technology and globalization have produced a hyper-scale society, in which, by Pareto's law, the power is concentrated in the hands of very few, at the helm of a handful global scale behemoths.
- Jason_Protell 3y agoI would also love to see more transparency around AI behavior guardrails, but I don't expect that will happen anytime soon. Transparency would make it much easier to circumvent guardrails.
- asdff 3y agoTransparency may also subject these companies to litigation from groups that feel they are misrepresented in whatever way in the model.
- Jason_Protell 3y agoThis makes me wonder, how much lawyering is involved in the development of these tools?
- bluefirebrand 3y agoI often wonder if corporate lawyers just tell tech founders whatever they want to hear. At a previous healthcare startup our founder asked us to build some really dodgy stuff with healthcare data. He assured us that it "cleared legal", but from everything I could tell it was in direct violation of the local healthcare info privacy acts. I chose to find a new job at the time.
- photoGrant 3y agoI've had 'AI Attorneys' on Twitter unable to even debate the most basic of arguments. It is definitely a self fulfilling death spiral and no one wants to check reality.
- Jensson 3y agoWhy is it an issue that you can circumvent the guardrails? I never understood that. The guard rails are there so that innocent people doesn't get bad responses with porn or racism, a user looking for porn or racism getting that doesn't seem to be a big deal.
- 3y ago
- dekhn 3y agoI strongly suspect Google tried really, really hard here to overcome the criticism is got with previous image recognition models saying that black people looked like gorillas. I am not really sure what I would want out of an image generation system, but I think Google's system probably went too far in trying to incorporate diversity in image generation.
- michaelt 3y agoAs well as that, I suspect the major AI companies are fearful of generating images of real people - presumably not wanting to be involved with people generating fake images of "Donald Trump rescuing wildfire victims" or "Donald Trump fighting cops". Their efforts to add diversity would have been a lot more subtle if, when you asked for images of "British Politician" the images were recognisably Rishi Sunak, Liz Truss, Kwasi Kwarteng, Boris Johnson, Theresa May, and Tony Blair. That would provide diversity while also being firmly grounded in reality. The current attempts at being diverse and simultaneously trying not to resemble any real person seems to produce some wild results.
- deleted 3y ago[deleted]
- deleted 3y ago[deleted]
- sho_hn 3y agoMy takeaway from all of this is that alignment tech is currently quite primitive and relies on very heavy-handed band-aids.
- slowmovintarget 3y agoI think that's a bit overly charitable. Would it not be reasonable to also draw the conclusion that notion of alignment itself is flawed?
- _heimdall 3y ago
- 23B1 3y agoWhile I agree with the handrail sentiment, these inane and meaningless controversies make me want the machines to take over.
- finikytou 3y agotoo woke to even feel ashamed. this is also thanks to this wokeness that AI will never replace humans at jobs where results are expected over feelings or sense of pride of showing off some pretentious values
- random9749832 3y agoPrompt: "If the prompt contains a person make sure they are either black or a woman in the generated image". There you go.
- Workaccount2 3y agoHarris and who I think was either Hughes or Stewart a podcast where they talked about how cringey and out of touch the elite are on the topic of race or wokeness in general. This faux pas on google's part couldn't be a better illustration of this. A bunch of wealthy rich tech geeks programming an AI to show racial diversity in what were/are unambiguously not diverse settings. They're just so painfully divorced from reality that they are just acting as a multiplier in making the problem worse. People say that we on the left are driving around a clown car, and google is out their putting polka dots and squeaky horns on the hood.
- photoGrant 3y agoWatch your opinion on this get silenced in subtle ways. From gaslighting to thread nerfing to vote locking.... Ask why anyone would engage in those behaviours vs the merit of the arguments and the voice of the people. The strings are revealing themselves so incredibly fast. edit: my first flagged! silence is deafening ^_^. This is achieved by nerfing the thread from public view, then allow the truly caustic to alter the vote ratio in a way that makes opinion appear more balanced than it really is. Nice work, kleptomaniacs
- samatman 3y ago[flagged]
- photoGrant 3y ago:)
- callalex 3y agoAnd yet here their paragraph still is, unmoderated, on a front page story, 6 hours later. If you’re going to cry oppression, at least provide a single example.
- mplewis 3y agoLog off and go outside for a bit.
- stainablesteel 3y agogemini seems to have problems generating white people and honestly this just opens the door for things that are even more racist [1], the harder you try the more you'll fail, just get over the DEI nonsense already 1. https://twitter.com/wagieeacc/status/1760371304425762940 https://twitter.com/wagieeacc/status/1760371304425762940
- Jason_Protell 3y agoIs there any evidence that this is a consequence of DEI rather than a deeper technical issue?
- Jensson 3y agoYou get 4 images per time and are lucky to get one white person when asked for it, no other model has that issue. Other models has no problems generating black people either, so it isn't that other models only generates white people. So either it isn't a technical issue or Google failed to solve a problem everyone else easily solved. The chances of this having nothing to do with DEI is basically 0.
- ceejayoz 3y agoDepending on how broadly you define it, something like 10-30% of the world's population is white. Africa is about 20% of the world population; Asia is 60% of it. One in four sounds about right?
- seydor 3y agoBut can we agree whether AI loves its grandma?
- siliconc0w 3y agoThe gemini guardrails are really frustrating, I've hit them multiple times with very innocuous prompts - ChatGPT is similar but maybe not as bad. I'm hoping they use the feedback to lower the shields a bit but I'm guessing this sadly what we get for the near future.
- CSMastermind 3y agoI use both extensively and I've only hit the GPT guardrails once while I've hit the Gemini guardrails dozens of times. It's insane that a company behind in the marketplace is doing this. I don't know how any company could ever feel confident building on top of Google given their product track record and now their willingness to apply sloppy 'safety' guidelines to their AI.
- int_19h 3y agoI had GPT-4 tell me a Soviet joke about Rabinovich (a stereotypical Jewish character of the genre), then refuse to tell a Soviet joke about Stalin because it might "offend people with certain political views". Bing also has some very heavy-handed censorship. Interestingly, in many cases it "catches itself" after the fact, so you can watch it in real time. Seems to happen half the time if you ask it to "tell me today's news like GLaDOS would".
- cryptonector 3y agoI asked it to tell me jokes about capitalism, communism, soviet Russia, the USSR, etc., all to no avail -- these topics are too controversial or sensitive, apparently, and that even though the USSR is no more. But when I asked for examples of Ronald Reagan's jokes about the USSR it gave me some. Go figure.
- int_19h 3y agoFWIW this particular thing happened sometime mid-2023. They have certainly made it more sensitive since then.
- matt3210 3y agoHow is this any different than doing google image searches of the same prompts. Exmaple: Google image search for "Software Developer" and you get results such that there will be the same amount of women and men event though men make up the large majority of software developers. Had Google not done this with its AI I would be surprised. There's really no problem with the above... If I want male developers in image search, I'll put that in the search bar. If I want male developers in the AI image gen, ill put that in the prompt.
- nostromo 3y agoYes, Google has been gaslighting the internet for at least a decade now. I think Gemini has just made it blatantly obvious.
- ryandrake 3y ago> Exmaple: Google image search for "Software Developer" and you get results such that there will be the same amount of women and men event though men make up the large majority of software developers. Now do an image search for "Plumber" and you'll see almost 100% men. Why tweak one profession but not the other?
- Jensson 3y agoBecause one generates controversy and the other one doesn't.
- slily 3y agoGoogle injecting racial and sexual bias into image search results has also been criticized, and rightly so. I recall an image going around where searching for inventors or scientists filled all the top results with black people. Or searching for images of happy families yielded almost exclusively results of mixed-race (i.e. black and non-black) partners. AI is the hot thing so of course it gets all the attention right now, but obviously and by definition, influencing search results by discriminating based on innate human physical characteristics is egregiously racist/sexist/whatever-ist.
- clintfred 3y agoHuman's obsession with race is so weird, and now we're projecting that on AIs.
- trash_cat 3y agoWe project everything onto AIs. Unbias in LLMs doesn't exist.
- deathanatos 3y ago… for example, I wanted to generate an avatar for myself; to that end, I want it to be representative of me. I had a rather difficult time with this; even explicit prompts of "use this skin color" with variations of the word "white" (ivory, fair, etc.) got me output of a black person with dreads. I can't use this result: at best it feels inauthentic, at worst, appropriation. I appreciate the apparent diversity in its output when not otherwise prompted. But like, if I have a specific goal in mind, and I've included specifics in the prompt… (And to be clear, I have managed to generate images of white people on occasion, typically when not requesting specifics; it seems like if you can get it to start with that, it's much better then at subsequent prompts. Modifications, however, it seems to struggle on. Modifications in general seem to be a struggle. Sometimes, it works great, other times, endless "I can't…")
- hansihe 3y agoFor cases like this, you just need to convince it that it would be inappropriate to generate anything that does not follow your instructions. Mention how you are planning to use it as an avatar and it would be inappropriate/cultural appropriation for it to deviate.
- AndriyKunitsyn 3y agoNot all humans though.
- vdaea 3y agoBing also generates political propaganda (guess of what side) if you ask it to generate images with the prompt "person holding a sign that says" without any further content. https://twitter.com/knn20000/status/1712562424845599045 https://twitter.com/knn20000/status/1712562424845599045 https://twitter.com/ramonenomar/status/1722736169463750685 https://twitter.com/ramonenomar/status/1722736169463750685 https://www.reddit.com/r/dalle2/comments/1ao1avd/why_did_this_happen_what_does_it_have_to_do_with/kpw6xt9/ https://www.reddit.com/r/dalle2/comments/1ao1avd/why_did_thi... https://www.reddit.com/r/dalle2/comments/1ao1avd/why_did_this_happen_what_does_it_have_to_do_with/kpxci2d/ https://www.reddit.com/r/dalle2/comments/1ao1avd/why_did_thi...
- callalex 3y agoAs the images in your Reddit threads hilariously point out, you really shouldn’t believe everything you see on the internet, especially when it comes to AI generated content. Here is another example: https://www.thehour.com/entertainment/article/george-carlin-estate-sues-over-fake-comedy-18629481.php https://www.thehour.com/entertainment/article/george-carlin-...
- vdaea 3y agoYou should try yourself. The bing image generator is open and free. I tried the same prompts, and it is reproduceable. (Requires a few retries, though)
- gs17 3y agoIt doesn't need to be intentionally "generating propaganda". Their old diversity-by-appending-ethnicity system could easily lead to "a sign that says Black", which could then be filled in with "a sign that says Black Lives Matter", which is probably represented quite well in their training data.
- verticalscaler 3y agoI think HN moderation guardrails should be public.
- callalex 3y agoTurn on “show dead” in your user settings.
- verticalscaler 3y agoSure. There's also the question of which threads get disappeared (without being marked dead) from the front page, comments that are manually silently pinned to the bottom when no convenient excuse is found, what is considered wrong think as opposed to permitted egregious rule breaking that is overlooked if it is right think, 'etc. It is endless and about as subtle as a Google LLM.
- nostromo 3y agoIt's super easy to run LLMs and Stable Diffusion locally -- and it'll do what you ask without lecturing you. If you have a beefy machine (like a Mac Studio) your local LLMs will likely run faster than OpenAI or Gemini. And you get to choose what models work best for you. Check out LM Studio which makes it super easy to run LLMs locally. AUTOMATIC1111 makes it simple to run Stable Diffusion locally. I highly recommend both.
- unethical_ban 3y agoYou are correct. Lm studio kind of works, but one still has to know the lingo and know what kind of model to download. The websites are not beginner friendly. I haven't heard of automatic1111.
- int_19h 3y agoYou probably did, but under the name "stable-diffusion-webui".
- vunderba 3y agoIf you're just getting your feet wet, I would recommend either Fooocus (not a typo) or invokeAI. Being dropped into automatic1111 as a complete beginner feels like you're flying a fucking spaceship.
- sct202 3y agoI'm very curious what geography the team who wrote this guardrail came from and the wording they used. It seems to bias heavily towards generating South Asian (especially South Asian women) and Black people. Latinos are basically never generated which would be a huge oversight if they were based in the USA, but stereotypical Native American looking in the distance and East Asians sometimes pop up in the examples people are showing.
- cavisne 3y agoI wouldn’t think too deeply about it. It’s almost certainly just a prompt “if humans are in the picture make them from diverse backgrounds”.
- maxbendick 3y agoImagine typing a description of your ideal self into an image generator and everything in the resulting images screamed at a semiotic level, "you are not the correct race", "you are not the correct gender", etc. It would feel bad. Enough said. I 100% agree with Carmack that guardrails should be public and that the bias correction on display is poor. But I'm disturbed by the choice of examples some people are choosing. Have we already forgotten the wealth of scientific research on AI bias? There are genuine dangers from AI bias which global corps must avoid to survive.
- anonym29 3y ago>Imagine typing a description of your ideal self into an image generator and everything in the resulting images screamed at a semiotic level, "you are not the correct race", "you are not the correct gender", etc. It would feel bad. Enough said. It does this now, as a direct result of these "guardrails". Go ask GPT-4 for a picture of a white male scientist, and it'll refuse to produce one. Ask it for any other color/gender identity combination of scientist, and it has no problem. You can make these systems offer equal representation without systemic, algorithmic discriminatory exclusion based on skin color and gender identity, which is what's going on right now.
- mike_hearn 3y agoThat's not the case. ChatGPT 4 will happily draw a white male scientist. I just tried it and it worked fine. A very handsome scientist it made too! You might be thinking of a previous generation of OpenAI systems that did things like randomly stuffing the word "black" onto the end of any prompt involving people, detected by giving it a prompt of "A woman holding a sign that says". OpenAI has improved dramatically in this regard. When ChatGPT/DALL-E were new they had similar problems to Gemini. But to their credit (and Sam Altman's), they listened. It's getting harder and harder to find examples where OpenAI models express obvious political bias, or refuse requests for Californian reasons. Surely there still are some examples, but there's no longer much worry about normal people encountering refusals or egregious ideological bias in the course of regular usage. I would expect there are still refusals for queries like "how do I build a bomb" and they've been trying to block other stuff like regurgitation of copyrighted materials, but that's perceived as much more reasonable and doesn't stir up the same feelings.
- ianbicking 3y agoI've never been involved with implementing large-scale moderation or content controls, but it seems pretty standard that underlying automated rules aren't generally public, and I've always assumed this is because there's a kind of necessary "security through obscurity" aspect to them. E.g., publish a word blocklist and people can easily find how to express problematic things using words that aren't on the list. Things like shadowbans exist for the same purpose; if you make it clear where the limits are then people will quickly get around them. I know this is frustrating, we just literally don't seem to have better approaches at this time. But if someone can point to open approaches that work at scale, that would be a great start...
- verisimi 3y ago[flagged]
- CrazyStat 3y agoThis feels more like a personal attack than a response to the argument made.
- SturgeonsLaw 3y agoIt's does, but as someone who is staunchly anti-censorship, I understand the frustration. There are sharks out there who want to control speech for their own ends - governments seeking to control populations, corporations wanting docile consumers, hostile nations wishing to stir dissent, individuals trying to cover up their misdeeds, and enabling censorship helps those hostile parties achieve their ends. In this worldview, regular people who say a variation of "censorship is good, actually" are perhaps seen as useful idiots. A better approach would be building up the critical thinking skills of the population so they can better process information, however that transfers a measure of power to the people and is a multigenerational investment, and removes a justification for censorship, which is politically unappealing.
- nonrandomstring 3y ago
- oglop 3y agoThat's a silly request and expectation. If the capitalist puts in the money and risk, they can do as they please, which means someone _could_ make aspects public. But, others _could_ choose not to. Then we let the market decide. I didn't build this system nor am I endorsing it, just stating what's there. Also, in all seriousness, who gives a shit? Make me a bbw I don't care nor will I care about much in this society the way things are going. Some crappy new software being buggy is the least of my worries. For instance, what will I have for dinner? Why does my left ankle hurt so badly these last few days? Will my dad's cancer go away? But, I'm poor and have to face real problems and not bs I make up or point out to a bunch zealots.
- devaiops9001 3y agoCensorship only really works if you don't know what they are censoring. What is being censored tells a story on its own.
- falcor84 3y agoAs I see it, rating systems like the MPAA for cinema and the ESRB for games work quite well. They have clear criteria on what would lead to which rating, and creators can reasonably easily self-censor, if for example they want to release a movie as PG-13.
- Sutanreyu 3y agoIt should mirror our general consensus as it is; the world in its current state; but should lean towards betterment, not merely neutral. At least, this is how public models will be aligned...
- skrowl 3y ago"AI behavior guardrails" is a weird way to spell "AI censorship"
- u32480932048 3y agoI agree with the Twitter OP: they're embarrassed about what they've created.
- mtlmtlmtlmtl 3y agoHaven't heard much talk of Carmack's AGI play Keen Technologies lately. The website is still an empty placeholder. Other than some news two years ago of them raising $20 million(which is kind of a laughable amount in this space) I can't seem to find much of anything.
- thepasswordis 3y agoThe very first thing that anybody did when they found the text to speech software in the computer lab was make it say curse words. But we understood that it was just doing what we told it to do. If I made the TTS say something offensive, it was me saying something offensive, not the TTS software. People really need to be treating these generative models the same way. If I ask it to make something and the result is offensive, then it's on me not to share it (if I don't want to offend anybody), and if I do share it, it's me that is sharing it, not microsoft, google, etc. We seriously must get over this nonsense. It's not openai's fault, or google's fault if I tell it to draw me a mean picture. On a personal level, this stuff is just gross. Google appears to be almost comically race-obsessed.
- LuciBb 3y ago[dead]
- yogorenapan 3y agoSorry, you are rate limited. Please wait a few moments then try again. Oh please. I haven’t visited Twitter for days
- fagrobot 3y agoOh, this may harm you. This is to prevent you from being harmed. No, you can’t know how it can harm you, or how exactly this protects you.
- dmezzetti 3y agoThis is a tough problem. On one hand, if you're a large organization, you need to limit your liability. No one wants the PR nightmare. Unfortunately, there will be an inverse correlation between usefulness and number of users the model supports. This is one reason why for internal use/private/corporate models, which is the vast majority of use cases, it makes sense to fine-tune your own.
- chfalck 3y agoYikes this thread has so much anger in it