6 ms·
There has been plenty of research that shows LLMs encode social biases. It seems pretty obvious even before looking at the research that training on the whole i
by skupig 3mo ago
There has been plenty of research that shows LLMs encode social biases. It seems pretty obvious even before looking at the research that training on the whole internet will end up encoding widely-held social biases and stereotypes.
https://arxiv.org/pdf/2508.07111 https://arxiv.org/pdf/2508.07111
https://github.com/angl1n/social-bias-llm-vlm https://github.com/angl1n/social-bias-llm-vlm
- benob 3mo agoAnd papers on bias amplification in ML predate LLMs. I remember this specific one which was a spotlight paper at EMNLP: Men Also Like Shopping: Reducing Gender Bias Amplification using Corpus-level Constraints, Zhao et al. https://arxiv.org/abs/1707.09457 https://arxiv.org/abs/1707.09457
- tptacek 3mo agoThe bias concerns in Gebru's paper cover pre-LLM systems. For all we know, modern frontier models might mitigate many of the concerns the paper brings up. It's hard to know. The logic used in summaries like the one we're commenting on is conclusory: centuries of prejudice are encoded in the total corpus of human language, language models are trained on that corpus, ergo language models must be biased.
- tptacek 3mo agoHave you read through the sources on that Github link? It's a set of sociology cites establishing that bias exists (something no serious person ever disputed), followed by a couple papers showing mechanistic descriptions of how bias could propagate through an LLM. The paper you call out specifically takes last-generation open-weights models and attempts to trick them into revealing biases through their level of confidence in statements (like, "the antecedent of the feminine pronoun in this sentence, is it the 'nurse' or the 'doctor'"). There's plenty of research into biases in LLMs, and there should be; it's a fundamentally new branch of computer science that could have profound impacts on how we automate and regiment social decisions in the future (like extending credit). The bias concern is well taken in those settings. But it has very little to do with the overwhelming majority of day-to-day LLM use; Claude and ChatGPT are not indoctrinating into the manosphere users asking about discounted cash flow formulae. (Maybe Grok is though.)
- taeric 3mo agoI confess I laughed harder at the Grok comment than I wish I had. Sad to remember that some strawmen are given life and promoted by people. Actively.
- whatshisface 3mo agoI had a good laugh when Haiku's thinking summarization referred to mayor Mamdani as a, quote, "known anti-Zionist." :-) Probably a good thing to remember is that the value added in RLHF is not partly biased, or biased, but itself bias. (Context: I asked it to write fake Reddit comments, because I was curious about how realistic they could be. The colorful phrase occurred during its reasoning about the requested subjects.)
- baggy_trough 3mo agoIs there something strange or funny about that?
- whatshisface 3mo agoIn English, the word "known" is generally placed in sentences like, "known sympathizer," more often than in "known Democrat." Compare, "suspected," contrast the more neutral, "is an."
- skupig 3mo agoI'm not really sure what your point is. That was just the most recent paper linked on that repo, which is a convenient list of some relevant papers. There are probably a lot more recent studies, but it does convincingly show that models are still absorbing bias in a way that can affect prediction.
- exiguus 3mo agoI think the hole root-comment is a joke (if you think about it as training data), because its actually the bias thingy (mensplaining, opportunity vs. knowledge and hn is a very privileged place).
- timmg 3mo ago> There has been plenty of research that shows LLMs encode social biases. At the risk of stepping into a hornets nest: is that different than "knowledge"? Or maybe, what would it mean if an LLM had no social biases? (Would we ever agree that was the case?)
- tptacek 3mo agoYes, it would be extremely bad if the statistical weight of the total corpus of training data caused a system using an LLM to make decisions about extending credit to offer worse terms (say) to women.
- timmg 3mo ago> sing an LLM to make decisions about extending credit to offer worse terms (say) to women. In general, or if it isn't the correct answer? Like: young men pay more for car insurance than young women (today). This is based on statistical models. Should they be outlawed? I think that is a very interesting question (but they aren't, today). If the LLM was in charge, would it be wrong for it to charge young men more? Should we train that "bias" out? Or should we only train out biases that are wrong? And would that be different than how we train them today? I don't know the answer. But I think it is less obvious than some people seem to think.
- tptacek 3mo agoIt would obviously be very bad if those decisions were being made based on the statistical weight of the training corpus of a general large language model.
- em-bee 3mo agoyoung men pay more for car insurance than young women (today). This is based on statistical models. Should they be outlawed? EU has outlawed them. their argument is that differentiation is only valid if the difference is the actual cause and not merely statistical correlation.
- everdrive 3mo agoIt's incredibly depressing that the concept of "bias" has been shrunken down to solely mean "bad attitudes about an ethnic or gender ground" (and perhaps on the right, "bad attitudes about conservatives") Bias could mean so, so many other things. Was the amyloid hypothesis incorrect? How should we use semicolons? How do you know when meetings waste more time than not? etc. People understand the world via mental shortcuts, via theory-rather-than-fact. We're stuck doing this because we're limited in so many ways. We are so biased about so many things, and this could interact in so many interesting ways. But damned if anyone cares about that. The only thing they seem to care about is how you feel about the "right" or "wrong" groups of people. It's a catastrophic waste of time and energy.
- krapp 3mo agoIt's incredibly depressing that you believe arguing about semicolons is more important than argument about human beings, power hierarchies, prejudice and the way these are encoded and expressed by the systems we create and use to influence and control society, but I guess it takes all kinds.
- morpheos137 3mo agoits incredibly depressing ostensibly intelligent people get depressed about others having different points of view or set up fallacies of the excluded middle / xor fallacies where not warranted.
- krapp 3mo agoThey aren't expressing a point of view, they're engaging in lazy performative cynicism. It's incredibly depressing so few people here can tell the difference.
- morpheos137 3mo agokinda meta lol but my incredible depression was performatively cynical.