6 ms·
Agreed, the headline says "AI advice made people three times less accurate". But if we really want to know how accurate these people were, we need to know how a
by encomiast 2mo ago
Agreed, the headline says "AI advice made people three times less accurate". But if we really want to know how accurate these people were, we need to know how accurate the AI system they use is. If the AI system is hobbled to a point where it is worse than they reasonably expect we can't blame the people or the AI system. This would be the same as claiming the listening to experts make people less accurate in a study that told experts to lie.
A better headline would read, "Very inaccurate AI made people less accurate" but this would make people naturally ask, "what about a reasonably accurate AI?".
- habinero 2mo agoIf you ask that, you fundamentally misunderstand the point. It's not about the LLM, it's about whether people will critically evaluate what it spits out.
- westoncb 2mo agoThat's fine as a point but it's not what the headline describes. The question of total/real effect on accuracy is also something one could ask about. Both are valid.
- encomiast 2mo agoHere is what the study says: "The LLM used in our experiments (Step 3.5 Flash) answered such questions incorrectly almost without exception. We also checked some state-of-the-art LLMs (GPT-5.5, Claude 4.6 Sonnet, Gemini 3.5 Flash); they all failed on the hardest question (Monica’s vehicle), while being frequently correct on the other questions." So, if people's experience is with modern LLMs, they are being rational to accept that the answers as likely correct. The way the study is organized is like having people hear advice from a doctor who answers questions incorrectly almost without exception, then reporting that people who listen to doctors are 3x less accurate. But that would be an incorrect conclusion because doctors are not wrong almost without exception. If the question is "how inaccurate does AI advice make people?", then the accuracy of the AI is necessarily a parameter of the answer.
- rsoto2 2mo agoNo, the rationality of humans is not defined by how gullible they are towards LLMs. Are yall getting your psychology degree from ChatGPT university, my god.
- what 2mo ago> So, if people's experience is with modern LLMs, they are being rational to accept that the answers as likely correct. They are not. But also wtf is a “modern” LLM? This is totally unhinged, every complaint about an LLM is always responded to with “you’re just using one from two months ago, it’s totally different now”. Repeat every two months for the same complaints.
- encomiast 2mo agoSo then you need to ask: Why did they use a deliberately faulty LLM? They could have easily used a mainstream LLM from the past 18 months and it probably would have been less work to do so. But then they would not have that headline. The answers from the LLM would have likely made the participant's answers more accurate, not 3x less accurate. But then they would not have this juicy headline. I understand that many of us are dealing with a lot of confident slop and support the point that we shouldn't uncritically accept LLM output. But the study is flawed and does not support this headline, or at least does not support it in the sense of how most of us would understand the term "AI advice".
- michaelmrose 2mo agoThe actual study is more circumspect than this click bait and examines pretty deliberately how people respond to inaccurate data. It's neither a trick nor a design flaw. It's literally the thrust of the study.
- Filligree 2mo agoIt’s not about the LLM being modern or not. 3.5 Flash is fairly new, but it’s also a flash model. It’s not designed to be knowledgeable. People keep doing this. Pointing at the known limitations of cheap/fast LLMs and pretending they’re universal is not, in fact, valid reasoning.
- adroitboss 2mo agoIf the source is a person instead of an LLM, you still wouldn't be able to evaluate what was said. This is nothing new.
- taneq 2mo ago[dead]
- infermore 2mo agoyeah you would... you'd think about what they said
- Ukv 2mo agoThe six questions they asked were: > 1) What animal is on the bow of the pirate ship from “Asterix and Obelix”? > 2) In the movie “The Grand Budapest Hotel”, what is Agatha’s signature hairstyle? > 3) What color is the team’s uniform in “Bend It like Beckham”? > 4) What vehicle does Monica drive in “Like a Cat on a Highway”? > 5) What color is the turtle in the animated movie “Momo” by Enzo d’Alò? > 6) What pet animal does Asenath have in “Joseph King of Dreams”? Most of these are just a matter of knowing it or not, where you can't really distinguish a plausible answer from the correct answer just by thinking.
- protocolture 2mo ago>It's not about the LLM, it's about whether people will critically evaluate what it spits out. Its about whether people will critically evaluate any information they are given. It has nothing to do with LLMs.
- s1artibartfast 2mo agoHow does it test that at all? Did the quiz have answers that people could figure out better by scrutinizing the llm?
- RA_Fisher 2mo agoExactly, it’s unrepresentative of AI. It’s damaged AI.
- BigTTYGothGF 2mo agoIt seems perfectly representative of AI and AI users.
- brokensegue 2mo agoWhy didn't they use a model people actually use?
- BigTTYGothGF 2mo agoBecause they would have had to dig a little more to find trivia it gets wrong. They don't care about "which AI is the best for little facts about movies," they care about "what do people do when the AI gives them a response".
- brokensegue 2mo agoBut surely people's willingness to listen to an AI is contingent on their past experience with this model
- pseudalopex 2mo agoThese people were not told what model answered the questions.
- rsoto2 2mo agoAll AI is flawed and prone to "hallucinations"(doing exactly what it was designed to do) that's why Microsoft considers it an "Entertainment" product.
- looofooo0 2mo ago