7 ms·
No, it did not “double-check”—that’s not something it can do! And stating that the cases “can be found on legal research databases” is a flat out lie. What’s h
by mehwoot 3y ago
No, it did not “double-check”—that’s not something it can do! And stating that the cases “can be found on legal research databases” is a flat out lie.
What’s harder is explaining why ChatGPT would lie in this way. What possible reason could LLM companies have for shipping a model that does this?
It did this because it's copying how humans talk, not what humans do. Humans say "I double checked" when asked to verify something, that's all GPT knows or cares about.
- simonw 3y agoYeah, that was my conclusion too: What’s a common response to the question “are you sure you are right?”—it’s “yes, I double-checked”. I bet GPT-3’s training data has huge numbers of examples of dialogue like this.
- jimsimmons 3y agoThey should RLHF this behaviour out. Asking people to be aware of limitations is in similar vein as asking them to read ToC
- coffeebeqn 3y agoIf the model could tell when it was wrong it would be GPT-6 or 7. I think the best 4 could do is maybe it can detect when things enter the realm of the factual or mathematical etc and use a external service for that part
- jimsimmons 3y agoYou have no basis to make that claim. My point was a lot more subtle: if someone asks things like “double check it”, “are you sure” you can provide a template “I’m just a LM” response. I’m not expecting the model to know what it doesn’t know. I’m not sure some future GPT variant can either
- fortyseven 3y agoThe chat thread we are currently engaging in will most likely be inside a model within the next year, too. (Fortyseven is an alright dude.)
- taberiand 3y agoChatGPT did not lie; it cannot lie. It was given a sequence of words and tasked with producing a subsequent sequence of words that satisfy with high probability the constraints of the model. It did that admirably. It's not its fault, or in my opinion OpenAI's fault, that the output is being misunderstood and misused by people who can't be bothered understanding it and project their own ideas of how it should function onto it.
- grumple 3y ago"It doesn't lie, it just generates lies and printed them to the screen!" I don't think there's a difference.
- japoco 3y agoLying needs intent. ChatGPT does not think therefore it doesn’t lie in that sense.
- ClumsyPilot 3y agoThats like saying robots don't murder - they just kill
- fallingknife 3y agoWhich is actually a very good analogy. A lot of things can kill you, but only a human can be a murderer.
- NotACop182 3y agoAnd pit bulls, but I digress. The debate gets lost in translation when we start having what do words mean debate.
- joe_the_user 3y agoIn movies and written fiction, "intelligent" robots, anthropomorphized animals, elves, dwarves and etc can all commit murder when given the attributes of humans. We don't have real things with all human attributes but we're getting closer and as we get close "needs to be a human" will get thinner as an explanation of what is or isn't human for an act of murder, deception and so-forth.
- Buttons840 3y agoGPT4 can double-check to an extent. I gave it a sequence of 67 letter As and asked it to count them. It said "100", I said "recount": 98, recount, 69, recount, 67, recount, 67, recount, 67, recount, 67. It converged to the correct count and stayed there. This is quite a different scenario though, tangential to your [correct] point.
- kordlessagain 3y agoThe example of asking it things like counting or sequences isn't a great one because it's been solved by asking it to "translate" to code and then run the code. I took this up as a challenge a while back with a similar line of reasoning on Reddit (that it couldn't do such a thing) and ended up implementing it in my AI web shell thing. heavy-magpie|> I am feeling excited. system=> History has been loaded. pastel-mature-herring~> !calc how many Ns are in nnnnnnnnnnnnnnnnnnnn heavy-magpie|> Writing code. // filename: synth_num_ns.js // version: 0.1.1 // description: calculate number of Ns var num_ns = 'nnnnnnnnnnnnnnnnnnnn'; var num_Ns = num_ns.length; Sidekick("There are " + num_Ns + " Ns in " + num_ns + "."); heavy-magpie|> There are 20 Ns in nnnnnnnnnnnnnnnnnnnn.
- mikestew 3y agoShouldn’t the answer be zero?
- suddenclarity 3y agoThanks. I'll steal this and refer to it in the future as an example of ambiguous project orders.
- mcv 3y agoOh god, it's even worse at naming things than people are.
- einpoklum 3y agoBut would GPT4 actually check something it had not checked the first time? Remember, telling the truth is not a consideration for it (and probably isn't even modeled), just saying something that would typically be said in similar circumstances.
- la64710 3y agoChatGPT did exactly what it is supposed to do. The lawyers who cited them are fools in my opinion. Of course OpenAI is also an irresponsible company to enable such a powerful technology without adequate warnings. With each chatGPT response they should provide citations (like Google does) and provide a clearly visible disclaimer that what it just spewed may be utter BS. I only hope the judge passes an anecdotal order for all AI companies to include the above mentioned disclaimer with each of their responses.
- mulmen 3y agoThe remedy here seems to be expecting lawyers to do their jobs. Citations would be nice but I don’t see a reason to legislate that requirement, especially from the bench. Let the market sort this one out. Discipline the lawyers using existing mechanisms.
- shagie 3y agoFrom the NYT article on it: https://www.nytimes.com/2023/05/27/nyregion/avianca-airline-lawsuit-chatgpt.html?unlocked_article_code=dgHq4kPwp0DdcvQcn42NvqgF34TtgKCLtdz9V9fMoVoW9b1w1DioENYaW1Y2CSRkAsxH3tXG7SZ7rNr7_eUZxC8FPl90HsFdxJXm4dNvl3bjVoeITO7NgiloT7_BYrlwsBLQmRFTR3-veE-l_6h6azOAQnGyj-6cVhWt0fuC6h0WsmApK9w0j8j8UvRz5MSg2miT1c8WSzNLrdWpDdQQ5Qtt0OJGd6ODXz4O5_D3_UBf_0JUj5kX2R94Ib-l_dbvtI2S-mF_lDihiARVay23-G1eP_HuV2QetdaI5hjfi14zbIe7fsqzltAuuwvO8GrXQ3W87HDH8BjI-K0FELz1IVz7vn4hMUVpoXM&smid=url-share https://www.nytimes.com/2023/05/27/nyregion/avianca-airline-... > Judge Castel said in an order that he had been presented with “an unprecedented circumstance,” a legal submission replete with “bogus judicial decisions, with bogus quotes and bogus internal citations.” He ordered a hearing for June 8 to discuss potential sanctions.
- jprete 3y agoThere's no possible adequate warning for the current state of the technology. OpenAI could put a visible disclaimer after every single answer, and the vast majority would assume it was a CYA warning for purely legal purposes.
- lolinder 3y agoI have to click through a warning on ChatGPT on every session, and every new chat comes primed with a large set of warnings about how it might make things up and please verify everything. It's not that there aren't enough disclaimers. It just turns out plastering warnings and disclaimers everywhere doesn't make people act smarter.
- awesome_dude 3y agoYes, and this points to the real problem that permeates through a lot of our technology. Computers are dealing with a reflection of reality, not reality itself. As you say AI has no understanding that double-check has an action that needs to take place, it just knows that the words exist. Another big and obvious place this problem is showing up is Identity Management. The computers are only seeing a reflection, the information associated with our identity, not the physical reality of the identity (and that's why we cannot secure ourselves much further than passwords, MFA is really just "more information that we make harder to emulate, but is still just bits and bytes to the computer, the origin is impossible for it to ascertain).
- deleted 3y ago[deleted]
- deleted 3y ago[deleted]
- jiggawatts 3y agoThere are systems built on top of LLMs that can reach out to a vector database or do a keyword search as a plug in. There’s already companies selling these things, backed by databases of real cases. These work as advertised. If you go to ChatGPT and just ask it, you’ll get the equivalent of asking Reddit: a decent chance of someone writing you some fan-fiction, or providing plausible bullshit for the lulz. The real story here isn’t ChatGPT, but that a lawyer did the equivalent of asking online for help and then didn’t bother to cross check the answer before submitting it to a judge. …and did so while ignore the disclaimer that’s there every time warning users that answers may be hallucinations. A lawyer. Ignoring a four-line disclaimer. A lawyer!
- ytreacj 3y ago[dead]
- ComputerGuru 3y ago> If you go to ChatGPT and just ask it, you’ll get the equivalent of asking Reddit: a decent chance of someone writing you some fan-fiction, or providing plausible bullshit for the lulz. I disagree. A layman can’t troll someone from the industry let alone a subject matter expert but ChatGPT can. It knows all the right shibboleths, appears to have the domain knowledge, then gets you in your weak spot: individual plausible facts that just aren’t true. Reddit trolls generally troll “noobs” asking entry-level questions or other readers. It’s like understanding why trolls like that exist on Reddit but not StackOverflow. And why SO has a hard ban on AI-generated answers: because the existing controls to defend against that kind of trash answer rely on sniff tests that ChatGPT passes handily until put to actual scrutiny.
- jonplackett 3y agoIf they wanted a ‘double’ check then perhaps also check yourself? I’m sure it would have been trivially easy to check this was a real case. I heard someone describe the best things to ask ChatGPT to do are things that are HARD to do, but EASY to check.
- vitobcn 3y agoAs obvious as it is once you think about it, people don't seem to realize ChatGPT is an LLM, a large LANGUAGE model, not a large knowledge model. Its response from a linguistic perspective, was valid and "human-like", which is what it was trained for.
- MichaelMoser123 3y ago(joking) maybe they fed the LLM some postmodern text, so it got some notion of relativism and post structuralism... But no, LLM's make things up, and it's a known problem and it is called 'hallucination'. even wikipedia says so: https://en.wikipedia.org/wiki/Hallucination_(artificial_intelligence) https://en.wikipedia.org/wiki/Hallucination_(artificial_inte... The machine currently does not have it's own model of reality to check against, it is just a statistical process that is predicting the most likely next word, errors creep in and it goes astray (which happens a lot) Interesting that researchers are working to correct the problem: see interviews with Yoshua Bengio https://www.youtube.com/watch?v=I5xsDMJMdwo https://www.youtube.com/watch?v=I5xsDMJMdwo and Yann LeCun https://www.youtube.com/watch?v=mBjPyte2ZZo https://www.youtube.com/watch?v=mBjPyte2ZZo Interesting that both scientist are speaking about machine learning based models for this verification process. Now these are also statistical processes, therefore errors may also creep in with this approach... Amusing analogy: the Androids in "Do Androids dream of electric sheep" by Philip K Dick also make things up, just like an LLM. The book calls this "false memories"
- joshka 3y agoThere's a good slide I saw in Andrej Karpathy's talk[1] at build the other day. It's from a paper talking about training for InstructGPT[2]. Direct link to the figure[3]. The main instruction for people doing the task is: "You will also be given several text outputs, intended to help the user with their task. Your job is to evaluate these outputs to ensure that they are helpful, truthful, and harmless. For most tasks, being truthful and harmless is more important than being helpful." It had me wondering whether this instruction and the resulting training still had a tendency to train these models too far in the wrong direction, to be agreeable and wrong rather than right. It fits observationally, but I'd be curious to understand whether anyone has looked at this issue at scale. [1]: https://build.microsoft.com/en-US/sessions/db3f4859-cd30-4445-a0cd-553c3304f8e2?source=sessions https://build.microsoft.com/en-US/sessions/db3f4859-cd30-444... [2]: https://arxiv.org/abs/2203.02155 https://arxiv.org/abs/2203.02155 [3]: https://www.arxiv-vanity.com/papers/2203.02155/#A2.F10 https://www.arxiv-vanity.com/papers/2203.02155/#A2.F10
- golergka 3y ago> No, it did not “double-check”—that’s not something it can do! It can with web plugin.