5 ms·
Your one example doesn't make all of LLMs a lie. It's like someone reading a National Enquirer article about "Hillary Clinton being an alien from outer space"
by thephyber 14d ago
Your one example doesn't make all of LLMs a lie.
It's like someone reading a National Enquirer article about "Hillary Clinton being an alien from outer space" (a real headline topic from decades ago) and drawling the conclusion that all journalism is "a lie". The user has to understand media literacy and be at least a little skeptical of the claims that are made, then cross reference with another source.
- evilfred 13d agoby what mechanism do you believe LLMs verify truth? they are amoral token generators
- fluoridation 14d agoOkay, but the anecdote states that every model repeated the pseudo-factoid about Foobar square, not just the 4 GB open source model equivalent of a tabloid.
- CamperBob2 14d agoWithout disclosing what you were prompting for, it's impossible to evaluate your claim.
- fluoridation 14d agoCorrect. We can either accept the claim or disregard it. The comment I replied to opted to accept it and then committed a fallacy, hence my response.
- nulbyte 13d agoI think the key phrase here is, "an obscure small town." There may only be a single mention of this place, hence the only one on which a response can be based. This says more about the user's understanding of LLMs than it does about LLMs.
- CamperBob2 13d agoThis says more about the user's understanding of LLMs than it does about LLMs. "Tell me everything you know about (obscure small town), (state). Only what's unique to (town), not commonly-known facts" is an excellent way to test for hallucinatory tendencies in a new model, in my experience. Likely the best I've found. Quality of results is almost linearly proportional to the size of the model in many cases. The largest models like K3 and GLM 5.3 will either confine their responses to known true facts about the town and its surroundings, or admit they don't have enough information to answer. Smaller ones will reliably make up hilarious or downright-strange things. Another good test is https://whatever.scalzi.com/2025/12/13/ai-a-dedicated-fact-failing-machine-or-yet-another-reason-not-to-trust-it-for-anything/ https://whatever.scalzi.com/2025/12/13/ai-a-dedicated-fact-f... , which still works on the newest models. Of the open-weight models available, only Kimi K3 will consistently admit it has no idea who Scalzi's novel is dedicated to. The rest still make up random stuff and present it confidently. TL,DR: progress is possible, and it has been made, but it's happening slower than many people think.
- mstaoru 13d agoReminds me of https://petergpt.github.io/bullshit-benchmark/viewer/index.v2.html https://petergpt.github.io/bullshit-benchmark/viewer/index.v...
- mstaoru 14d agoIt kind of does, mathematically. It doesn't label confidence. If 0.0001% of answers is a lie, without knowing which parts are a lie exactly, you cannot trust any of them. If you need to independently verify every fact, why not just gather facts yourself in the first place. Let's say, a mathematical concept of lie. I still use them every day, of course.
- alansaber 13d ago"why not just gather facts yourself" vs "i use them every day", the duality of man. But yes, I appreciate that LLM output is theoretically completely untrustworthy- but when in practice I observe that it's around 90% accurate, I have to rely on my own internal calibration for how useful it is (depends on type, nature of task ofc)
- mstaoru 13d agoAs they say, the less you know, the better LLM output is. :D
- vetler 13d ago> Your one example doesn't make all of LLMs a lie. It's not a lie though, because the truth isn't guaranteed by the mechanism that generates the answer.
- deleted 13d ago[deleted]