9 ms·
I feel like I'm going nuts. There are other commenters saying this is a good practice they've also done for other injuries. You are saying you are an actual ra
by rafterydj 3mo ago
I feel like I'm going nuts.
There are other commenters saying this is a good practice they've also done for other injuries. You are saying you are an actual radiologist and immediately clock the problems with its advice.
I have seen this pattern over and over again. Anytime someone is an actual expert at anything, AI output appears insufficient or incomplete or outright misleading. It is only when you do not know what the AI is being asked to do is it likely you will find the output helpful.
This is itself alarming to me, but no one else seems to find this to be quite damning for the AI services being offered, preferring instanced to be wowed by the convenience and speed at which they can be delivered unreviewed and unproven information.
- redsocksfan45 3mo ago[dead]
- highfrequency 3mo agoSeems natural enough. There will always be complexity and nuance that is missed by an AI model or person - the world is just super detailed. The more expertise you have the more you will be aware of that nuance. That doesn't mean the model or person is not useful as a starting point.
- sbarre 3mo ago> Anytime someone is an actual expert at anything, AI output appears insufficient or incomplete or outright misleading Yes, this is exactly so. AI is able to confidently sound plausible enough to convince laypersons or anyone who isn't very familiar with the subject matter, which is a big part of the mass-appeal "magic" of ChatGPT and other similar tools. It's like having a know-it-all friend (who also makes shit up to bridge their own knowledge gaps). In many non-advanced non-specialized situations, AI is right enough to be at best useful or at worst not harmful (usually landing in the middle somewhere). But speaking for myself, in areas where I consider myself quite proficient, I can very easily spot the subtle inconsistencies and naive conclusions that AI responses provide, and I have to guide/steer/correct it a lot to get good results when the subject matter is complex enough.
- nlawalker 3mo ago> no one else seems to find this to be quite damning for the AI services being offered, preferring instanced to be wowed by the convenience and speed at which they can be delivered unreviewed and unproven information "Be wowed by the convenience and speed", or merely "take advantage of the mere availability"? What most people find to be damning about expert advice is that they simply can't get it anywhere, at any cost that they can afford.
- whatever1 3mo agoSo if you want to do a surgery but you don’t see any surgeons around you ask a grocery butcher to have his way?
- sxg 3mo agoIn certain circumstances, the answer is yes. If an airplane's pilots are incapacitated, do you simply give up and crash the plane because there are no other pilots on board? Or would you rather have someone on the ground try to coach a passenger into at least attempting to land the plane?
- ChrisMarshallNY 3mo agoAs long as that passenger didn’t have the fish.
- acheron 3mo agoYes, I remember, I had lasagna.
- frereubu 3mo agoThat's an extreme edge case, which I don't think is in the context of the concerns in this thread.
- sxg 3mo agoThe specific case doesn't matter--it's meant to make you think about the general question throughout this thread: when an expert isn't available, should non-experts use AI (or other tools) to help themselves? Sometimes the answer is yes because the potential benefits outweigh the potential harms (if any harms exist). But sometimes the answer is no because misleading/incorrect advice can cause a net harm.
- newsclues 3mo agoLLM is not necessarily an expert system. Once there are expert systems for law, healthcare, accounting, governance… https://en.wikipedia.org/wiki/Expert_system https://en.wikipedia.org/wiki/Expert_system
- microgpt 3mo agoDidn't they try that in the 80s and 90s but discover the real world is too variable for that to work?
- parineum 3mo ago> I have seen this pattern over and over again. Anytime someone is an actual expert at anything, AI output appears insufficient or incomplete or outright misleading. AI isn't even the first instance of this phenomenon, news articles are like this as well. https://en.wiktionary.org/wiki/Gell-Mann_Amnesia_effect https://en.wiktionary.org/wiki/Gell-Mann_Amnesia_effect
- appplication 3mo agoThis is the root of AI psychosis. There’s a lot of unpack here, and I won’t go too deep because you can’t really have a discussion with affected folks because their fundamental basis is not evidence, it’s belief. It is weirdly religious in a way, because if you were to present contrary evidence (e.g. experts in a field weighing in about how plausible sounding responses are bunk), you would only be told you don’t believe enough in the long term potential and capabilities. Don’t get me wrong, I think we all agree capabilities will eventually improve (and farther-future capabilities could reasonably surpass experts), but really is unclear if the current transformer architectures with their probabilistic/hallucinatory outputs will plateau before they surpass current experts abilities in all promised fields.
- lazide 3mo agoI don’t think they will improve, there is too much incentive to poison the datasets going forward. A lot of the models up to this point have been benefitted - like Google did - from essentially ‘pre SEO’ internet. Now the same tools are being used to generate nigh infinite good sounding bullshit, which poisons the dataset in all sorts of hard to detect ways. To add insult to injury, the human experts are also not as. Naive, and have many incentives to poison their own input in subtle ways too.
- rvnx 3mo agoHuman doctors use LLMs to diagnose too OpenEvidence claims "More than 40% of U.S. physicians use it daily, and it handled around 20 million clinical consultations per month. Over 100 million Americans were treated by a doctor using it in 2025." https://www.cnbc.com/2026/01/21/openevidence-chatgpt-for-doctors-doubles-valuation-to-12-billion.html https://www.cnbc.com/2026/01/21/openevidence-chatgpt-for-doc...
- something98 3mo agoThis is a very misleading statement; most of those physicians are using LLMs to transcribe notes from visits and/or for billing purposes (e.g., proper billing codes).
- sxg 3mo agoI see your argument, but it's not exactly news that an expert found a flaw in a popular tool. You could say the same about Wikipedia--experts have tons of issues with it, but Wikipedia still provides value to non-experts. The most likely alternative to Wikipedia for non-experts is simply not trying to learn anything new. Similarly with LLMs, you can't just write them off entirely because they sometimes provide misleading or incorrect advice. The positive utility maximizing view is to learn when you need to call in an expert. I recently moved in to a new house and have used Claude extensively to figure out basic things (e.g., adjusting the garage door height, how to mount a TV). However, when the HVAC suddenly stopped working, I gave Claude a shot for an hour and tried some non-destructive fixes, but then realized I had to call in an HVAC expert.
- ohyes 3mo agoThe free alternative to Wikipedia is the library, not “don’t learn anything new ever”. I find Claude is surprisingly similar to a confident but incorrect coworker, with the benefit that Claude will reevaluate when I correct it.
- bflesch 3mo agoClaude will do everything to retain you as a user, because that's one of their most important metrics.
- ohyes 3mo agoExcellent point my colleague has the exact opposite incentive.
- sxg 3mo agoI used the phrase "most likely alternative" intentionally. The library is where people should go to get answers in a world without Wikipedia, but the vast majority of people won't. So in practice, most non-experts either learn from Wikipedia or don't try to learn anything at all.
- 3mo ago
- kryogen1c 3mo agoOn the flip side of this problem, novel best practices lag the medical standard of care, other human failures like corruption and competing priorities notwithstanding. For example, we had to advocate for certain practices during the birth of our first child that became routine during our second several years later. So, neither side is guaranteed correct, doctor or citizen researcher (which did not include LLMs in my case, for the record). The truest answer is also the most useless one, applicable to all fields: it depends. The real question is: if you embrace being a layman, whom do you trust more: LLMs/the internet or experts, like doctors? I think the answer is pretty clearly experts.
- beering 3mo agoTFA doesn’t actually state where the bit about shockwave therapy came from and it wasn’t the main point of the article. The concern was about being given useless therapies. The homeopathic analgesic is concerning, at least to me. I.e. nothing this radiologist said was related to the LLM’s advice.
- jstummbillig 3mo agoNo, not anytime someone is an actual expert at anything, AI output appears insufficient. That is why experts in various fields use AI. Then to say "Aha, but all of that is AI psychosis" makes obviously no sense: Why would we trust experts when they offer critique but not when they say "this is helpful"? Overall: People are not insane. AI makes mistakes and, often, fails completely. AI also helps them do things better, quicker, increasingly so. The jaggedness of AI is confusing and real.
- torben-friis 3mo agoHow many times have you seen an expert go "yeah these results are good consistently enough for a non expert to trust them without expert assistance"? There is a huge difference between having a chance of a good result, which can be useful for experts able to filter out the bullshit, and consistent success. I would generate code as a helper, I would never allow a guy from marketing to merge unreviewed AI code.
- hectdev 3mo agoThat's what I would like to call job security. When you know how to read what is wrong, you can easily catch the mistakes and correct it. AI gets you there faster by doing a lot of things right and you correct the mistakes.
- tpmoney 3mo agoI had a realization recently that the problem with "AI isn't consistently good enough" is that experience is probably not sufficiently distinguishable from the experience most non-experts have with computer systems all the time. As an industry we've been promising people for decades that if they put all their data into our special softwares they can get all sorts of information back out that will make life easier for them, reveal new insights and otherwise improve their understanding. But the unspoken caveat has always been that you have to put the right data into the right places, in the right format, in the right way and then you have to ask the right questions, in the right syntax, with the right tools. And if you get any one of those parts wrong, you're not going to get the right answers (or possibly even any answer at all). How many people have had their excel worksheet that they (or someone else they asked/employed) built for some task that has been working fine for the last year suddenly stop working or start throwing out nonsense numbers because some input changed? Or how many people have experienced their system seemingly throw out meaningless garbage because daylight savings changed right at the moment the report was being run? Or spent months operating on wrong data because the person who wrote the query misplaced a parenthesis and the query was searching for "(foo AND bar) OR baz" and not "foo AND (bar OR baz)". For most people, the computer and the programs they use to do their jobs are magical black boxes that most of the time produce mostly the right answers and sometimes get things very very wrong with no indication of what has changed. Which is effectively the same experience they will have with an AI, but now instead of needing to figure out some arcane excel pivot table and VBA script, they can just dump some raw data and a "natural language" question into the AI. And that's not counting the fact that their experience with looking information up online is about the same as well. How many absolutely confident wrong takes have you encountered online for things you're an expert in? How many of those wrong takes have come straight from supposedly trustworthy sources like news companies or even other people in the field? For most people, using a computer has always come with the asterisk that you should always be aware that the source you're reading could be very wrong, that the output is only correct assuming all the inputs and all the parts processing that input are also correct and that everything you do should be accompanied by vetting by experts, whether those experts were software developers or domain experts. For most people the only thing that's changed with AI is that it's a one stop shop for their "probably directionally right, almost certainly wrong in the details" access to the digital oracles.
- silisili 3mo agoThis is natural and even logically expected. It's just Gell-Mann amnesia in action. The world has more people spouting on things than it has people knowledgeable in said things. Apply that to the Internet at large, and realize where LLMs got their training. They're basically ConfidentlyIncorrect personified.
- tomaskafka 3mo agoYes. The PM’s “with AI I know enough to be dangerous, haha” means “I’m actually dangerous and I don’t realize”
- meindnoch 3mo agoWe're past the point of Gell-Mann amnesia. This is full blown Gell-Mann psychosis.
- suttontom 3mo agoYour instinct is correct, and in a lot of cases it's true. However, I've heard from enough doctors by now (a cardiologist, psychiatrist, and epidemiologist/former physician) that they use medical LLMs and find them extremely helpful, mostly as a way to either bring up knowledge they'd forgotten about or as a way to learn something new and then verify it. I'm extremely skeptical about LLMs in general and the connection to Gell-Mann Amnesia is apt, but I wouldn't necessarily write them off completely like that. There are experts using the models that find them genuinely helpful in their field.
- GTP 3mo agoProbably this is the point, and it's a point that has been brought up a lot of times in the past, maybe less in recent times: you need to know the things you're applying an LLM to. In this way, you can keep the good outputs while having the expertise to discard the bad ones.
- meowface 3mo agoI may be missing something, but I think it's unclear that the parent poster here is necessarily actually contradicting anything the AI said. It may depend on the exact information the OP wrote to Claude and GPT. The full transcripts would be needed. (Though there is definitely a separate point that a doctor would generally better know all the right questions to ask, while current LLMs may be making certain assumptions.) The LLM may have, from its "perspective", implicitly thought the OP was telling it that he had strong reason to believe there was no calcification and was not considering the bigger picture of possibly receiving an incomplete/poor assessment from the medical staff. In fact, the issue here may be the LLM overly trusting doctors vs. trusting its own expertise.
- je42 3mo agoThe question is how far is AI off compared to the professional that we have access to. World best experts are not accessible to most of us. :(
- qnleigh 3mo agoTotally agree. I'm a scientist, and like most scientists I have some specialized skills that most of my colleages don't. AI has empowered them to learn and build things that they might have otherwise needed me for. But there have been quite a few cases where it led them very far down a wrong path. This has started happening way more often in the last few months.* We've known since the beginning that AIs confidently say incorrect things. But now that they can speak confidently about very complex topics, and mostly say correct things, we are letting our guard down and lots of subtle falsehoods are slipping through. *In one case, I was able to put things back on track because the AI suggested my colleague talk to me; somehow it figured out we were co-workers.
- bitlad 3mo ago>very far down the wrong path. Absolutely agree. Have seen this first hand
- aspenmartin 3mo agoRight but hallucination rates have been consistently decreasing every model iteration. It's about error rates. As also a fellow scientist, I also will mess something up. Humans have an error rate. Once that error rate is low enough, it doesn't matter that it's > 0, it matters that it's low enough to be trustworthy and useful. Coding agents of 2024-25 had error rates too large; you couldn't meaningfully vibe code anything and needed a ton of oversight. It's still true but FAR less so, and this is after like a year of iteration.
- grayhatter 3mo ago> This is itself alarming to me, but no one else seems to find this to be quite damning for the AI services being offered, preferring instanced to be wowed by the convenience and speed at which they can be delivered unreviewed and unproven information. Welcome to the club? This new awareness you've found over the true quality of LLM based GenAI output has been what "all the haters" have been mad about for-ever. That the output of LLMs are clearly defective, and merely have found a cute trick towards making humans think they're less defective than they are actually measured to be. And the corresponding anger and frustration to push the risks of genai output out onto others, while also aggressively pushing it as a feature you should be using already. You're behind don't you know, and whatever other lie I have to tell to trick you into enough FOMO to pay me 200USD/mo so I can sell FOSS back to you. An LLM can only output the mean next likely token, and then add a bunch of extra noise on top of that so it feels interesting and not repetitive. None of this is new, the problem is, 50% of humans are below the mean, but have no idea. So when an LLM tells them some lie: well, it sounds so helpful! It's impossible for someone who sounds this helpful to lie to me, liars never sound confident! It must be PERFECT! I'm gonna tell everyone how perfect it is. so the bottom 0-33% think LLMs are fantastic tools that make nearly 0 mistakes in comparison to the bottom 33%. 33-66%-ish aren't sure, some times it's great, but it will make that random mistake sometimes, but I can catch most (or all of them depending on ego). and the 66%+ are angry about how many people are getting tricked by something so obviously low quality, or are lucky enough to not have to care.
- orangecat 3mo agoAn LLM can only output the mean next likely token, and then add a bunch of extra noise on top of that so it feels interesting and not repetitive. So when an LLM was asked to analyze the unit distance conjecture, it just spat out a bunch of average-or-random tokens that coincidentally happened to correspond to a valid proof that had eluded humans for decades?
- grayhatter 3mo ago> So when an LLM was asked to analyze the unit distance conjecture, it just spat out a bunch of average-or-random tokens that coincidentally happened to correspond to a valid proof that had eluded humans for decades? yes https://en.wikipedia.org/wiki/Texas_sharpshooter_fallacy https://en.wikipedia.org/wiki/Texas_sharpshooter_fallacy How many problems and/or times did it make up completely random bullshit with no basis in reality? Random noise looks really cool or impressive when it's right, but if and only if, you're willing to ignore all the times it was wrong. I'm not.
- Hikikomori 3mo agoIt's like reading news articles. Seems reasonable until you read an article about something you know, then you see how wrong they can be.
- deleted 3mo ago[deleted]
- mattgreenrocks 3mo agoYou're not. This site was also bullish on using LLMs as therapists, which defeats the very point of them, and reflects a lack of knowledge on what exactly therapists do for people. More on topic: if the article's author arrived at a definitively negative result would this have shown up on HN?
- rapatel0 3mo agoYou shouldn’t expect frontier models to work on medical imaging. There is much more that goes into building a medical imaging product. First and foremost is data. Medical imaging datasets are not prevalent one the public internet at the scale necessary to have good performance on medical imaging tasks especially MRI. Also the labels are super noisy. This is completely different than asking for general medical reasoning which is more derived from papers, public standards and textbooks. Text exists at the right scale but images don’t.
- pwg 3mo ago> Anytime someone is an actual expert at anything, AI output appears insufficient or incomplete or outright misleading. The term for when the press "gets it wrong" is Gell-Mann Amnesia (https://en.wiktionary.org/wiki/Gell-Mann_Amnesia_effect https://en.wiktionary.org/wiki/Gell-Mann_Amnesia_effect). In that case, when you have personal knowledge of the facts, or know the specific domain area, you can see where the reporter mixed things up. AI is no different, it's just a bunch of matrix math substituting for "the reporter" regurgitating what it was previously told. So the Gell-Mann Amnesia effect would apply just the same. If you have domain knowledge, you immediately see where the AI got it wrong. When you do not have domain knowledge, you have less chance of seeing where the AI was wrong.
- Aurornis 3mo ago> I have seen this pattern over and over again. Anytime someone is an actual expert at anything, AI output appears insufficient or incomplete or outright misleading. It is only when you do not know what the AI is being asked to do is it likely you will find the output helpful. I always recommend people try asking LLMs a lot of questions on something they know first. Programmers should start by asking LLMs to work on a codebase they’re familiar with first. You’re overstating the problem, though. Even for an expert the LLM will get a lot of things right and can be helpful under a watchful eye. The real problem is knowing how to identify when it’s on the right track and when you need to correct it, because both cases are presented with the same tone and confidence. An expert can better identify when the LLM output doesn’t sound plausible. Someone unfamiliar with the topic will think everything it says looks correct.
- jefffoster 3mo agoAI is an expert in everything you are not.
- sevenzero 3mo ago>AI output appears insufficient or incomplete or outright misleading It has been like this since the rise of "AI". The only people enthusiastic about it are usually the ones hoping to make a profit in one way or another.
- stringfood 3mo agowhat is happening is that the gap between what the experts and AI know is getting smaller each year. this year sure radiologists are mocking AI's ability to interpret MRI results, but they are a lot better at that this year than last. In five years perhaps radiologists will truly appreciate AI, but I am not holding my breath because radiologists are notoriously slow to adapt to changes in medical science compared to other specialists like anesthesiologists or surgeons
- jrockway 3mo agoI came here to post this as my experience. AI is magical when I apply it to something I know nothing about. It far exceeds my expectations every single time. I know nothing, but here is a report with animated graphics explaining exactly what I asked it to explain! In fields where I'm an expert... it makes a lot of silly mistakes that are annoying and I feel like they would just cascade if I didn't correct them early. (I still think it's a net win, but... I watch it and it watches me, and we both do better work. I'd even apply the "magical" adjective when it does stuff I hate but know how to do, like edit Helm charts. What would normally be 20 minutes of me griping about YAML indentation is just a correct diff in seconds. I'll take it!) So with that in mind, I tend to distrust output that I can't verify. If a doctor was recommending surgery and I thought the plan was too aggressive, I'd get a second opinion. I don't expect Claude Code to have much medical diagnostic ability, as that is really not what the model is trained for, and I know how it performs on work that it's trained and fine-tuned for. That is not to say the output is wrong and that it can't have diagnostic value, just that I personally wouldn't feel safe trusting it. Wrap up the same model with fine-tuning in the domain and a harness that reminds Claude to do a lot of sanity checks, perhaps with a human in the loop to guide it back onto the rails when it gets hyperfixated on something that doesn't matter? That could very much be a useful AI product.
- aerodexis 3mo agoI am personally very excited by the development of a medical AI-harness that would 1. operates against a well-defined DB of medical studies 2. intakes my basic demographics, vital signs and medical history 3. quantify uncertainty wrt a specific diagnostic (it's own or one received by a healthcare professional) 4. specify medical tests that can be executed, and how they can be obtained 5. provide scripts for interacting with healthcare professionals/functionaries I would imagine that the thing would need distinct operating modes: 1. A diagnosis generator 2. A diagnosis evaluator/critiquer 3. A patient educator
- scosman 3mo agoI dunno. I know a lot of software engineering experts. AI isn't always right, but neither are the people, and it's getting better and better. Software is one domain where it excels because of structured training data and simulation environments, so I'm well aware it's better here than other areas. Still there's somewhere balanced between saying every time it's "insufficient or incomplete or outright misleading" and "just trust AI". AI's a useful source of information/reasoning/research, but know you need to validate it's answers for important decisions.
- gofreddygo 3mo agoThis is true in broader contexts too. Bunch of experts can't agree on something fundamental which is hard to prove/ disprove, and they have strong opinions on the topic. AI is much worse.
- stymaar 3mo ago> I have seen this pattern over and over again. Anytime someone is an actual expert at anything, AI output appears insufficient or incomplete or outright misleading AI assistant are industrializing the Gell-Mann amnesia effect.
- serf 3mo ago>I have seen this pattern over and over again. Anytime someone is an actual expert at anything, AI output appears insufficient or incomplete or outright misleading media is awash at the moment with experts chiming in to support AI, saying their fields are being revolutionized, etc. it seems unsurprising to me that the laymen opinion would follow the loudest media trumpets.
- xivzgrev 3mo agoWell that's part of the problem. AI is not accountable - if you take its advice and hurt yourself, who is responsible? A real doctor is accountable. They might both "know" a lot of things but implicitly the party who is accountable is going to be more trustworthy. And I don't see that going away until AI companies must be licensed for application x and can lose their license / be sued if engaging in malpractice.
- david-gpu 3mo agoLast week I went to a highly-specialized tertiary clinic about further treatment for a rare medical condition that I was diagnosed and treated for as a child. The two very specialized doctors I met there confirmed a diagnostic mistake that a specialist had made ten years ago. The only reason I pursued a second opinion, ten years later, was because Google Gemini had explained to me that the specialist ten years ago had performed the wrong type of test for my condition. Do these LLMs make mistakes? They sure do, I see it all the time. But they can also help people make breakthroughs. And this isn't the only time that Gemini has helped me diagnose long-term health issues, either. I am not advocating to trust anything they say blindly, but they can be a great place to form new hypotheses and learn the right terms to look for when you are unfamiliar with a subject.
- wasabi991011 3mo agoCan you elaborate on how you use Gemini to diagnose long term health issues? Considering doing the same for myself, but I have no idea what is too much vs too little information, and generally the type of prompt engineering to do.
- david-gpu 3mo agoSome folks are not going to like what I am about to say, but what I do is write down as much information that I think may be relevant as possible, trying to avoid leading the witness with any of my preconceived ideas of what may be going on. At the end, I encourage them to ask me questions to get a more complete picture of what may be going on. After a couple of rounds of that, a picture will start to emerge. The AI will make a few XYZ hypotheses of what may be going on, some of which will make more sense to you than others. This is when you can start searching some of those terms in places like pubmed.ncbi.nlm.nih.gov, including for example like diagnostic criteria for XYZ. One of the ways I often use these AIs, not just in the context of finding possible diagnoses, is requesting them to make the case for and against hypothesis XYZ based on the data you have personally collected. Again, it's not about fully buying every thing that comes out of them, but it can help you consider angles or possibilities that did not occur to you, or that you had previously accepted/discarded without sufficient evidence. Think of them as that quirky acquaintance that knows a little bit about everything but sometimes misremembers, rather than as a god-like oracle. And don't do all this in a single session/context. Start a new context every now and then, because otherwise it tends to go in circles as these AIs are biased towards agreeing with whatever it is you said most recently. Intentionally challenge yourself, re-evaluate the existing data from other perspectives. Sometimes what you learn is not pleasant, but as more data becomes available, you learn to accept it. Good luck.
- baxtr 3mo agoThis is a serious issue for young people I think. I have seen outputs that look good but the actual content is bad. If you’re inexperienced in a field you can’t see it because AI makes anything look right. I have gotten very good results with AI but you can’t take the first answer at face value. You need to be suspicious and challenging until you tweak out the right answer over time.
- dang 3mo ago(We detached this subthread from https://news.ycombinator.com/item?id=48709121 https://news.ycombinator.com/item?id=48709121.)