5 ms·
I’m happy to see attention in this area of research. I don't see this referenced in the paper, but a related paper worth reading: "The Medium is the Message: Ho
by kierangill 1mo ago
I’m happy to see attention in this area of research. I don't see this referenced in the paper, but a related paper worth reading: "The Medium is the Message: How Non-Clinical Information Shapes Clinical Decisions in LLMs" [0].
> Through the perturbation of patient messages, we evaluate whether LLM behavior remains consistent, accurate, and unbiased when non-clinical information is altered. […] Our findings reveal notable inconsistencies in LLM treatment recommendations and significant degradation of clinical accuracy in ways that reduce care allocation to patients. […] Our perturbations reflect realistic patient messages from electronic formatting errors and/or simulate patient groups that would be impacted by a wide adoption of patient-AI systems (female patients, non-binary patients or those who use gender-neutral pronouns, patients with health anxiety, patients with a more dramatic disposition, patients with less technological aptitude, and patients with limited English proficiency, etc.)
We’re all peering down the kaleidoscope of a trillion parameter model. It’s no surprise gentle nudges in inputs (grammar, language proficiency, cultural norms) yield different outcomes, despite the intent not changing. It’s one thing to generate crap code, it’s another to generate crap medical advice.
[0] https://dl.acm.org/doi/10.1145/3715275.3732121 https://dl.acm.org/doi/10.1145/3715275.3732121