6 ms·
It's not a strawman. There are many fundamentally unpredictable things where we can't make the benchmark be 100% accuracy. To make it more concrete on work I a
by dontreact 3y ago
It's not a strawman. There are many fundamentally unpredictable things where we can't make the benchmark be 100% accuracy.
To make it more concrete on work I am very familiar with: breast cancer screening. If you had a model that outperformed human radiologists at predicting whether there is pathology confirmed cancer within 1 year, but the accuracy was not 100%, would you want to use that model or not?
- frumper 3y agoIt's a strawman because they aren't comparable to AI detection tests. A screening coming back as possible cancer will lead to follow up tests to confirm, or rule out. An AI detection test coming back as positive can't be refuted or further tested with any level of accuracy. It's a completely unverifiable test with a low accuracy.
- dontreact 3y agoYou are moving the goalposts here. The original claim I am responding to is "A tool that gives incorrect and inconsistent results shouldn’t have any part of a decision making process." I agree that there are places where we shouldn't put AI and that checking whether something is an LLM or not is one of them. However I think the sentence above takes it way too far and breast cancer screening is a pretty clear example of somewhere we should accept AI even if it can sometimes make mistakes.
- frumper 3y agoThe thread is about tools to evaluate LLMs. Please re-read my comment in that light and generously assume I'm talking about that.