6 ms·
I quickly read through some of the claims that the articles make about the studies. While I agree wholeheartedly that there are a lot of false claims about med
by dontreact 5y ago
I quickly read through some of the claims that the articles make about the studies.
While I agree wholeheartedly that there are a lot of false claims about medical AI swirling about, I don’t trust the authors of this paper to tell us where it’s happening based on the mistakes I’ve seen here
“ The remaining studies used enrichment leading to breast cancer prevalence (ranging from 7.4%26 to 73.8%37), which is atypical of screening populations. Five studies used reading under “laboratory” conditions at risk of introducing bias because radiologists read mammograms differently in a retrospective laboratory experiment than in clinical practice. Only one of the studies used a prespecified test threshold which was internal to the AI system to classify mammographic images.”
I’m very close to the authors of McKinney et. al and I happen to know that 2/3 of these claims are either false or specious:
1. The operating point was decided ahead of time before the reader study that was conducted in the paper. I was literally there and saw Scott choose it and I saw his methodology for doing so. So this claim that they did not use a prespecified threshold I have first hand evidence of it being false.
2. Enrichment:
Enriching for positives is absolutely a standard practice and it has mathematically 0 impact on computed metrics. It’s not possible to conduct a reader study without enrichment because cancer is so rare that you would never be able to recruit readers to your study. It just doesn’t make sense at this stage of research to not do enrichment because regardless of what you do the study will still be retrospective.
The paper also shows that results on the unenriched data is basically the same.
I think the understanding that we had at Google and that was mentioned in the paper is that these types of retrospective studies are not sufficient to deploy the AI. It’s probably a good thing this paper is stressing those points.
We need prospective studies, but practically speaking it makes sense to first do really solid retrospective studies. Otherwise a hospital system is not going to just let you run your AI on their scans.
Google is proceeding to do this sort of testing with the work that this paper is just falsely thrashing:
https://news.northwestern.edu/stories/2021/02/artificial-intelligence-breast-cancer/ https://news.northwestern.edu/stories/2021/02/artificial-int...
I feel that this article has good intent (stressing the importance of prospective trials rather than just retrospective studies)
However, the execution is significantly flawed and the outcome seems to be in this particular forum people believing that AI will never work for this application.
These things take time and there is a pipeline of testing and trials that takes many years to go through.
1. Research and development
2. Retrospective testing: as a replacement system (so you don’t have to figure out the UX of how to make it help the human)
3. Retrospective testing: as a helping system (generally scientific papers have been less interested in this step and it would be nice to see more interest here)
4. Prospective testing as an assistive system.
I think the Google system is on step 4 now, and this paper is looking at evidence published from stage 2 and claiming the system does not work.
It’s probably a fair claim because it has not completed all the testing to be deployed. There is as of yet no intent to deploy it as a replacement system, but rather as some sort of assistive system so that prospective evidence can be collected.
However the way they’ve done this is by misrepresenting the paper, and I really think they’ve overstated their case here.