7 ms·
Very interesting article and statistical analysis, but I really don't see how it concludes that the DK effect is wrong based on the analysis. The fact that the
by andersource 4y ago
Very interesting article and statistical analysis, but I really don't see how it concludes that the DK effect is wrong based on the analysis. The fact that the DK effect emerges with _completely random data_ is not surprising at all - in this case the intuitive null hypothesis would be that people are good at estimating their skill, therefore there would be strong a correlation between their performance and self-evaluation of said performance. If the data weren't related, then this hypothesis isn't likely, which is exactly what DK means. And indeed if you look at the plots in the article (of the completely random data), they depict a world in which people are very bad at estimating their own skill, therefore, statistically, people with lower skills tend to overestimate their skills, and experts tend to underestimate it.
Also wanted to point out that in general there is no issue with looking at y - x ~ x, this is called the residual plot, and is specifically used to compare an estimate of some value vs. the value itself.
That being said, the author seems very confident in their conclusion, and from the comments seems to have read a lot of related analyses, so I might be missing something. ¯\_(ツ)_/¯
- uldos 4y agoUnskilled people are more random with their self assessment that skilled. It has nothing to do with unskilled people thinking that they know everything.
- andersource 4y ago> Unskilled people are more random with their self assessment than skilled This is a statistical claim, supported by the DK graph (but not the random data thought experiment from the article). > It has nothing to do with unskilled people thinking that they know everything This reads to me as a claim about the psychological reason for the statistical pattern, which I don't think is either supported nor contradicted by data, both in the article and in the original paper.
- BlueTemplar 4y agoNot the D-K graph, the Nuhfer et al. graph.
- andersource 4y agoThe one reproduced in the article doesn't show the density of points, so it's hard to conclude anything from it. Figure 4 from the Nuhfer et al. paper does seem, to me at least, to support DK's conclusions.
- BlueTemplar 4y agoIt does show the lack of extreme values for higher-skilled people, surely this has some statistical significance ? Especially in a situation where you would expect the distributions to be of the same type ? Unless they had messed up in failing to normalize the number of points per group, and so this might come from the law of large numbers failing + sheer randomness failing to create extreme values on higher-qualified, but lower population groups ?
- deleted 4y ago[deleted]
- notahacker 4y ago"Unskilled people often think they know everything" is pretty much the colloquial use of 'Dunning Krueger' to label people who insist that they know better than the experts from a position of relative ignorance But I think that's consistent with the statistical pattern: if the distribution of self-assessment [amongst unskilled people] of their relative abilities is random or near random, it logically follows that the set of unskilled people includes a lot of people who significantly overestimate their ability at something. Dunning and Kruger don't really talk about the propensity of excellent test performers to underestimate their skill as much (though the Nuher study results somewhat justify their original focus on the ignorant by finding that more skilled groups like professors and graduate students make smaller average prediction errors of their test scores than undergrads). Dunning and Krueger's contention in the original article is that "incompetence robs people of their ability to realise they're incompetent". Similarity of the prediction errors to a random walk isn't a rebuttal of that (although it's a fair critique of the presentation) because the null hypothesis is that people who find a test particularly difficult shouldn't be [almost] as likely to believe they achieved above average performance as the people who aced it. There might be other reasons for that (like the test being pretty easy for all participants and raw scores in a fairly narrow range, or test takers wrongly assuming their lack of understanding was being compared against the general population rather than other smart undergraduates) but in general people ought to be able to incorporate knowing that they didn't know how to answer a lot of questions into their self-assessment of how they performed.
- _dain_ 4y ago>The fact that the DK effect emerges with _completely random data_ is not surprising at all - in this case the intuitive null hypothesis would be that people are good at estimating their skill, therefore there would be strong a correlation between their performance and self-evaluation of said performance. If the data weren't related, then this hypothesis isn't likely, which is exactly what DK means. DK effect is not that low skill people are overconfident and high skill people are underconfident. It is specifically that low skill people are more overconfident than high skill people are underconfident. i.e. if someone's estimated skill is true_skill+bias+noise, then bias_lowskill > -bias_highskill. This is very clear in the original DK paper, they specifically focus on the supposed metacognitive deficiencies of low-skill people. The article argues that the graphs supposedly demonstrating this fact, can also be generated from a model that does not have this difference, i.e. where bias_lowskill == bias_highskill. EDIT: My characterization of the article is not correct, see here[1] for a visualization of the point I'm trying to make. [1] http://emilkirkegaard.dk/understanding_statistics/?app=Dunning_Kruger http://emilkirkegaard.dk/understanding_statistics/?app=Dunni...
- andersource 4y agoThat would make sense, except that graphs generated from random data show identical bias for overestimation and underestimation, as can be seen in the article. And this is opposed to graphs from the DK paper, which show a smaller underestimation bias for experts than an overestimation bias for low-skill people. (Of course that alone doesn't prove anything, just saying that to my understanding, nothing in the article contradicts my interpretation of DK).
- nbernard 4y ago> The article argues that the graphs supposedly demonstrating this fact, can also be generated from a model that does not have this difference, i.e. where bias_lowskill == bias_highskill. But as I understand it, it doesn't: In the graph generated using random data, the lines intersect in the middle (bias_lowskill == bias_highskill), whereas in DK's paper they intersect in the upper right (so bias_lowskill != bias_highskill).
- 4y ago
- IceDane 4y agoI'm definitely not even close to a statistician, but I'm also having a hard time accepting this analysis. I'll admit that part of it also comes from personal experience, at work and elsewhere. I've met some catastrophically incompetent people were completely oblivious to their own incompetence, and this has very often felt like that the more incompetent they were, the more likely they were to be try to do stuff that was waaaay out of their comfort zone, which would make even experienced, competent people tread carefully. But even ignoring personal experiences, I'm not convinced by the arguments either. I understand what they are saying, but I don't see how this disproves the DK effect. Even if everyone is equally bad at estimating their own skill, so that their estimate is essentially a completely random variable, then we would expect the self-assessment score average to be around 50. If I understand it correctly, this is essentially what figure 9 is demonstrating. But that figure still says that worse performers are then likely to overestimate their own ability, just as much as it says that better performers are bad at it. If we look at the original DK figure and contrast it with figure 9 with random data, then I think one way of interpreting the differences is that, yes, worse performers are indeed bad at self-assessment, but they're just kind of bad at it as if their self-assessment is a completely random variable. It then seems to keep being essentially random but as people's skills improve, the distance between their score and their self-assessment becomes a bit tighter.. so in conclusion: most people are pretty bad at self-assessment, but skilled people are a bit less so. The end result is still that people in the bottom quartiles are going to over-estimate their own ability. I don't know, maybe this is way out in the weeds. Please school me.
- js8 4y agoYou might be biased towards avoiding catastrophic risk. Somebody who doesn't know what they are doing without knowing it is more dangerous than somebody who knows what they are doing yet taking precautions as if they don't.
- andersource 4y agoThis is my understanding as well.
- bouncycastle 4y agoThe problem is that they calculated each person’s ‘self-assessment error’ with the actual test score. This error is the difference between a person’s self assessment and their test score, This is like comparing x - y to x, and if you do this, you will get a correlation no matter what.
- omnicognate 4y agoThe fact that the statistical artifact is seen in completely uncorrelated data is only shown as a demonstration that it is not itself evidence of the claimed effect. To gather evidence that the effect doesn't actually exist you need a new experiment, not just a new analysis of the same data, because the the different levels of actual skill need to be established separately from establishing the error in skill self-assessment. In the original experiment the test results were used to establish both, which is not sufficient. But the article presents the results from just such a new experiment. In this one they used university education level (sophomore through to professor) and measured skill self-assessment level within those groups. Higher education level (a good proxy for skill on the test used, which was about science literacy) was found to be associated with more accurate skill self-assessment but the bias of lower-skilled people overestimating their skills was not observed. That's just one study of course, but it sounds like a much better designed one than the original and does constitute actual evidence that the Dunning-Kruger effect doesn't exist.
- andersource 4y ago> The fact that the statistical artifact is seen in completely uncorrelated data is only shown as a demonstration that it is not itself evidence of the claimed effect I don't understand this part. "Completely uncorrelated data" is usually taken to represent the null hypothesis, but that's not the case here. In the DK paper, the implicit null hypothesis is "people of all skill levels are good at estimating their performance". In this case the "completely uncorrelated data" matches an alternative hypothesis, "people's skills have nothing to do with their ability to estimate their performance in tasks testing that skill". This hypothesis doesn't outright contradict the DK proposed hypothesis (and is certainly not the DK null hypothesis), so getting similar results is unsurprising to me, and I'm not sure that we learn from it anything about the DK results. As for the other study cited, the figure shown in the article doesn't give a lot of information on density, and looking at the paper itself, figure 4 does actually seem to show that self-assessment gradually shifts left with increasing level of education. (Edited for accuracy).
- omnicognate 4y ago
- alecbz 4y agoThe author’s confidence is itself an indication that they’re more likely to be wrong. Kidding. Well, half-kidding, I did kind of find the tone a bit biting and dismissive, especially towards one of the commenters that were pointing out exactly what you did. It’s an interesting question to ask whether ask whether the uniformly random data “really” exhibits DK or not, and whether that’s interesting. A world where people have 0 ability to assess their own skill and resort to making uniformly random guesses at it is kind of interesting, and of course in such a world more skilled people would end up on average underestimating themselves and vice versa. But I think the author’s right that obviously nothing psychological is happening here. There’s the psychological effect of no one being able to assess themselves, but the fact that unskilled people overestimate themselves in this world has nothing to do with the fact that they are unskilled.
- andersource 4y ago> But I think the author’s right that obviously nothing psychological is happening here. There’s the psychological effect of no one being able to assess themselves, but the fact that unskilled people overestimate themselves in this world has nothing to do with the fact that they are unskilled. If the results from DK were similar to the random data results, I'd agree. But the DK results do show some correlation between skill and self-assessment ability.
- civilized 4y agoMy guess is that the spread in self-evaluation is largest at low skill level and decreases as skill level increases. This would produce results more similar to what D-K actually observed, and is much more plausible than postulating that highly skilled people have no more idea of their skill level than low skilled people.
- diputsmonro 4y agoWhat strikes me about that graph is that the entire group is likely to be more skilled than the general population. The selection only of people who have the interest and means to attend higher education seems like a narrow window at the furthest edge of the true graph. So I wonder what it would look like if we included people of all education levels and social strata? My gut, based on this article, is that it would look generally the same but with a larger spread at the lower end of the scale. But I don't think we can truly say we've disproved the Dunning-Kruger effect without a more varied dataset.
- leto_ii 4y ago> therefore there would be strong a correlation between their performance and self-evaluation of said performance. If the data weren't related, then this hypothesis isn't likely, which is exactly what DK means. DK doesn't mean no correlation, it means inverse correlation. It's the correct analysis at the bottom that shows what no correlation actually looks like (at least no correlation in tend, there is heteroskedasticity). > a world in which people are very bad at estimating their own skill, therefore, statistically, people with lower skills tend to overestimate their skills, and experts tend to underestimate it. Be careful here, the conclusion you drew doesn't actually follow. > y - x ~ x, this is called the residual plot You're giving x and y meaning that they don't have. In the article these are uncorrelated random variables - the plot of y-x ~ x will always look that way. That's however not the case if you're plotting y_hat - y ~ y_hat for a y_hat taken out of a model. That won't be a random variable in your setup. Edit: note on heteroskedasticity
- tacitusarc 4y ago> > a world in which people are very bad at estimating their own skill, therefore, statistically, people with lower skills tend to overestimate their skills, and experts tend to underestimate it. > Be careful here, the conclusion you drew doesn't actually follow. How does that not follow? It's just regression to the mean.
- dangerlibrary 4y agoI believe the god-emperor was implying that it is possible to imagine a world where people are bad at estimating their own skill, but in the other direction - people who are good at something drastically overestimate how good they are, and vice-versa.
- jkqwzsoo 4y ago> a world in which people are very bad at estimating their own skill, therefore, statistically, *people with lower skills tend to overestimate their skills, and experts tend to underestimate it*. I think the correct conclusion is that if a cohort is not good at estimating their own skill, you can conclude that the variance in their predicted self-assessment will be high (since they’re concluding things without strong evidence), but you can’t assume that their estimates will be biased without additional evidence. That is, indeed, what was shown in the last panel of the article: freshman (and undergraduates in general) are much worse at assessing their own ability than professors, but no group has a strong bias towards over- or under-assessing their own ability (recalling the “my guesses are much better than your guesses” from “The Death of Expertise”).
- bitshiftfaced 4y agoCheck out Nuhfer et al 2016, who had a different explanation for why Dunning Kruger wasn't true. Dunning Kruger effect: lower performers overestimate their ability, and higher performers underestimate their ability. How did they find that? They asked participants to take a test and then had them do a self-assessment. Both were standardized from 0-100. They rated a participant's self-assessment accuracy by "self-assessment minus test score." What's wrong with that method? You can't arrogantly self-assess as though you got a 130, and you can't humbly say that you got -50. Because of the standardization, you're bound by 0 and 100. This method makes it almost impossible for higher performers to overestimate their ability and for lower performers to underestimate. What they actually found was that higher performers tend to be better at self-assessment. Lower performers are less accurate, but in both directions (not just overconfident).
- brnaftr361 4y agoAuthor includes reference to Nuhfer, related studies can be found here: https://digitalcommons.usf.edu/numeracy/vol9/iss1/art4/ https://digitalcommons.usf.edu/numeracy/vol9/iss1/art4/ https://digitalcommons.usf.edu/numeracy/vol10/iss1/art4/ https://digitalcommons.usf.edu/numeracy/vol10/iss1/art4/
- ohwellhere 4y agoIt depends on what one means by the "Dunning-Kruger effect." I had the same impulse that the analysis did not disprove DK, but after sitting with it for overlong I agree with the analysis. I think there are two competing DK effect definitions that are being conflated, one descriptive and one explanatory: 1. DK shows that less skilled people overestimate their ability, and highly skilled people underestimate it 2. DK shows that people's estimation of their ability is causally determined by their actual ability I believe you are claiming, correctly, that the article does not disprove the first definition that explains the observation, but I think the article is trying to disprove the second definition that explains why it occurs. In other words: Yes, there is an observable Dunning-Kruger effect in the sense that we're bad at self evaluation. Is that effect attributable to one's actual level of competence? The evidence for that appears to be a statistical artifact, and further experiments seem to disprove that conjecture. I'm not a statistician or a psychologist.
- kenjackson 4y ago> Also wanted to point out that in general there is no issue with looking at y - x ~ x, this is called the residual plot, and is specifically used to compare an estimate of some value vs. the value itself. This article seemed very unconvincing -- and this part noted above, early on in the article set the tone that I felt like the author didn't know what they were doing. And even after reading it all, I felt like the standard lay use of DK remained valid. This just felt like the type of thing I would have thought about as an undergrad, started to write it, and then realized it didn't make sense halfway through it. Or maybe I just missed something as well...
- Accujack 4y ago>the author seems very confident in their conclusion They are, but honestly all that can be concluded safely IMHO is that the original D-K graph doesn't support that the widely discussed "effect" which their conclusion describes exists. Therefore unless there is more evidence from some other subsequent study there may not be any evidence for it at all, and if that's the case then there's potentially no proof it exists. However, even if you prove that their data is not evidence, that doesn't actually say anything about whether the effect exists or not, just that the D-K paper isn't evidence of such an effect. I don't think that's enough for the author to conclude that "DK is autocorrelation". A more careful conclusion would be that "the DK data do not support DK's conclusion"... but of course that's much less likely to attract click throughs.
- hgomersall 4y agoThis is it. You need to write down the model then perform the inference. With some rudimentary munging you can get a super simple linear model.
- usefulcat 4y agoIMO the most interesting thing is not so much that you can get DK from noise, it's that the Nuhfer study was utterly unable to replicate the DK effect. If DK is real, there should have been at least a hint of it visible in the Nuhfer study.
- longtimegoogler 4y agoYup. THat's what I was going to say. There data suggests that everyone kind of estimates their ability similarly so that more skilled people underestimate there ability (impostor's syndrome) and less skilled people overestimate the their abilities.
- pdonis 4y ago> I really don't see how it concludes that the DK effect is wrong based on the analysis. Neither do I. Basically what the article actually shows is that these two statements are equivalent: (1) People with low test scores tend to overpredict their test scores, while people with high test scores tend to underpredict their test scores. (2) People's predictions of their test scores are uncorrelated (or more precisely very weakly correlated [1]) with their actual test scores. This is not a statement that the D-K effect is wrong. It's just restating what the D-K effect is in different words. All the talk about "autocorrelation" is just another way of saying that, if people's predictions of their test scores are only weakly correlated with their test scores, then people with low test scores will have to overpredict their test scores (because there's virtually no room to underpredict them--there's a minimum possible test score and their actual score is already close to it), and people with high test scores will have to underpredict them (because there's virtually no room to overpredict them--there's a maximum possible test score and their actual score is already close to it). But the real question is: why are x and y so weakly correlated? Why are people's predictions of their test scores so weakly correlated with their actual test scores? That is not what one would intuitively expect. That is the question the D-K effect raises, and the author not only doesn't answer it, he doesn't even see it. Also, this statement in the description of the Nuhfer research doesn't make sense: "What’s important here is that people’s ‘skill’ is measured independently from their test performance and self assessment." Um, the test performance is the people's "skill". And in the original D-K research, it was "measured independently" from the people's self-assessment (their prediction of their test performance). [1] Notice that in the "uncorrelated data" graph, Figure 10, the red line is basically horizontal. That's what you get when x and y are uncorrelated. But in the original D-K graph, Figure 2, the thick black line is not horizontal--it slopes upward. That's what you get when x and y are weakly correlated. If the author had put in a weak correlation between x and y in his own experiment, he would have gotten a graph that looked like Figure 2. But of course that still would do nothing to explain why x and y are so weakly correlated, which is the actual question.
- hgomersall 4y agoMy issue with the thought process is the general picking of uniform distributions to show random data that follows DK. Uniform distributions are kind of uninteresting in this case because they're pretty artificial. It's not like competence is actually measured in the range 0-100 in anything that actually matters.
- uoaei 4y ago> the intuitive null hypothesis would be that people are good at estimating their skill, therefore there would be strong a correlation between their performance and self-evaluation of said performance That's not really the intended interpretation of "null" in "null hypothesis". "Null" does not mean "contrary to the effect you're testing". "Null" means "do not assume dependencies anywhere" and so your description is backwards.
- andersource 4y ago"Null hypothesis" comes down to agreeing on a prior that is "reasonable". Most of the time, that indeed means not assuming dependencies, e.g. when testing the outcomes of a medical treatment. But that's not always the case, e.g. does it seem reasonable to you to assume a priori no dependency between a person's age and their height? Does the result "People get taller until the age of 20" merit a journal article? There's a very strong correlation there after all. As I've written elsewhere this all comes down to the prior, based on my life experience the prior "people are generally capable of self-assessing their performance" is much more likely than "people have absolutely no ability to self-assess their performance." To state it differently, when you assume no dependencies anywhere, you've already jumped to a conclusion that is more far-fetched to me than the results of DK. Do you really think people have zero ability to self-assess their own performance? In all domains and all contexts, as this is your "null hypothesis"?
- uoaei 4y agoI invite you to ponder on the etymology and actual definition of the word "null". It does not refer to some antecedent or to some notion of "reason", it refers as explicitly as possible to "not-x", contra whatever effect generates "x". You are conflating the construction of the null hypothesis with a more liberal notion of hypothesis testing (science) per se.
- hgomersall 4y agoNull means "that thing which I've decided to attach special significance to". There's in general no reason to prefer a null hypothesis over any other hypothesis.
- a-dub 4y agoi immediately jump to think that replacing all the data with noise is a pretty good null hypothesis (at least for the analysis). is that not true?
- rawgabbit 4y agoYou're changing the definition of the null hypothesis. What you are essentially saying is that there is no need to perform experiments and studies, I can arbitrarily grab data from anywhere and if it fails to support the hypothesis... then the hypothesis is wrong.
- a-dub 4y agoif your analysis produces an effect when fed uniform random data, then yes, no need to perform experiments or studies, because all your results are null. right? so the hypothesis would be something like "the analysis shows a relationship between the two variables" and the null hypothesis would be something like "the analysis shows a relationship between two uniform random variables" and in this case that null is shown and accepted because no such relationship exists by definition. right? (unless it's like, "they have the same entropy", or something) i'm very rusty with this stuff, so clarification would be much appreciated!
- a-dub 4y agonote that i have not carefully been through the claimed analysis here (and specifically have doubts about the data v. error plot) but if the claim that the analysis produces the effect with random inputs is assumed, then the whole thing can be rejected at step 0, right?
- rawgabbit 4y agoIn the case of DK, the hypothesis is that there is a bizarre relation between perceived skill and actual skill. The null hypothesis is that is there no relation or correlation between perceived skill and actual skill. A researcher would perform an experiment. If researcher observes statistically significant results. Then the researcher can reject the null hypothesis, and say the theory is valid. If researcher does not observe statistically significant results. They can only say their experiment doesn't support the theory.
- msrenee 4y agoPlease feel free to tell me why my interpretation is wrong. I understand stats just well enough to get myself in trouble. The line for actual ability is basically x=y. If you scored 10%, you're in the bottom quartile. If you scored 100%, you're in the top quartile. That line isn't really data, just something for comparison. The perceived ability line is the one that utilizes the data. It seems to show that once you average out what everyone rated themselves, it ends up kind of in the middle between ~55-70%. So the people who scored 10% assumed, on average, they would score around 55%. The people who scored 100% assumed, on average, that they would score about 75%. That makes the average expected score much higher than the actual score on the low end and somewhat lower than the actual score on the high end. I'd interpret this as the bottom quartile thinks they're average and the top quartile thinks they're a bit above average. So basically everyone thinks they're average-ish, but the people who did worst on the test were the most wrong about that. But then again, I can't remember what the questions on the test were even about and just the single graph isn't terribly useful to argue over because it's missing all of the context of the paper. Now that I've sat and interpreted the graph using my own set of notions about what the numbers mean and what the graph actually shows, I feel like this ought to be used as one of those life lessons about how quotes and diagrams outside of their context within a paper are the epitome of the phrase "lies, damn lies, and statistics." Statistics aren't always lies, but they're incredibly easy to bend to your own biases and assumptions.
- krick 4y agoYou've been corrected about "what DK means" in the other comments, but this is not quite the point of the post. This is not about if DK (as expressed in English words) is true or not — in fact, author points out in the beginning that it's one of these "everybody knows it's like that" ideas (as it often is with social psychology). The point is, that the original DK paper is bullshit. At least, this plot is. And people tend to miss it, until they start to carefully read the labels and think about the caveats. In fact, as presented here it looks like it shouldn't even be accepted as a valid study, this is outright deceptive, maliciously so. If there is assumed to be a correlation between x & y, how about we start by plotting x against y then? I know, it may be messy. It almost certainly will be. Because of that, I personally won't even be offended (but some people might) by you removing the outliers and producing the unnaturally clean version of the plot in the end to highlight the main idea. Then some statistical tests to make the results quantified. But here we see nothing, it really is just comparing x to x. IMO, this is pretty much the invariant of most of the problems of academic research in the last God-knows-how-many decades (maybe always was, I don't know). Computer science papers without the code. Data science papers without the data. Yeah-yeah, I've heard hundreds of excuses why researchers do it like that. But it's pointless, such "research" shouldn't be accepted by anybody. Either you make your findings actually public by providing everything to replicate every single step of your study (which is supposed to be the point), or you just don't publish anything and keep the research proprietary (I mean, obviously it's never black and white, there always will be concerns about test-subject anonymity, etc. — but it's ridiculous to discuss that when the accepted standard even in "proper" sciences are 20 pages of dense text which might never even get to the point of the study, i.e., actually showing the data to any extent.)
- andersource 4y agoI strongly disagree, not necessarily with everything (e.g. I don't have access to the raw data from the DK experiment, don't know how well they performed all the analysis leading to the plot). But the plot itself is not inherently deceptive, and, unlike implied in the article, is not equivalent to "just comparing x to x". The plot essentially shows the actual performance vs. self-assessment of performance, compared to what we would expect if there were perfect correlation.