5 ms·
Twenty-nine teams use same dataset, find contradicting results [pdf]
- dang 11y agoA blog post giving background is at http://www.nature.com/news/crowdsourced-research-many-hands-make-tight-work-1.18508 http://www.nature.com/news/crowdsourced-research-many-hands-....
- jdp23 11y agoThe blog post is a great overview as well as useful context, thanks for sharing it. TL:DR summary: Scientific results are highly contingent on subjective decisions at the analysis stage. Different (well-founded) data analysis techniques on a fairly simple and well-defined problem can give radically different results. It's very interesting research -- a great real-life example supporting the models Scott Page et. al. use for the value of cognitive diversity. The thrust of the blog post is about where crowdsourcing analysis can be helpful (as well as reasonable caveats about where it might not apply), which is certainly an interesting question. Obvioulsy, there are a lot of other implications to this as well.
- nkurz 11y agoOff-topic, but you seem well positioned to answer: Why do you say "TL:DR" here when summarizing a short blog post that you enjoyed? Clearly the meaning has diverged from the original abbreviated insult of "Too long; didn't read", but I don't understand what people mean when they use it today. Why did you phrase it this way? Are you a native English speaker? If not intended to be derogatory, does the dissonance bother you?
- justinlardinois 11y agotl;dr stated off as a way of saying "this is too long and therefore requires too much effort for me to read it." Then that gave rise to people accompanying long reads with a "tl;dr version," which is usually a one or two sentence summary. Now that the latter is common and understood people just write tl;dr and then follow it with the summary, with the understanding that those who are unwilling to read the full version will read that instead.
- jdnier 11y agoHow do you see it as derogatory? I'm a native English speaker and have never thought of it that way. I didn't click the link, but did appreciate his short summary – and upvoted him for it. ;)
- taneq 11y ago'tl;dr' is often a troll response to a long post that that someone has obviously spent a lot of time on. Bonus troll-points if the long post was in response to another troll. Example: poster1: only retards use vi, notepad rules poster2: huge list of reasons why vi is better than notepad poster1: lol tldr
- ZenoArrow 11y ago> "'tl;dr' is often..." But not always. You have to consider how it was used to tell whether it was meant with ill will, the term tl;dr on its own doesn't necessarily tell you enough.
- dalke 11y agoI was quite peeved when I first saw a "tl;dr" comment concerning one of my blog posts. My thought was, and still is somewhat, "if you didn't read it, how can you say it was too long for what it needed to cover? Why do you feel the need to tell others that you have the attention span of a fly?" We already have terms like "summary", "digest", and even "précis"; why create a new term imbued with snark?
- nkurz 11y agoI also appreciated the summary; my question was just about the phrasing. In it's literal usage, it's saying that the article had nothing useful to say: http://knowyourmeme.com/memes/tldr http://knowyourmeme.com/memes/tldr. That clearly wasn't the case here, so I was wondering why the author choose to use it. I realize that meaning has changed over time, but I was wondering what meaning he (and others) intend when it is used.
- jdp23 11y agoLike 'justinlardinois I just use "TL:DR summary" to mean a "short summary." I didn't mean it as derogatory, and it doesn't feel dissonant to me -- although given how you view it I can see why you do. And yes, I'm a native English speaker.
- hnmcs 11y ago> I don't understand what people mean when they use it today By now, it often is used to be friendly. There's a subconscious acknowledgement that long words take people's time. The speaker can even be talking about his own work and tell everyone "tldr version: " at the top. Google uses it a lot in their own docs: https://developers.google.com/s/results/?q=tldr https://developers.google.com/s/results/?q=tldr You can also choose to pronounce it "teal deer" and use images of a green deer animal to signify the same. ... In some situations, if you want, you can say it to be mean.
- sndean 11y agoFiveThirtyEight did a write up of this paper (part 2): http://fivethirtyeight.com/features/science-isnt-broken/ http://fivethirtyeight.com/features/science-isnt-broken/ On the bright side, if you look at the 95CI for the 29 studies, almost all of them overlap.
- hmate9 11y agoLies, damned lies, and statistics https://en.wikipedia.org/wiki/Lies,_damned_lies,_and_statistics https://en.wikipedia.org/wiki/Lies,_damned_lies,_and_statist... Statistics can be manipulated surprisingly easily.
- justinlardinois 11y agoThere are three kinds of lies. There are also three kinds of comments I see in this thread: > "This is interesting, here's some thoughts and ideas that further contribute to this subject" > "This is interesting, here's a link to some further writing on this subject" > "The entire concept and discipline of statistics is bullshit." Par for the course here at Hacker News.
- duaneb 11y agoHey, the null hypothesis is powerful and valuable. I, for one, and happy that all three types are well-represented; all three are healthy in moderation. I also think that the quote fits in quite nicely here, it's not a wholesale rejection of statistics.
- justinlardinois 11y agoThe quote itself isn't, but just posting a link to the Wikipedia article without explaining how they think it applies here is pretty much a wholesale rejection of statistics.
- jdp23 11y ago4. > "About TL:DR ...:" Also par for the course here :)
- joe_the_user 11y ago"The primary research question tested in the crowdsourced project was whether soccer players with dark skin tone are more likely than light skin toned players to receive red cards from referees." This seems like a topic where one indeed typically winds-up with a multitude of competing conclusions. Among other factors for we have: * Pre-existing beliefs on the part of researchers. * Lack of sufficient data. * Difficulty in defining hypothese (is there a skin tone cut-off or should one look for degrees of skin tone and degrees of prejudice, should one look all referees or some referees). Given this, I'd say it's a mistake to expect just numeric data at the level of complex social interactions to be anything like clear or unambiguous. If studies on topics such as this have value, they have to involve careful arguments concerning data collection, data normalization/massaging, and only then data analysis and conclusions. But a lot of the context comes from prevalence shoddy studies that expect you can throw data in a bucket and draw conclusions, further facilitated having those conclusions echoed by mainstream media or by the media of one's chosen ideology.
- entee 11y agoThis paper is awesome because it transparently folds the analytical approach into the experiment being conducted. There are two kinds of scientific study: those where you can run another (ideally orthogonally approaching to the same question) experiment along with rigorous controls, and those where you can't. The first type is much less likely to have results vary based on analytical technique (effectively the second experiment is a new analytical technique). Of course it does happen sometimes and sometimes the studies are wrong, still more controls and more experiments are always more better. However, studies were you're limited by ethical or practical constraints (i.e. most experiments involving humans) don't have that luxury and therefore are far more contingent on decisions made at the analysis stage. What's awesome with this paper is it kind of gets around this limitation by trying different analytical methods, effectively each being a new "experiment" and seeing if they all reach the same consensus. Interestingly, very few features in the analysis were shared among a large fraction of the teams, (only 2 features were used by more than 50% of teams) which suggests that no matter the method, the result holds true. A similar approach to open data and distributed analysis would be a really great way to eliminate some of the recent trouble with reproducibility in the broader scientific literature.
- SilasX 11y agoReminds me of the idea (Robin Hanson's, I think?) to add an extra layer of blindness to studies: during peer review, take the original data, and write a separate paper with the opposite conclusion. Randomize which reviewers get which version. Your original paper is then only accepted if they reject the inverted version.
- gwern 11y agoI think you misremembered it: http://www.overcomingbias.com/2007/01/conclusionblind.html http://www.overcomingbias.com/2007/01/conclusionblind.html http://www.overcomingbias.com/2010/11/results-blind-peer-review.html http://www.overcomingbias.com/2010/11/results-blind-peer-rev... Nothing about accepted only if they rejected the reversed version; just that the pro & con versions be supplied (first post), or a paper sans conclusions/results (second post).
- DiabloD3 11y agoSo, does this mean every team used improper methodology? Or can we meta-review the results and figure out what's really going on?
- DanBC 11y agoIt makes it really hard to work out what's happening, especially if you want the result to match existing standards. For a real world example of this see deworming schoolchildren. People looking at the educational effects of deworming children reach different conclusions because some of them use a medical model and some of them use an economics model. http://www.cochrane.org/news/educational-benefits-deworming-children-questioned-re-analysis-flagship-study http://www.cochrane.org/news/educational-benefits-deworming-... http://www.cochrane.org/CD000371/INFECTN_deworming-school-children-developing-countries http://www.cochrane.org/CD000371/INFECTN_deworming-school-ch... Talked about in this More or Less episode: http://www.bbc.co.uk/programmes/b0659q1f http://www.bbc.co.uk/programmes/b0659q1f http://www.theguardian.com/society/2015/jul/23/research-global-deworming-programmes http://www.theguardian.com/society/2015/jul/23/research-glob...
- krick 11y agoI understand how tempting it is in our age of big data and all that stuff to perceive this as some curious new phenomena, but it really is not. This is precisely the reason why we've come up with some criteria for "science" quite a while ago. And in fact, all this experiment is pretty meaningless. So, for starters: 29 students get the same question on the math/physics/chemistry exam and give 29 different answers. Breaking news? Obviously not. Either the question was outrageously bad worded (not such a rare thing, sadly), or students didn't do very well and we've got at most 1 correct answer. Basically, we've got the very same situation here. Except our "students" were doing statistics, which is not really math and not really natural science. Which is why it is somehow "acceptable" to end up with the results like that. If we are doing math, whatever result we get must be backed up with formally correct proof. Which doesn't mean of course, that 2 good students cannot get contradicting results, but at least one of their proofs is faulty, which can be shown. And this is how we decide what's "correct". If we are doing science (e.g. physics) our question must be formulated in a such way that it is verifiable by setting up an experiment. If experiment didn't get us what we expected — our theory is wrong. If it did — it might be correct. Here, our original question was "if players with dark skin tone are more likely than light skin toned players to receive red cards from referees", which is shit, and not a scientific hypothesis. We can define "more likely" as we want. What we really want to know: if during next N matches happening in what we can consider "the same environment" black athletes are going to get more red cards than white athletes. Which is quite obviously a bad idea for a study, because the number of trials we need is too big for so loosely defined setting: not even 1 game will actually happen in isolated environment, players will be different, referees will be different, each game will change the "state" of our world. Somebody might even say that the whole culture has changed since we started the experiment, so obviously whatever the first dataset was — it's no longer relevant. Statistics is only a tool, not a "science", as some people might (incorrectly) assume. It is not the fault of methods we apply that we get something like that, but rather the discipline that we apply them to. And "results" like that is why physics is accepted as a science, and sociology never really was.
- platz 11y agoPhysics uses statiatics all the time, e.g. detecting the higgs boson at cern. Do you have a formal proof thay each time they fired the accelerator it was going to be i.i.d.?
- LunaSea 11y agoOf course it's social "sciences".
- DanBC 11y agoAlso medical treatment: https://news.ycombinator.com/item?id=10387375 https://news.ycombinator.com/item?id=10387375
- josh5555 11y agoSerious question - what if players of a certain skin color break the rules more frequently than another skin color? "Bias" is extremely subjective and complex. Was this possibility taken into account?