10 ms·
How to avoid P hacking
- gregwebs 1y agoThis is one of the most disturbing articles I have seen related to reproducibility because it seems to imply that scientists don’t already know this.
- a_bonobo 1y agoAs a biologist all the field wants is p < 0.05. What it actually means is unnecessary. It's a hurdle to pass to have another paper on your CV.
- aaron695 1y ago[dead]
- p4ul 1y agoIf the conclusion is "be transparent", I'm strongly supportive. And moreover, I would be even more supportive if we found a way to change the incentives for tenure and promotion such that reproducibility was an important factor in how we make decisions about grants, tenure, and promotion.
- analog31 1y agoJust make it even more cutthroat than it already is. Replacing one hackable incentive system with another will just produce a new set of hacks. Disclosure: I left academia before I had to worry about any of this.
- neilv 1y ago> As any gambler knows, if you roll the dice often enough, eventually you’ll get the result you want by chance alone You never count your results, when you're sitting at the lab bench, there will be time enough for counting, when the experiments are done.
- boulos 1y agoNicely done. Since many folks may not know the original song: https://en.m.wikipedia.org/wiki/The_Gambler_(song) https://en.m.wikipedia.org/wiki/The_Gambler_(song) (And TIL, this wasn't original to Kenny Rogers!)
- neilv 1y agoI almost did this verbatim quote of the lyrics, which paralleled the article's sentence, and is relevant to P-hacking, but it's the wrong advice: Every gambler knows That the secret to survivin' Is knowin' what to throw away And knowin' what to keep
- saghm 1y agoI don't know, maybe knowing when to "hold them" versus "fold them" and "walk away" would be a valuable skill here. The phrasing sounds off in the part you quite because in poker you only can play a given hand once, and after you've lost, you need to draw an entirely new dataset and start fresh.
- TorKlingberg 1y agoIt depends on what Poker variant you're playing. These days Texas hold 'em is dominant, but Five-card draw used to be very common, especially for informal games.
- saghm 1y agoI'm not sure I understand the relevance. My point is that you can only play a hand once, regardless of the variant, and after that you deal a new hand rather than getting to go back and make different decisions on the same hand.
- smallmancontrov 1y agoIt might be below the fold, but it looks like they're missing the most important p-hacking strategy of all: the dogshit null hypothesis. It's very reliable and it's the most common type of p-hacking that I see. It's easy to create a dogshit null hypotheses by negligence or by "negligence" and it's easy to reject a dogshit null hypothesis by simply collecting enough data as it automatically crumbles on contact with the real world -- that's what makes it dogshit. One might hope that this would be caught by peer review (insist on controls!) but I see enough dogshit null hypotheses roaming around the literature that these hopes are about as realistic as fairy dust. In practice, the dogshit null hypothesis reins supreme, or more precisely it quietly scoots out of the way so that its partner in crime, the dogshit alternative hypothesis, can have an unwarranted moment in the spotlight.
- deleted 1y ago[deleted]
- aw1621107 1y ago> looks like they're missing the most important p-hacking strategy of all: the dogshit null hypothesis Would you mind giving an example(s) of such and how it differs from a "good" null hypothesis?
- gms7777 1y agoNull hypotheses are often idealized distributions that are mathematically convenient and are often over-simplifications of the distributions we'd expect if there were truly no effect (because the expected distributions are either intractable to work with, or irregular and unknown). So for example, suppose you want to detect if there's unusual patterns in website traffic -- a bot attack or unexpected popularity spike. You look at page views per hour over several days, with the null hypothesis that page views are normally distributed, with constant mean and variance over time. You run a test, and unsurprisingly, you get a really low p-value, because web traffic has natural fluctuations, it's heavier during the day, it might be heavier on weekends, etc. The test isn't wrong -- it's telling you that this data is definitely not normally distributed with constant mean and variance. But it's also not meaningful because it's not actually answering the question you're asking.
- shoo 1y agosee also: Andrew Gelman's blog > The problem with p-hacking is not the "hacking," it’s the "p." Or, more precisely, the problem is null hypothesis significance testing, the practice of finding data which reject straw-man hypothesis B, and taking this as evidence in support of preferred model A. https://statmodeling.stat.columbia.edu/2021/09/30/the-problem-with-p-hacking-is-not-the-hacking-its-the-p/ https://statmodeling.stat.columbia.edu/2021/09/30/the-proble... See also this post from 2014 with a discussion of Confirmationist and falsificationist approaches to reasoning in science: https://statmodeling.stat.columbia.edu/2014/09/05/confirmationist-falsificationist-paradigms-science/ https://statmodeling.stat.columbia.edu/2014/09/05/confirmati... > I understand falisificationism to be that you take the hypothesis you love, try to understand its implications as deeply as possible, and use these implications to test your model, to make falsifiable predictions. The key is that you’re setting up your own favorite model to be falsified. > In contrast, the standard research paradigm in social psychology (and elsewhere) seems to be that the researcher has a favorite hypothesis A. But, rather than trying to set up hypothesis A for falsification, the researcher picks a null hypothesis B to falsify and thus represent as evidence in favor of A. > As I said above, this has little to do with p-values or Bayes; rather, it’s about the attitude of trying to falsify the null hypothesis B rather than trying to trying to falsify the researcher’s hypothesis A. > Take Daryl Bem, for example. His hypothesis A is that ESP exists. But does he try to make falsifiable predictions, predictions for which, if they happen, his hypothesis A is falsified? No, he gathers data in order to falsify hypothesis B, which is someone else’s hypothesis. To me, a research program is confirmationalist, not falsificationist, if the researchers are never trying to set up their own hypotheses for falsification. > That might be ok—maybe a confirmationalist approach is fine, I’m sure that lots of important things have been learned in this way. But I think we should label it for what it is. See also: Andrew Gelman and Eric Loken's 2014 "garden of forking paths" paper: https://sites.stat.columbia.edu/gelman/research/unpublished/forking.pdf https://sites.stat.columbia.edu/gelman/research/unpublished/...
- pizlonator 1y agoThe worst part about this: > Running experiments until you get a hit Is that it's literally what us software optimization engineers do. We keep writing optimizations until we find one that is a statistically significant speed-up. Hence we are running experiments until we get a hit. The only defense I know against this is to have a good perf CI. If your patch seemed like a speed-up before committing, but perf CI doesn't see the speed-up, then you just p-hacked yourself. But that's not even fool proof. You just have to accept that statistics lie and that you will fool yourself. Prepare accordingly.
- throwanem 1y agoWhy is this bad for you? You're optimizing software, not trying to describe reality. Monte Carlo and Drunkard's Walk are fine.
- analog31 1y agoYou're churning the user experience for no reason. Maybe constant optimization churn is one of the reasons why UIs are so bad.
- throwanem 1y agoPerf, though? If a perf optimization changes the UI noticeably other than by making it smoother or otherwise less janky, someone is lying to someone about what "performance" means. Likely though that be, we needn't embarrass ourselves by following the sad example. No, UIs churn because when they get good and stay that way, PMs start worrying no one will remember what they're for. Cf. 90% of UI changes in iOS since about version 12.
- appleaday1 1y agoI thought languages such as Rust and flamegraphs and etc were supposed to help us avoid doing all this testing and optimization right? Like I use the built in analysis tools that come with cargo and such and what I have on my os, tools like cutter or reverse engineering tools. Even on python I use the default or standard profiling and optimization tools, I wonder sometimes if I am not doing something enough if the default tools thats recommended should cover most edge cases and performance cases right?
- cypherpunks01 1y agoLike the old saying goes, "It is difficult to get a researcher to stop P hacking, when his career depends on his not stopping P hacking."
- bjornsing 1y agoYeah that was kind of my feeling too while skimming through this: ”Good luck with that…” It’s not a knowledge problem. It’s a vales and incentives problem.
- WhitneyLand 1y agoIt is an old saying, and I’m not sure there’s much use to it as it feels like a mitigation. No doubt the system needs to change, but lots of careers benefit from cheating or unethical behavior. It doesn’t rationalize it or force a choice on anyone.
- eviks 1y agoThe irony of the article appearing in the "career" section when following its advice means you'll not have a career
- gwerbret 1y ago> Stopping an experiment once you find a significant effect but before you reach your predetermined sample size is classic P hacking. Although much of the article is basic common sense, and although I'm not a statistician, I had to seriously question the author's understanding of statistics at this point. The predetermined sample size (statistical power) is usually based on an assumption made about the effect size; if the effect size turns out to be much larger than you assumed, then a smaller sample size can be statistically sound. Clinical trials very frequently do exactly this -- stop before they reach a predetermined sample size -- by design, once certain pre-defined thresholds have been passed. Other than not having to spend extra time and effort, the reasons are at least twofold: first, significant early evidence of futility means you no longer have to waste patients' time; second, early evidence of utility means you can move an effective treatment into practice that much sooner. A classic example of this was with clinical trials evaluating the effect of circumcision on susceptibility to HIV infection; two separate trials were stopped early when interim analyses showed massive benefits of circumcision [0, 1]. In experimental studies, early evidence of efficacy doesn't mean you stop there, report your results, and go home; the typical approach, if the experiment is adequately powered, is to repeat it (three independent replicates is the informal gold standard). [0]: https://pubmed.ncbi.nlm.nih.gov/17321310/ https://pubmed.ncbi.nlm.nih.gov/17321310/ [1]: https://pubmed.ncbi.nlm.nih.gov/16231970/ https://pubmed.ncbi.nlm.nih.gov/16231970/
- coolcase 1y agoSounds like a variable cost experiment. Each observation cost x$. Like an A/B split on Google ads. Why keep paying for A when you know B is better already.
- rrr_oh_man 1y agoGoogle Optimize used to tell you to let an experiment run for one-two weeks (?), exactly because early strong results tend to not don't hold up in the long run. -> https://en.wikipedia.org/wiki/Regression_toward_the_mean https://en.wikipedia.org/wiki/Regression_toward_the_mean
- 1y ago
- parpfish 1y agoI was heavily encouraged to do what would later be called “p-hacking”, but it looked different from what they describe here. This article describes p-hacks for people that aren’t into math/stats. I always ended up p hacking because I was into stats methods. Somebody would say “here’s an old dataset that didn’t work out, I bet you can use one of those new stats methods you’re always reading about to find a cool effect!”, and then the fishing expedition takes off. A couple weeks later you show off some cool effects that your new cutting edge results were able to extract from an old, useless dataset. But instead of saying “that’s good pilot data, let’s see if it holds up with a new experiment”, you’re told “you can publish that! Keep this up and maybe you’ll be lucky enough to get a job someday!”
- AstralStorm 1y agoThe practice you describe is called data dredging though. The thing about it is that you do not know enough experimental design details to make sure it was all on the up, especially worse the older the dataset gets. Normally when doing that you need a multiple comparison corrections and conservative stats. That won't get you published though, or if you do get published you won't get noticed except by someone running a meta analysis. Perhaps not even then. Usually you end up with negative results from reanalysis, evidence of tampering or small effect sizes. And this does not that reliably detect dataset manipulation, p hacking on the part of experimenters or accidental violations of the protocol, not even necessarily if the data collection included measures to prevent it. In short: you cannot 100% trust any dataset you did not make. Not even as part of the team that makes it.
- nlitened 1y agoIf you "dredge" any data set (even the one you can 100% trust) over and over with random hypotheses until p-value is <0.05, you will eventually (actually, pretty quickly) support some false hypothesis. That's why "data dredging" is also p-hacking.
- karma_fountain 1y agoYes, as I understand it there is bias inherent in any dataset due to the fact it is a sample. Data dredging is just looking for that bias. You could do that, but then you'd have to confirm with a new experiment.
- notpushkin 1y ago> You have full access to this article via your institution. Huh. I’m not on a university connection or anything. Is it just open access?
- spinf97 1y ago> Ending the experiment too early > Running experiments until you get a hit But if I'm running an experiment how do I know how many time to run it.
- remus 1y agoBefore you start your experiment, you calculate how many samples you need based on the estimated effect size you're looking for and how small you want your confidence interval to be. Small effect with high confidence => more samples Big effect with low confidence=> less samples
- analog31 1y agoIn the physical sciences you can often estimate the noise level in a null measurement -- or even measure it. You often do this just to get your setup working before doing something like wasting a precious specimen on a "this time for real" measurement.
- zipy124 1y agoThe Bonferroni correction part of this article is the most important. The amount of papers that don't account for this is shocking, comparing 20 variables with a 0.05 confidence interval is extremely annoying, as you end up having to do analysis on all papers data yourself to correct for it to see if it is still significant or not.
- pcrh 1y ago>If you need statistics, you did the wrong experiment. ~Ernest Rutherford.
- biofox 1y ago>If you don't need statistics, you did the wrong experiment. ~Psychologists >What are statistics? ~Computer scientists
- nlitened 1y agoPsychologists are notoriously bad at statistics though
- perrygeo 1y agoIt's not that they suck at statistics. It's that their statistics and experimental designs are artificially stuck in the dark ages. This is forced on the world by the academic publishing industry - you publish this way, or you perish. The completely unsurprising result is a reproducibility crisis that undermines the entire field. Check out "Bernoulli's Fallacy" for a good overview. My theory isn't that Psychologists are bad at statistics. It's that the remaining problems involve lots of messy interactions and messy data that all but require statistical techniques. We just don't have the tools to extract obvious causality amidst such complexity.
- BeetleB 1y agoNot really - it just shows up so much in psychology because they need statistics much more than, say, physics. Most physics programs in the US do not even teach statistics as a subject.
- bossyTeacher 1y agomedicine, biomedical, economics, cancer biology have similar issues hence the reproducibility crisis in those fields
- WhitneyLand 1y agoReading this article tbh causes second hand embarrassment. Ostensibly it’s targeting professional scientists using the brand of a prestigious journal, yet it has a vibe of explaining ethics and common sense to school kids. We’ve come to the point of having to explain to PhDs why cherry picking data is bad. I’m not criticizing the article, rather bemoaning the fact that it’s needed. Of course the problem is not just with the much maligned social sciences, it’s physics and computer science too. The controversy around Microsoft’s topological qubits, a super complex topic, in part involved the most basic kind of this nonsense, something like including 4 samples of 20 measured in the paper iirc. The community needs to get its shit together. The world we’re living in now, the post truth era, is the result of many factors but this is one of them. The loss of faith in science is partially a self-inflicted wound.
- andrewla 1y agoThe article cuts off for me so I do not know if they talk about this, but preregistration has to be part of the conversation moving forward. And it has to have teeth -- withdrawn studies have to have a reputational risk that affects the credibility of future studies, even if it means publishing a retrospective or a null result in a minor journal.
- some_random 1y agoImagine if Nature simply didn't publish obviously p-hacked papers. Perhaps that would do more than a blog post.
- dimal 1y agoWon’t these just make it less likely that you can publish your work, and end up damaging your career in the short term? As opposed to getting published, having a career, with a long tail risk of being found out later? And you could mitigate that risk by publishing research that doesn’t really matter, so no one ever checks.
- freehorse 1y agoThere are many more or less obvious ways that people do p-hacking without even realising it. A classic one is looking at eg an eeg topographic plot, notice which areas or channels within an area seem to be more promising, and running stats and follow ups on these. There are of course degrees of these: people may have preregistered which area (let's say prefrontal cortex for example) but leave open which channels (because it is a bit hard to make that exact guesses anyway). There are methods to deal with this (eg cluster permutation analysis) but often people seem to think that they have to choose between averaging between too many channels, thus risking smoothening out and decreasing an existing effect, or cherry-picking channels based on visual inspection of the data, which means artificially increasing an existing effect or even creating an artifactual one. Because people do not actually run a test to pick the channels, they just visually inspect the data, they do not actually realise this is p-hacking. The problem is that determining the researcher's degrees of freedom is not an easy task, and not one that can just be formalised in a p-adjustment technique. There is a huge spectrum of practices around these degrees of freedom, that may happen during any stage of the data processing, that range from obviously to subtly sketchy and problematic. And believe me that often people who do that think that they actually have good practices, and others do p-hacking. Imo the main way to actually avoid this issue is actually being transparent with all the decisions one makes, even if this can reduce the faith on one's results (which actually should be the point of it, if that's the case!). A lot of time shit happens, and often it is hard to predict everything in advance in a preregistration. If the incentive was to just play safe then not much innovation and method experimentation would occur. It is easy to talk about preregistration as panacea in fields with long ago established practices, but much harder when the state of the art wrt both methods and theory may change wildly even in 2 years that may take to run a study. I believe we need better frameworks for rigorous exploratory research. The only paper I have seen to actually take this idea seriously is this one [0], but I believe a lot of research would more honestly fit in such a framework, and not everything should be conceptualised within a hypothesis testing framework. Method-wise, closed testing procedures also seem very interesting for such research (and can work both actually inferentially, but also for extracting hypotheses for further testing), such as [1]. [0] https://pmc.ncbi.nlm.nih.gov/articles/PMC7098547/ https://pmc.ncbi.nlm.nih.gov/articles/PMC7098547/ [1] https://openpharma.github.io/CTP/articles/closed_testing_procedure.html https://openpharma.github.io/CTP/articles/closed_testing_pro...
- ivansavz 1y agoNon-paywall link: https://archive.is/IJcOI https://archive.is/IJcOI