6 ms·
This article doesn’t identify any statistical problems. It just points out that it would be possible to get a misleading result if you handle your data incorrec
by benchaney 5y ago
This article doesn’t identify any statistical problems. It just points out that it would be possible to get a misleading result if you handle your data incorrectly. It provides no actual evidence that such a mishandling is actually occurring.
- concinds 5y agoThe problem isn't only mishandling of data by statistical analysts, it is merely sufficient that government data be improperly recorded (as was famous last year, when people saw Covid statistics miraculously jump down on weekends and back up on Mondays in several countries) or that vaccinated/unvaccinated populations were misclassified (for example, by only defining people as vaccinated 2 weeks after they had their second dose, i.e. 3 months after the first dose, a rather long time). The article shows that artefacts you'd expect to see if data was improperly recorded, is actually visible in the real government figures. Whether due to the issues discussed, or due to other causes would require deeper analysis. But the onus then goes on people doing statistical analyses to understand, and correct for any mistakes in the originally compiled government data, and not blindly put numbers in Excel without deeply understanding the data they're feeding into it. Any good undergrad statistics prof would cringe at that.
- benchaney 5y agoThe issue is that there is no actual evidence that the data is being recorded improperly. Anyone can come up with a just so story to describe how the data could be corrupted, but writing such a blog post does not in fact put a burden on the people doing actual analysis to refute the claim.
- concinds 5y agoNot how it works. Burdens of proof work in courts of law, but you can't make a statistical analysis without addressing whether the data being used is sound. You'll see a bunch of first-year undergrad papers rely on mediocre data, and then do all sorts of straight-from-textbook statistical analysis of the data, with p-values computed to 5 digits, and then discuss the data's mediocrity in a "Discussion" section. That's a common trope. But if the data is mediocre, you don't have a paper in the first place! Your p-values are junk, so are the abstract and conclusion, and you can't hand-wave it away by bringing up easily-addressed problems with the data in the Discussion section; should have addressed them in the first place. A great physics prof said: "any figure without an uncertainty is meaningless". So is any statistical analysis that doesn't validate the data it relies on. But all that's besides the point, since the article I linked gives good reasons to believe the original government data does display issues in the first place.