Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
christopheraden
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
christopheraden
7y ago
The Ashley Madison Breach comes to mind. If the core demographic cares about not wanting their data on the platform to get out, they will vote with their feet. That said, I think this example is not the norm, and most people probably won&#x
2.
▲
by
christopheraden
7y ago
> Have fun explaining to your spouse why your household's TV is showing more dating site ads than that of their friends. The targeting is sometimes only _so_ good. While it works in aggregate well, sometimes there are laughable targ
3.
▲
by
christopheraden
8y ago
> I feel like this argument should be a class of fallacy. Lots of things have "existed" for a long time, but that doesn't mean previous iterations were effective. This is pretty close to Survivorship Bias ( https:/&#x
4.
▲
by
christopheraden
9y ago
While it is a policy issue at its core, changing law and policy moves at a glacial pace, and it's not even a certainty that it'll get changed at all (the "nothing to hide" defense gets brought up a lot on these matters,
5.
▲
by
christopheraden
9y ago
Sure, using hypothesis tests could pick out some of the structured examples in the Datasaurus, but in practice, things are often more subtle. Goodness of Fit tests to check for normality, in particular, are a little bit thorny, lacking powe
6.
▲
by
christopheraden
10y ago
A couple questions. > OLS works fine in classification problems. And it has advantages. Do you have more explanation of these advantages? I read through the link you sent, and a bit more about linear probability models. Such things were
7.
▲
by
christopheraden
11y ago
What about something ala Neovim? I've only ever looked at Tex from the perspective of a user (I don't program my own macros too often), so I don't know how hard it'd be, but why not a language overhaul?
8.
▲
by
christopheraden
11y ago
I use SAS professionally at my job, and R in all my academic/hobby work. R has a couple packages that give similar functionality as PROC SQL (about 95% of my SAS workflow, since it's far nicer than data steps for a lot of things).
9.
▲
by
christopheraden
11y ago
>What kind of negative effect might result from a bunch of unqualified high school teachers teaching CS poorly? Is some exposure better than none regardless of teaching quality? Is this problem similar enough to math that we can draw on
10.
▲
by
christopheraden
11y ago
I agree it's a barrier to have such crazy prices, but there are free resources available, especially on a topic as popular as Grammar of Graphics. Hadley Wickham (in my mind, synonymous with the concept, since he implemented Wilkinson&
11.
▲
by
christopheraden
11y ago
Excellent! I guess now is the time for me to finally make the move over to El Capitan, since your usage seems pretty similar to mine. Thanks!
12.
▲
by
christopheraden
11y ago
Can anyone comment if the GM fixes the battery problems Beta 1 had? I was on the first beta awhile back, and my 2012 rMBP got about an hour battery life (I usually get closer to ~3-4) and was always hot to the touch. The same thing happened
13.
▲
by
christopheraden
11y ago
On twitter, the common way to refer to R is rstats, which seems to work okay. See: https://twitter.com/hashtag/rstats
14.
▲
by
christopheraden
11y ago
Tufte-Latex, mentioned in the article, is a really nice template that produces some gorgeous Latex handouts with very little effort (I've used it a couple times when I wanted something to not look like the standard Latex article class)
15.
▲
by
christopheraden
11y ago
Fixxer, I totally agree that it's great that one be capable of doing these things, but sometimes it's not as important as other things that could be taught. Like acbart, sometimes I want to teach why/when to use a statistical
16.
▲
by
christopheraden
11y ago
Very well deserved! With Anaconda, I just tell my students to download a quick installer, slap on iPython or PyCharm, and it's ready to go. It's one less thing to worry about! The installation is dead-simple, and is almost exactly
17.
▲
by
christopheraden
11y ago
Definitely agree. Part of the problem is that data science doesn't have nearly the same formalism in its definition that statistics does. What's the difference between BI's, Data Miners, Data Analysts, Data Scientists, etc? T
18.
▲
by
christopheraden
11y ago
I guess mine is sort of an anti-food-related math concept, then! https://en.wikipedia.org/wiki/No_free_lunch_theorem
19.
▲
by
christopheraden
11y ago
I come from the world of biostatistics, where diagnostic tests are usually measured in terms of Sensitivity (probability of Predicting Evil, given actually Evil, same as Precision) and Specificity (probability of predicting Not Evil, given
20.
▲
by
christopheraden
11y ago
They hint at withholding data ("The same logic applies when it comes time to splitting our data into training and validation sets"), though they don't outright mention cross-validation. Seems like a bit of an oversight for an
21.
▲
by
christopheraden
11y ago
Most likely http://en.wikipedia.org/wiki/Design_of_experiments
22.
▲
by
christopheraden
11y ago
Thank you for posting the data--it makes it easier for us to follow along at home. I've written something up where I used tenure instead of the ranks. http://christopheraden.github.io/SickTime.html . Here’s my concern
23.
▲
by
christopheraden
11y ago
Depends on whether you represent sick-time as continuous or discrete. Zipf's Law takes support on the integers, whereas the Exponential takes support on all positive reals.
24.
▲
by
christopheraden
11y ago
This is an interesting application of statistics! I'm surprised how well the sick-time ranking correlates with the sick times. Intuitively, I'd imagine we'd expect correlation between the ranks and the raw values (your X axis
25.
▲
by
christopheraden
11y ago
Try graphing y = -1 * log(x) and imposing a limit on the upper bound of x and you'll get close to what he has. Perhaps that's the angle he's coming from. He provided the fitted equation further down in the featured article, a
26.
▲
by
christopheraden
12y ago
But then the requirement that the rv's be jointly normal is violated. The "jointly normal + uncorrelated" combination is special. There aren't too many other named distributions that have the property that uncorrelated i
27.
▲
by
christopheraden
12y ago
Could you clarify this statement? Independence is defined in terms of distributions (the joint distribution can be split up into a product of marginals), so I'm not sure how "the way a set of data is distributed" and "ca
28.
▲
by
christopheraden
12y ago
>But if you look at the effect sizes, you see the five studies found nearly the same answers -- its just two of them didn't quite cross the threshold for significance. This is why I wish meta-analysis was introduced much earlier tha
29.
▲
by
christopheraden
12y ago
You are correct about the tissue samples. CCR mostly collects treatment, demographic, and disease information. I didn't get the feeling from their website that FlatIron was collecting tissue samples, either, though.
30.
▲
by
christopheraden
12y ago
It seems like they are doing more than that, though. Cancer data has been collected (by law) and stored for decades. I'd be curious to see if FlatIron was using historical records from Cancer Registries (California's Cancer Regist
More ›