Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
tysam_and
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
tysam_and
2y ago
I sort of wish that we would move on from the "grokking" terminology in the way that the field generally uses it (a magical kind of generalization that may-or-may-not-suddenly-happen if you train for a really long time). I general
2.
▲
by
tysam_and
2y ago
Nope! Plenty of people get drowsy effects from non-drowsy antihistamines. It is different for everyone (though, again, I am not a doctor!)
3.
▲
by
tysam_and
2y ago
Heyo! Have been doing this for a while. SSMs certainly are flashy (most popular topics-of-the-year are), and it would be nice to see if they hit a point of competitive performance with transformers (and if they stand the test of time!) Ther
4.
▲
by
tysam_and
2y ago
This is incorrect enough as to be dangerous (IMPE, I am not a doctor). They are non-drowsy because they do not cross the blood brain barrier effectively as I understand. Second and third generation antihistamines are fantastic.
5.
▲
by
tysam_and
2y ago
Or even $7.99, that or $8.99 is sort of a nice line between signaling "very cheap game" and "potentially short but enjoyable experience for the evening worth the gamble to find out if so". I can't speak to it in gen
6.
▲
by
tysam_and
2y ago
yeah it's been crazy to see how things have changed and im really glad that theres still interest in optimizing things for these benchmarks. ;P keller's pretty meticulous and has put in a lot of work for this from what i understan
7.
▲
by
tysam_and
2y ago
Yeah, I saw the work from @Sree_Harsha_N, though that accuracy plot on the Adam/SGD side of things is very untuned, it was about what one could expect from an afternoon of working with it, but as far as baselines go most people in the
8.
▲
by
tysam_and
2y ago
hey dont forget about david me and keller (he is currently the champ and has good pareto configs for not just 94 but also 95 and 96 % : https://github.com/KellerJordan/cifar10-airbench )
9.
▲
by
tysam_and
2y ago
This is a pretty hyped-up optimizer that seems to have okay-ish performance in-practice, but there are a number of major red flags here. For one, the baselines are decently sandbagged, but the twitter posts sharing them (which are pretty hy
10.
▲
by
tysam_and
3y ago
Funding is a huge one as well. Funding is the wheel that drives the project (source, have been hanging around the project people for a little while). If you know anyone that would help chip in for the Phase 2 of the project (scaling up, ple
11.
▲
by
tysam_and
3y ago
I get the feeling you may not have read the paper as closely as you could have! Section 8 followed by Section 2 may look a tiny bit different if you consider it from this particular perspective.... ;)
12.
▲
by
tysam_and
3y ago
Yes! This is a consequence of empirical risk minimization via maximum likelihood estimation. To have a model not reproduce the density of data it trained on would be like trying to get a horse and buggy to work well at speed, "now just
13.
▲
by
tysam_and
3y ago
I wish that this worked out in the long run! However, watching the field spin its wheels in the mud over and over with silly pet theories and local results makes it pretty clear that a lot of people are just chasing the butterfly, then afte
14.
▲
by
tysam_and
3y ago
I appreciate the effort that went into this visualization, however, as someone who has worked with neural networks for 9 years, I found it far more confusing than helpful. I believe it was due to trying to present all items at once instead
15.
▲
by
tysam_and
3y ago
Some of the topics in the parent post should not be a major surprise to anyone who has read https://people.math.harvard.edu/~ctm/home/text/others/shanno... ! If we do not have read the foundations of the
16.
▲
by
tysam_and
3y ago
This message confused me on a few dimensions, so I translated it a bit: "State subjective perspective as objective fact. Cast shame upon the OP for not pre-aligning with said belief. Put the responsibility on the OP to prove that they
17.
▲
by
tysam_and
3y ago
This is, among other things, a very natural consequence of some of the equations surrounding and involved in Shannon's original noisy channel capacity theorem, where the noise is (in many ways) conditioned upon the structure of the mod
18.
▲
by
tysam_and
3y ago
Yes! Playing through the rote action exchange can be rather exhausting, especially if I've already bridged that connection and know the person -- there's not much reason for it, and it can be exhausting! Unfortunately, with where
19.
▲
by
tysam_and
3y ago
I mean, again, that's not really the point that I was making. I'm talking about the foundational emotional need of connection, not everyone connects well in that manner, the quality of the response to the question doesn't alw
20.
▲
by
tysam_and
3y ago
Well they can find alternative methods then that are less frazzling, there are fewer things worse than not feeling seen due to only answering questions! I know it can be good, but sometimes the questions can legitimately get in the way of c
21.
▲
by
tysam_and
3y ago
I really hope this stays top comment.
22.
▲
by
tysam_and
3y ago
I think it's honestly quite hard to know, as it's really (generally speaking, AFAIPK) impossible to directly compute the KC in most cases, only really from the feasibility standpoint we can check that it's slower than some ot
23.
▲
by
tysam_and
3y ago
Minor potential performance benefit -- it looks like you might be able to fuse the x_proj and dt_proj weights here as x_proj has no bias. This is a thing that's possibly doable simply at runtime if there's any weight-fiddling reqs
24.
▲
by
tysam_and
3y ago
Oh my gosh, another one-file PyTorch implementation. This is fantastic. I'd like to hope that some of my previous work (hlb-CIFAR10 and related projects, along with other influences before it like minGPT, DawnBench, etc.) has been able
25.
▲
by
tysam_and
3y ago
Another visualization I would really love would be a clickable circular set of possible prediction branches, projected onto a Poincare disk (to handle the exponential branching component of it all). Would take forever to calculate except on
26.
▲
by
tysam_and
3y ago
Highly, highly disagree. If it became obsolete, then y'all were doing the new shiny. The fundamentals don't really change. There are several different streams in the field, and there are many, many algorithms with good staying pow
27.
▲
by
tysam_and
3y ago
You might want to take another look at Shannon's paper, lol, this statement is quite contradictory. Probability _is_ the backbone of information theory, dude! It's quite incredible.
28.
▲
by
tysam_and
3y ago
Yeah, if you read Hinton's long backlog... It's mindblowing. So many fresh concepts, just left there in the dust. Highly encourage. There's a goldmine in there, I thinksies. <3 :')))))))))
29.
▲
by
tysam_and
3y ago
Yeah, it's a very silly article with wrong mathematical reasoning. Hinton is quite obviously talking about a much more information-theoretic approach to the process, but he's phrasing it in people-friendly terms. What's a lit
30.
▲
by
tysam_and
3y ago
This article is unfortunately complete mathematical rubbish. The author appears throughout to show a strong lack of understanding about the mathematics behind what Hinton was saying and the math behind LLMs, and tries to rebut it with casua
More ›