Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
dkislyuk
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
dkislyuk
5mo ago
I agree that "it has nice derivatives" is a great empirical reason to use a specific function in ML, but it doesn't sufficiently prove that it's the best function to use. And even if a derivative term looks more complex,
2.
▲
by
dkislyuk
5mo ago
(meant to say, scale-invariance of probability ratios, or shift-invariance of the inputs)
3.
▲
by
dkislyuk
5mo ago
Something that really helped me grasp the foundational relevance of the softmax is to justify from first principles why e^x shows up in the preferred mapping function in the numerator (1). The stated problem of mapping raw inputs/score
4.
▲
by
dkislyuk
5mo ago
Softmax is defined over an arbitrary vector of raw real numbers. Stating that those inputs are "logits" is applying post-hoc semantics to what the model is learning. One of the key properties of a softmax is scale invariance, (e.g
5.
▲
by
dkislyuk
7mo ago
The Paul Cooper production is great. The Rest Is History also just finished a long series (spread out in three seasons, starting on episode 421) on the Punic wars, similarly well done.
6.
▲
Writing Textbooks for Oneself
(dkislyuk.com)
3 points
by
dkislyuk
8mo ago
|
0 comments
7.
▲
by
dkislyuk
1y ago
Presenting information theory as a series of independent equations like this does a disservice to the learning process. Cross-entropy and KL-divergence are directly derived from information entropy, where InformationEntropy(P) represents th
8.
▲
by
dkislyuk
1y ago
From Walter Isaacson's _Steve Jobs_: > One of Bill Atkinson’s amazing feats (which we are so accustomed to nowadays that we rarely marvel at it) was to allow the windows on a screen to overlap so that the “top” one clipped into the
9.
▲
by
dkislyuk
1y ago
Pinterest | Hybrid @ {San Francisco, New York, or Seattle} | Full-time + internships Pinterest’s Advanced Technologies Group (ATG) is an ML applied research organization within the company, focusing on large-scale foundation models (e.g. mu
10.
▲
by
dkislyuk
1y ago
This is a great characterization of self-information. I would add that the `log` term doesn't just conveniently appear to satisfy the additivity axiom, but instead is the exact historical reason why it was invented in the first place.
11.
▲
by
dkislyuk
1y ago
I think commodification is directly tied to a perceived drop in quality. For example, if the barriers to making a video game keep going down, there will be far more attempts, and per Sturgeon's law, the majority will be of low quality.
12.
▲
by
dkislyuk
2y ago
Presumably the book from this thread by Charles Petzold will be a great canonical resource, but originally there was a quote by Howard Eves that I came across that got me curious: > One of the anomalies in the history of mathematics is t
13.
▲
by
dkislyuk
2y ago
Yes, but such a property was not available to Napier, and from a teaching perspective, it requires understanding exponentials and their characterizations first. Starting from the original problem of how to simplify large multiplications see
14.
▲
by
dkislyuk
2y ago
I found that looking at the original motivation of logarithms has been more elucidating than the way the topic is presented in grade-school. Thinking through the functional form that can solve the multiplication problem that Napier was faci
15.
▲
by
dkislyuk
2y ago
Pinterest | San Francisco, New York, or hybrid/remote (US-only) | ML Engineer / Applied Research Scientist | Full-time Pinterest’s Advanced Technologies Group (ATG) is hiring for an engineering position on our visual modeling team
16.
▲
by
dkislyuk
2y ago
Pinterest Advanced Technologies Group | Staff Engineer, iOS and applied ML | US remote or hybrid in SF/NY | Full-time We’re looking for strong engineers to help us build consumer AI products within Pinterest’s Advanced Technologies Gro
17.
▲
by
dkislyuk
2y ago
Yes, exactly. ViTs need O(100M)-O(1B) images to overcome the lack of spatial priors. In that regime and beyond, they begin to generalize better than ConvNets. Unfortunately, ImageNet is not a useful benchmark for a while now since pre-train
18.
▲
by
dkislyuk
3y ago
Rocket Men by Robert Kurson tells the Apollo 8 story in a captivating manner. Some of the passages are quite dramatic but it's justified given the litany of firsts accomplished by the mission.
19.
▲
by
dkislyuk
3y ago
In the current world, deep learning with homogeneous computation graphs, tuned with backprop, has won the Hardware Lottery [1]. This is unfortunate for research outside of that area, but just looking at the momentum of development it seems
20.
▲
by
dkislyuk
3y ago
Moravec's paradox is the usual counterargument given to this line of reasoning. We've had far less progress in embodied robotics, where a robot has to interact with the real world in any kind of generalized, tactile way, compared
21.
▲
by
dkislyuk
3y ago
Uber sold off its own self-driving division back in 2020: https://arstechnica.com/cars/2020/12/uber-sells-self-driving... so this kind of partnership has probably been looming for a while.
22.
▲
by
dkislyuk
3y ago
The one distinction I would add with neural networks is that it's not just a recursive tree traversal that one would get when evaluating an arithmetic statement, but an actual graph: a computation node can have gradients from multiple
23.
▲
by
dkislyuk
3y ago
As another commenter said, viewing a neural network as a computation graph is how all automatic differentiation engines work (particularly reverse-mode where one needs to traverse through all the previous computations to correctly apply the
24.
▲
by
dkislyuk
4y ago
The LLM / foundation model industry will happily stay in the United States and ignore the stifling regulation of the EU if necessary. There is so much market share to capture domestically at the moment.
25.
▲
by
dkislyuk
4y ago
> These platforms are for casuals. At least in road and trail running, this doesn't line up with my observation. There are plenty of elite athletes, at least in the US, who are posting their training on Strava (Jim Walmsley, CJ Albe
26.
▲
by
dkislyuk
4y ago
> New techniques, like diffusion models, shrink down the costs required to train and run inference. Since the postscript mentions that the article was co-written by GPT-3, perhaps this is one of the generated lines? :) (compared to the p
27.
▲
by
dkislyuk
4y ago
The hope is that the Abel Prize becomes the "Nobel Prize of Mathematics", as it is directly modeled after it.
28.
▲
by
dkislyuk
4y ago
Obsidian is a tool like no other: a "second brain" that you have complete ownership over. Being able to link together knowledge or concepts across contexts is a superpower. A recent example: the AlphaTensor paper comes out with a
29.
▲
by
dkislyuk
4y ago
Obsidian is an amazing piece of software. Perhaps PKM as a cottage industry is overdone, and the author's point that everybody's system is highly personal and should not be replicated is valid. That said, Obsidian has allowed me t
30.
▲
by
dkislyuk
4y ago
Obligatory _Sounds Easy_: https://danluu.com/sounds-easy/ Doing anything on Twitter scale has significant technical challenges.
More ›