Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
psb217
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
1.
▲
by
psb217
3mo ago
"if you haven't read them you also shouldn't cite them" -- this is wildly incorrect in an academic context. If I'm using ResNets, I should cite the original ResNet paper, even if I haven't read it. If I'm
2.
▲
by
psb217
4mo ago
It seems like they're doing RL to minimize the reconstruction error when going through the: activation -> encoder -> "verbal" description of activation -> decoder -> reconstructed activation loop. Depending on how
3.
▲
by
psb217
8mo ago
Yeah, I assume it was partly chosen since the problem structure provides some convenient hooks for selectively introducing subtle and less subtle inefficiencies in the baseline algorithm that match common optimization patterns.
4.
▲
by
psb217
8mo ago
Per your point 4, some current hyped work is pushing hard in this direction [1, 2, 3]. The basic idea is to think of attention as a way of implementing an associative memory. Variants like SDPA or gated linear attention can then be derived
5.
▲
by
psb217
11mo ago
Yes, you can get good compression of a long sequence of "base" text tokens into a shorter sequence of "meta" text tokens, where each meta token represents the information from multiple base tokens. But, grouping a fixed
6.
▲
by
psb217
11mo ago
The trick is that the vision tokens are continuous valued vectors, while the text tokens are elements from a small discrete set (which are converted into continuous valued vectors by a lookup table). So, vision tokens can convey significant
7.
▲
by
psb217
1y ago
That past work will pay off even more when you start looking into diffusion and flow-based models for generating images, videos, and sometimes text.
8.
▲
by
psb217
1y ago
I think there's an implicit assumption here that interaction with the world is critical for effective learning. In that case, you're bottlenecked by the speed of the world... when learning with a single agent. One neat thing about
9.
▲
by
psb217
1y ago
But, if empirically our current system for net wealth creation tends to also produce wealth concentration, it makes sense to consider ways of modifying the system to mitigate some of the wealth concentration while maintaining as much of the
10.
▲
by
psb217
1y ago
Most of the people pursued in these "AI talent wars" are folks deeply involved in training or developing infrastructure for training LLMs at whatever level is currently state-of-the-art. Due to the resources required for projects
11.
▲
by
psb217
1y ago
Comparing the process of research to tending a garden or raising children is fairly common. This is an iteration on that theme. One thing I find interesting about this analogy is that there's a strong sense of the model's autoregr
12.
▲
by
psb217
1y ago
I think you misunderstood what I meant about setting a high bar. First, passing the bar is a necessary but not sufficient condition for superintelligence. Secondly, by "fair for" I meant it's fair to set a high bar, not that
13.
▲
by
psb217
1y ago
I don't think current models are capable of making abstract links across domains. They can latch onto superficial similarities, but I have yet to see an instance of a model making an unexpected and useful analogy. It's a high bar,
14.
▲
by
psb217
1y ago
I'd say superintelligence is more about producing deeper insight, making more abstract links across domains, and advancing the frontiers of knowledge than about doing stuff faster. Thinking speed correlates with intelligence to some ex
15.
▲
by
psb217
1y ago
You wouldn't get 5 years to noodle -- maybe 1 or 2 at best. You're competing for your next thing against other smart folks who are going hard on maximizing publication rate and grant winning in their current thing. To continue wit
16.
▲
by
psb217
1y ago
One challenge with this line of argument is that the base model assigns non-zero probability to all possible sequences if we ignore truncation due to numerical precision. So, in a sense you could say any performance improvement is due to sh
17.
▲
by
psb217
1y ago
Yeah. It's easy to get over 3000 total daily calories if you have, eg, an hour of cycle commute per day and then add some purposeful gym or running on top.
18.
▲
by
psb217
1y ago
The best way to hit 3000 is cycling. A reasonably fit (70kg-100kg) cyclist should burn 600-800 cal/hr riding at a moderate pace, so 3000 is a 4-5hr ride. It wouldn't be unusual for an enthusiastic amateur cyclist to hit that 1-2x&
19.
▲
by
psb217
1y ago
To be fair, the "trick" part of the kernel trick involves implicitly transforming the data into a higher dimensional space and then fitting a linear function in that space. Ie, you're transforming the inputs so that a linear
20.
▲
by
psb217
1y ago
Offhand, I don't know any specific examples for LLMs. In general though, if you google something like "automated curriculum design for reinforcement learning", you should find some relevant references. Some straightforward sc
21.
▲
by
psb217
1y ago
That depends a bit on the length of the RL training and the distribution of problems you're training on. You're correct that RL won't get any "traction" (via positive rewards) on problems where good behavior isn
22.
▲
by
psb217
1y ago
I think racism accounts for a bigger chunk than you're leaving for it here.
23.
▲
by
psb217
1y ago
Not to mention other aspects of the overall visual experience, eg, everything about scene dynamics, object interactions, etc. A bigger compute budget is always welcome.
24.
▲
by
psb217
2y ago
Autoregressive vs non-autoregressive is a red herring. The non-autoregressive model is still susceptible to exponential blow up of failure rate as the output dimension increases (sequence length, number of pixels, etc). The final generation
25.
▲
by
psb217
2y ago
Natural, cluttered environments are a lot tougher to deal with. This near future-y minimalist environment has the dual benefits of looking stylish and being much closer to whatever they were able to simulate at scale for training the models
26.
▲
by
psb217
2y ago
I think the ubiquity of general disarray and confusion surrounding this agency is relevant to the discussion of whether or not we can trust their intentions and claims. So, further evidence of the extent of this general disarray and confusi
27.
▲
by
psb217
2y ago
The ability to inject your preferred biases into the system that people use for finding or generating nearly all information they consume on a day-to-day basis is extremely powerful. Eg, if all "term papers" produced by this plagi
28.
▲
by
psb217
2y ago
The people deciding to execute layoffs are generally immune to those layoffs.
29.
▲
by
psb217
2y ago
If Henry Ford didn't exist, do you think no one would have tried mass manufacturing? If Jeff Bezos didn't exist, do you think no one would have scaled up an online store? These techno business overlord types occasionally nudge our
30.
▲
by
psb217
2y ago
The problem is in leaving the topic of what to teach open. To some people, this may feel like freedom. To others, in the context of an interview where the purpose is to judge the candidate, it will just lead to a bunch of stress from trying
More ›