Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
jaidhyani
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
18 ms
·
1.
▲
by
jaidhyani
5mo ago
I used to work for Meta. I quit largely because of intense frustrations with the company. Meta has made a lot of mistakes, overlooked a lot of harms, and made a lot of short-sighted, selfish choices. Many things about the world are worse th
2.
▲
by
jaidhyani
5mo ago
Said company is literally in court against said government at the moment, after said government attempted to designate it too dangerous to do business with.
3.
▲
by
jaidhyani
1y ago
Approximately no one in the community thinks this. If you can go two days in a rationalist space without hearing about "Chesterton's Fence", I'll be impressed. No one thinks they're 100% rational nor that this is a
4.
▲
by
jaidhyani
3y ago
Compare the trajectory of the US to other industrialized countries. The best charts I could find on this are from an admittedly-biased think tank, but the sources it's pulling from are well-regarded and neutral: https://www.
5.
▲
by
jaidhyani
3y ago
https://ourworldindata.org/renewable-energy Quick stats for the US: In 2022, 11.3% of energy was generated by renewables (hydropower, solar, wind, geothermal, bioenergy, wave, and tidal). It's been growing at just unde
6.
▲
by
jaidhyani
3y ago
> His actions made perfect sense from his utilitarian Effective Altruist worldview. They don't. Everyone in EA (AFAICT) has been pretty clear about this. Lying and undermining trust and institutions does tremendous lasting harm. I a
7.
▲
by
jaidhyani
3y ago
I will never cease to wonder at how so many people can blame so much on people trying to take a rigorous approach to world improvement, up to and including "a narcissistic con-man claimed to do trying to do X, and I can imagine a scena
8.
▲
by
jaidhyani
3y ago
I am begging people to stop confusing "I was unable to get LLM X to do Y using strategy Z" with "All LLMs are categorically unable to do Y".
9.
▲
by
jaidhyani
3y ago
As the other commenter said, this is incorrect. The input was a sequence of legal moves (not even "real" moves - most of the training data was synthetically generated with "generate legal moves" as the only constraint).
10.
▲
by
jaidhyani
3y ago
Alternatively, the prior on "this is not possible" is very low because RLHF & Friends have targeted metrics that, inadvertently or not, discourage that outcome.
11.
▲
by
jaidhyani
3y ago
Smallpox eradication
12.
▲
by
jaidhyani
3y ago
Going to be extremely hard to quantify, ransomware peddlers aren't famous for their meticulous public record-keeping. You could try to sift through all the transactions on the public blockchain and try to classify the ransomware ones,
13.
▲
by
jaidhyani
3y ago
Could have gone with "More Comprehensive Metrics Are All You Need"
14.
▲
by
jaidhyani
3y ago
This is a weird future.
15.
▲
by
jaidhyani
3y ago
GPT-LikeSubscribeAndRingThatBell
16.
▲
by
jaidhyani
3y ago
This is true in general but not in the use case they presented. If they had explained why a normalized distribution is useful it would have made sense - but they just describe this as pick-the-top-answer next-word predictor, which makes the
17.
▲
by
jaidhyani
3y ago
Prediction happens at the very end (sometimes functionally earlier, but not always) - most of what happens in the model can be thought of as collecting information in vectors-derived-from-token-embeddings, performing operations on those vec
18.
▲
by
jaidhyani
3y ago
It depends on the values of the vectors. (4, 4) + (3, 3) results in a new vector (7, 7) which is further away from both contributing vectors than either one was to each other originally. Additionally, negative coefficients are a thing.
19.
▲
by
jaidhyani
3y ago
The original paper is very good but I would argue it's not well optimized for pedagogy. Among other things, it's targeting a very specific application (translation) and in doing so adopts a more complicated architecture than most
20.
▲
by
jaidhyani
3y ago
I endorse all of this and will further endorse (probably as a follow-up once one has a basic grasp) "A Mathematical Framework for Transformer Circuits" which builds a lot of really useful ideas for understanding how and why transf
21.
▲
by
jaidhyani
3y ago
That's true, but they didn't go into any other applications in this explainer and were presenting it strictly as a next-word-predictor. If they are going to include final softmax, they should explain why it's useful. It would
22.
▲
by
jaidhyani
3y ago
Thanks, that's a really useful intuition!
23.
▲
by
jaidhyani
3y ago
TIL. Man, I'm behind on my paper reading.
24.
▲
by
jaidhyani
3y ago
I am the guy asked and I endorse this guy's endorsements.
25.
▲
by
jaidhyani
3y ago
The way the article presents this is misleading. The attention mechanism builds a new vector as a linear combination of other vectors, but after the first layer these have also all been altered by passing through a transformer layer so it m
26.
▲
by
jaidhyani
3y ago
Skimming it, there are a few things about this explanation that rub me just slightly the wrong way. 1. Calling the input token sequence a "command". It probably only makes sense to think of this as a "command" on a model
27.
▲
by
jaidhyani
3y ago
I think we're on the verge of few-shot believable voice impersonation. Between that, realtime deepfake videos, and AIs being more than good enough to solve CAPTCHAs, it seems like we're at most a few years from having no means of
28.
▲
by
jaidhyani
3y ago
When a dangerous exploit is discovered, best practice is to give good actors playing defense a heads up and time to prepare before publicly revealing the exploit, and preferably only revealing the exploit once critical systems have been har
29.
▲
by
jaidhyani
4y ago
The content moderators were fired some time ago.
30.
▲
by
jaidhyani
4y ago
The Ghostwriter expansion pack was next fucking level. Clues hidden in documents to solve a mystery - one of my favorite early childhood software experiences.
More ›