Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
gwd
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
gwd
5d ago
This sounds a lot to me like people in the 90's complaining that computers were destroying chess. Thirty years later, chess is more popular than it ever was, and chess players are better than they ever have been. I wouldn't be s
2.
▲
by
gwd
7d ago
It's a neural network rather than a bunch of hard-coded rules. That turns out to make a big difference. Actually, there's this interesting snippet from the release page: > These techniques have been applied to hundreds of bill
3.
▲
by
gwd
7d ago
One way would be to calculate a cost per game, factoring in both electricity and an amortized cost of the hardware, maybe having a penalty too for extra time run (e.g., if focusing only on hardware depreciation and electricity, 1 minute of
4.
▲
by
gwd
7d ago
My understanding is that AlphaZero only really existed for a year or two; there's no objective way to compare it at the moment. Leela Zero tried to open-source that work, but Stockfish incorporated a number of improvements from AlphaZe
5.
▲
by
gwd
7d ago
Two things, both from the system prompt [1]: > Your context window is limited to roughly 69000 tokens. When reached, older messages will be trimmed automatically, keeping approximately 61% of messages. Fable wasn't trained to be eff
6.
▲
by
gwd
10d ago
> For nuclear weapons it has become quite clear that even for small players, being in the race and having at least a few nukes is far more rational than having none. Ukraine found out the hard way that giving them up in exchange for prom
7.
▲
by
gwd
10d ago
> Lying to preserve a childhood myth like Santa Claus. FWIW from the very beginning, I told my son that Santa Claus, the Tooth Fairy, and the Easter Bunny were just a game we all played, and it's seemed just as fun to me. I don
8.
▲
by
gwd
13d ago
Indeed, but I actually kind of mis-spoke here. The question is less about having an experience to be aware of, but the ability to accurately reflect internal state. Even humans need to learn how to read their own internal state (e.g., sayi
9.
▲
by
gwd
14d ago
> His answer is yes, but only after it has really lived life, experienced heartbreak, and so on. The thing about this is that we're always encouraging people to read, because it gives them access to experiences and exposure to ideas
10.
▲
by
gwd
14d ago
Two comments on this, trying to take a "which hypothesis fits the evidence" approach. First, an LLM describing its own experience is not actually proof that it has any experience to be aware of, any more than an LLM confidently as
11.
▲
by
gwd
15d ago
Maybe, "Free as in free WiFi?" Like WiFi, the models you can use for free online aren't the highest quality, and can be pulled any time. The models used in TFA are halfway in between the traditional "free as in beer&quo
12.
▲
by
gwd
15d ago
> Brevity means less output tokens, which doesn’t really align with the AI vendors incentives Actually, I think Jeavon's Paradox [1] means the opposite. If doing X is $100, you may only use it to do X, but not Y, Z, or W. If doing
13.
▲
by
gwd
20d ago
> I suspect, as we continue forward, humans will slowly start to adopt the language of LLMs, or at least certain language quirks that come from interacting with LLMs. What's somewhat interesting to me is that "load-bearing"
14.
▲
by
gwd
21d ago
I think "cameo" is the correct term. He didn't wait in line all day with 1000 others to be an extra; he was asked to be in the movie because of how easily recognizable he is, and he plays himself. https://en.wikip
15.
▲
by
gwd
26d ago
This. The author talks about the fact that now they can try far more experimental optimizations than they could before. But they already had an architecture with efficiency in mind, and were trying optimizations, to begin with. The kind
16.
▲
by
gwd
26d ago
Even at the time, the example was almost certainly more representative of the concept than meant to be a specific thing that happened all the time. Then as now, people have animals; animals sometimes do bad things; when is the owner respon
17.
▲
by
gwd
26d ago
As others have said, most crimes require intent. Although I think there is a concept of "criminal negligence", I think you at least have to know you were doing something wildly dangerous. One can imagine a future where users are,
18.
▲
by
gwd
28d ago
> Since people started using Claude now I get full context of everything… And I think this is where one key argument on the page itself I think fails: > The person on the other side has the same tools you do. Yes, the person on the ot
19.
▲
by
gwd
1mo ago
The question here is not to what product management advice to the Gemini team. The discussion here is whether the watermarking is noticeable. The easy thing to do here would be to have 1000 questions, randomly assigning one half to an LLM
20.
▲
by
gwd
1mo ago
It says they observed no difference in people clicking thumbs up or down. There are loads of other behavior that they didn't observe; like, say, switching to a different LLM.
21.
▲
by
gwd
1mo ago
"DEI Officer" is literally a job title [1]. Obviously "DEI people" would be the kind of people who advocate for what DEI officers do. [1] https://www.prospects.ac.uk/job-profiles/equality-diversity-
22.
▲
by
gwd
1mo ago
> After ~30 or so commits it apparently started instructing subagents to copy the "existing verbose comment style of the codebase" - a verbose style it initiated. Yes, the "Y would make more sense, but the doc says do X...
23.
▲
by
gwd
1mo ago
Here's an actual output from Claude from a conversation about rewording a document to make it more readable: > Start with §1 (Overview) as the register-calibration piece. It's small, it's the section where the skimmability
24.
▲
by
gwd
2mo ago
Got downgraded from Fable to Opus after asking about the Great Oxidation Event [1], presumably because it started thinking about cyanobacteria, which made the monitor afraid I was trying to make a biological weapon or something. [1] https:
25.
▲
by
gwd
2mo ago
There was an interview I watched with Cory Doctorow just a few weeks ago, where someone asked him, "What kind of evidence could be presented to you to make you believe that LLMs were now ready to replace people at jobs entirely?"
26.
▲
by
gwd
2mo ago
Do the math: The number of professorial positions is not growing (or not very quickly). Under steady-state, each professor only needs to train one single tenure-track PhD in their entire career . (Or maybe 10 professors need to train 11-
27.
▲
by
gwd
2mo ago
Leveling this up: Calculate sun angles at different times of the year to simulate light entering the house. Find out if some room is going to be unbearably hot in the summer because it's got too much full sunlight; make sure there ar
28.
▲
by
gwd
2mo ago
> Exploiting multiple zero-day vulnerabilities autonomously to escape containment is pretty nuts and the first story of this kind that I've heard. But this also feels like bragging under the guise of transparency. I mean, does it ha
29.
▲
by
gwd
2mo ago
...and even though they've technically found the result through the non-intended route (breaking out of OpenAI's harness and into Huggingface's servers), they can then pretend they found the original vulnerability. Simil
30.
▲
by
gwd
2mo ago
My interpretation was that the Damore thing was the threshold where the author finally stopped identifying as a "rationalist". After that point, they considered several times writing a post describing that event. The first tim
More ›