Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
MrCheeze
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
MrCheeze
2mo ago
That is in fact Anthropic's entire reason for existence, the belief that AGI is too dangerous to be controlled by OpenAI/Sam Altman. It naturally follows that it would also be too dangerous to be in the hands of literally everyone
2.
▲
by
MrCheeze
2mo ago
I mean, surely the main motivation for "use an LLM to rewrite a huge project in a new language" was excitement about the shiny new tech that made it possible.
3.
▲
by
MrCheeze
3mo ago
TBF, they did it first with ada/babbage/curie/davinci. "Sol" is a much weaker branding, though.
4.
▲
by
MrCheeze
7mo ago
Does anyone understand why LLMs have gotten so good at this? Their ability to generate accurate SVG shapes seems to greatly outshine what I would expect, given their mediocre spatial understanding in other contexts.
5.
▲
by
MrCheeze
7mo ago
The Claude Plays Pokemon stream with a minimal harness is a far more significant test of model intelligence compared to the Gemini Plays Pokemon stream (which automatically maintains a map of everything that has been seen on the current map
6.
▲
by
MrCheeze
7mo ago
Notably 45 out of the 50 days of improvement were in two specific dungeons (Silph Co and Cinnabar Mansion) where 4.5 was entirely inadequate and was looping the same mistaken ideas with only minor variation, until eventually it stumbled by
7.
▲
by
MrCheeze
7mo ago
In my experience with the models (watching Claude play Pokemon), the models are similar in intelligence, but are very different in how they approach problems: Opus 4.5 hyperfocuses on completing its original plan, far more than any older or
8.
▲
by
MrCheeze
9mo ago
This writeup on the underground puzzle is worth reading, it's a pretty baffling "puzzle" design. https://pokemow.com/Gen2/ShutterPuzzle/ That said, it's definitely Gem's fault that it stru
9.
▲
by
MrCheeze
9mo ago
There were no such writeups, 99% of the discussion about difficulties in Crystal were in twitch and discord chats where Google doesn't scrape. (It hadn't yet gotten the public attention that Claude and Gemini's runs of Pokemo
10.
▲
by
MrCheeze
9mo ago
It's hard to say for sure because Gemini 3 was only tested with this prompt. But for Gemini 2.5, which is who the prompt was originally written for, yes this does cut down on bad assumptions (a specific example: the puzzle with Farfetc
11.
▲
by
MrCheeze
11mo ago
Exactly what I was going to post. Optimizations like loop unrolling slow down the N64 because keeping the code size small is the most important factor. I think even compilers of the time got this wrong, not just modern ones.
12.
▲
by
MrCheeze
1y ago
The busy beaver function is interesting precisely because you _can't_ come up with any computable function that grows faster.
13.
▲
by
MrCheeze
1y ago
Claude almost universally reacts to everything with a positive exclamation as its first sentence, regardless of whether it's good or bad. If you don't believe me, just watch https://www.twitch.tv/claudeplayspokem
14.
▲
by
MrCheeze
1y ago
With 2, the real problem is that approximately 0% of the OpenAI employees actually believed in the mission. Pretty much every single one of them signed the letter to the board demanding that if the company's existence ever comes into
15.
▲
by
MrCheeze
2y ago
Are you thinking of Bismuth's "Speedrunning as a gateway to scientific endeavours", perhaps? https://www.youtube.com/watch?v=w8_1lQ2KH50
16.
▲
by
MrCheeze
2y ago
As an n=1 data point, that was my exact situation for a while. Also a lot of the people who put out high effort stuff are college students, which works for the same reason. More interestingly and more surprisingly, some of the people who wo
17.
▲
by
MrCheeze
2y ago
I've wondered myself why there's so little overlap between these two closely related interests of mine. Some of it seems to be the "But I don't want to cure cancer. I want to turn people into dinosaurs." effect, whe
18.
▲
by
MrCheeze
2y ago
How long until we get to the point where models know that LLMs get this wrong, and that it is an LLM, and therefore answers wrong on purpose? Has this already happened? (I doubt it has, but there ARE already cases where models know they are
19.
▲
by
MrCheeze
2y ago
In conversational language, "All my hats..." implies that the speaker has at least two hats, which theoretically means that the sentence could be a lie from them having exactly one hat (even a green one). However, in practice, I
20.
▲
by
MrCheeze
2y ago
https://xcancel.com/legit_rumors/status/1861448164084978157#...
21.
▲
by
MrCheeze
2y ago
https://www.scottaaronson.com/writings/bignumbers.html
22.
▲
by
MrCheeze
2y ago
Two reasons why I don't particularly believe him: 1) Altman's companies have had similar clauses before: https://news.ycombinator.com/item?id=40396787 2) The entire OpenAI board debacle started because Sam wanted
23.
▲
by
MrCheeze
2y ago
If, like me, you are suddenly curious what would happen if you added a small fourth body: https://youtu.be/WrahPSY9pf0
24.
▲
by
MrCheeze
3y ago
I have just been informed that my above comment is false, the CLIP-L is in fact referring to OpenAI's, despite that also being the name of an OpenCLIP model.
25.
▲
by
MrCheeze
3y ago
One of the diagrams says they're using CLIP-G/14 and CLIP-L/14, which are the names of two OpenCLIP models - meaning they're not using OpenAI's CLIP.
26.
▲
by
MrCheeze
3y ago
You should never ask an LLM to answer questions about itself. The answer is guaranteed to be hallucinated unless Google specifically finetuned it on an answer of that question. The answer it gave you is meaningless. (But also, coincidenta
27.
▲
by
MrCheeze
3y ago
So the one thing this article doesn't explicitly justify is whether the ten-byte zlib string is truly the shortest possible. You could imagine that it might be possible to hand-craft a DEFLATE block which is only three bytes instead of
28.
▲
by
MrCheeze
3y ago
It has been done - first by OpenAI (MuseNet, which is no longer available) and later by Stanford (Anticipatory Music Transformer): https://nitter.net/jwthickstun/status/1669726326956371971