Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
versteegen
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
versteegen
15d ago
Ugh, people are still saying the Codex limits are more generous. They're not, Claude's are over 2x higher, have been for months! [1] It's just that Claude uses far more tokens, 2-3x is common. Except sometimes GPT will use ju
2.
▲
by
versteegen
24d ago
> Ok the rear red turn signal lights were already a trend in the US before Tesla. But please reverse this too. I'm surprised. But it shouldn't be up to Tesla; it's illegal and they never get on the road in countries with p
3.
▲
by
versteegen
1mo ago
I urge you to reconsider your beliefs. You are missing something important because you are thinking in terms of low-dimensional statistics. Deep learning doesn't just fit data, it finds features (abstractions) of the data.
4.
▲
by
versteegen
1mo ago
IME using 5.6 Luna and DS V4 Flash, I notice that although they are excellent at programming, even Opus-like in the way they try to debug, the thing they are worst at is inferring user intent and making good decisions with little informatio
5.
▲
by
versteegen
1mo ago
I think this is the best and most useful way to measure model intelligence. In my experience it's what really sets apart the capable models from the best. A small model can be RL trained to be extremely good at programming or narrow pr
6.
▲
by
versteegen
1mo ago
If doesn't correspond cleanly. I can see why you draw the link, because LZ compression will replace words with symbols but BPE is a non-contextual entropy encoding while LZ is contextual and adaptive and that makes it very different. I
7.
▲
by
versteegen
1mo ago
That conception of knowledge is interesting, but I think using the label 'knowledge' for it is very problematic, it's too far from common definitions. The fact that you have to carve out an exception for mathematics already s
8.
▲
by
versteegen
1mo ago
There is a distinction between a compressor for a fixed dataset and one for an unknown population from which we have a sample. The optimal compressor for the sample may be the single best guess for the population, but that's not what S
9.
▲
by
versteegen
1mo ago
> For me it’s the massive amount of resources it takes to produce and run one It's amazing that LLM pretraining is both extremely data inefficient at learning concepts and cognitive functions from the training data compared to human
10.
▲
by
versteegen
1mo ago
You misread. "Pain and suffering" not "death". Of all the pain and suffering in the world, a vast amount of it really is our own fault. Famines and wars shouldn't happen. And if you see a country border with poverty
11.
▲
by
versteegen
2mo ago
To save anyone else the trouble: discussion there is not really worth looking at (largely a flame war), except: the author of this disproof seems to be a crank, and the disproof's been refuted.
12.
▲
by
versteegen
2mo ago
...but we're talking about compaction, and opencode's compaction is (or was) terrible. I've seen so many horrible problems that I keep it disabled (with an envvar flag, because even the config flag to turn it off was broken).
13.
▲
by
versteegen
2mo ago
Wow, remarkable. Clearly Aum is a different league from the lone-wolf "Fort Detrick guy", treat the risks separately. But I'll take these questions as rhetorical. (See my reply to the sibling comment.) I can't answer the
14.
▲
by
versteegen
2mo ago
I didn't argue "must be regulated". I'm arguing AI is powerful (at achieving things, and hence has dual-use dangers). Many people won't even admit that, which is the part that really annoys me: they don't even
15.
▲
by
versteegen
2mo ago
Thank you for taking this seriously enough to write this, and anyone else likewise. But this is attacking a strawman, amateur bioterrorists. AI is a force multiplier in the hands of an expert. If it took a team before, maybe a single malici
16.
▲
by
versteegen
2mo ago
The old observation that people tend to define AI as whatever computers can't do yet is as true as ever. It's getting a bit absurd, moving from demanding "general intelligence" to replicating human cognitive phenomenolog
17.
▲
by
versteegen
2mo ago
Isn't it ~$3000 per week? Extrapolating from the current limit on Pro plans.
18.
▲
by
versteegen
2mo ago
Disagree. Actually, in API cost equivalent, the $20/mo ChatGPT Plus plan gives you ~$100 of usage, while $20/mo Claude Pro gives you >$250 of usage (I measure at ~$300 in my last week), though that is currently +50% for the nex
19.
▲
by
versteegen
2mo ago
> It looks like it wrote a python script to generate test cases in our file format for testing. Just... you know, as a side quest. On the one hand, agents have done this sort of thing for a year+, if you pushed them to check their work.
20.
▲
by
versteegen
2mo ago
Algebra is useful because graphs are algebraic objects, and a lot of CS is about graphs, in particular search/planning. But no, I never saw rings mentioned except for generating functions, which are used for analysing recurrence relati
21.
▲
by
versteegen
2mo ago
Having studied CS and maths to post-grad, I think you exaggerate. Although a CS course might use these tools, they didn't in my experience go into explaining or defining them. The only use of linear algebra I can remember was in analys
22.
▲
by
versteegen
2mo ago
That was a tendency of 5.4 and earlier, OpenAI specifically worked to avoid it in 5.5 and I find it happens rarely know. It really felt like 5.4 had been intentionally trained to stop and check, I believe it wasn't the system prompt.
23.
▲
by
versteegen
2mo ago
"Thinking" seems to be a political term now, people have completely different definitions of it, based on how they wish the world to be organised, and find defining it differently offensive.
24.
▲
by
versteegen
2mo ago
Being correct doesn't give you licence to use insults.
25.
▲
by
versteegen
2mo ago
Not half, only on the order of 10% of the cost of inference is electricity.
26.
▲
by
versteegen
2mo ago
Pretty sure there isn't. HN has lots of hellbanned bots lately (I have showdead on [1]) but I never notice ones pushing a political slant. But I suppose nation states conducting a campaign on Hacker News would be much too cunning to m
27.
▲
by
versteegen
2mo ago
I always run opencode with envvar `OPENCODE_DISABLE_AUTOCOMPACT=1` because I discover a new horrible bug in its autocompact every time I don't... including the config .json option to turn off autocompact not working. Sad but not surpri
28.
▲
by
versteegen
2mo ago
BTW the quotas for Go have very recently changed, now only $15 for some models instead of $60. Which is not actually a difference for DS4 Pro, because they lowered the token pricing 4x at the same time (to match the change in official prici
29.
▲
by
versteegen
2mo ago
To summarise the full results table further down the page (which doesn't render on the page for me!): Kimi K3 beats each model (out of 35 benchmarks, excluding missing): vs Fable 5 : 12/35 (34%) (ties: 1) vs GPT
30.
▲
by
versteegen
2mo ago
It's a custom 3 bit quant of Llama 3.1 8B and other shortcuts. The quant is not good. Their newer arch switches to standard 4 bit quants, should be far better!
More ›