Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
WhitneyLand
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
WhitneyLand
7d ago
By that logic we should also consider the case of cutting the number of layers in half because that would also reduce hidden state between token generation. In your generalized example I think the concern is when the additional evaluation e
2.
▲
by
WhitneyLand
7d ago
No. It’s not at all by definition hidden reasoning. Looping transformers uses additional calculations (repeating layers) to generate a token. Reasoning (in this context) is test time generation of multiple tokens that allow a model to have
3.
▲
by
WhitneyLand
12d ago
I don’t know that it’s that shocking, remember Go it’s not solved game, so the the limits of what’s really possible is not known in all cases. For example, we don’t even know whether perfect White play can possibly overcome two correctly pl
4.
▲
by
WhitneyLand
12d ago
If you’re wondering how they wrote to the wiki having only GET ability… Basically it was a bug in the wiki code. They transferred the POST form parameters to GET URL parameters, and wiki internally doesn’t distinguish between the two.
5.
▲
by
WhitneyLand
14d ago
There are important gaps in that hot take. For example, it's not even close to Opus 5 on Terminal-bench 4.0, 19.1% vs. 51.8%.
6.
▲
by
WhitneyLand
18d ago
Why do people say document database when they really just mean json database?
7.
▲
by
WhitneyLand
19d ago
“It's no where near those numbers” “get inflated in anecdotes” Just because one study has the number 34% in it, does not prove the numbers I gave were wrong and does not mean they are anecdotal. This is something that’s been studied fr
8.
▲
by
WhitneyLand
19d ago
So, this is not supported by the data when people are asked. Sexual effects alone hit about 50–70% people. Then you could face nausea, insomnia, profuse sweating when you’re still, emotional blunting, and weight gain. I don’t want to discou
9.
▲
by
WhitneyLand
21d ago
That’s only true regarding the one sentence about the singularity not being a point. People like to reduce papers to a simple hot take, but the paper is more than that, offers viewpoints that are non-standard and speculation about new possi
10.
▲
by
WhitneyLand
1mo ago
Why would you bring up fraud in the midst of science research in the US being burned to the ground? Fraud is not the reason it is happening. Even the people who are making the cuts in this case have clarified the purported reason and it has
11.
▲
by
WhitneyLand
1mo ago
The DeepSeek team is so strong, very impressive. Imagine if they had GPU resources of western labs.
12.
▲
by
WhitneyLand
1mo ago
They chose to compare against Open AI’s mid tier model Terra instead of Sol and still lost some benchmark against it. They left Opus in and got beat in all but one benchmark. Nothing wrong with trying to improve, but why the marketing games
13.
▲
by
WhitneyLand
1mo ago
Nowhere in the paper do they mention the reasoning level or budget used for the experiments? You’ve got to be kidding me. That one variable could make a huge difference in the results. I can’t understand why they would leave that out.
14.
▲
by
WhitneyLand
1mo ago
Another headline of “model runs on x”, which usually means “let’s list how much you give up to run on x”. Dumbed down quantization? No. Full intended inference weights preserved, so far so good. Slow performance? No again. Looks like you co
15.
▲
by
WhitneyLand
1mo ago
How do you figure that? When they just loaded the weights alone, it was taking 156GB in vLLM. After warm-up and adding a KV cache pool, it took over 200GB. And this implementation is already cutting down the 1M token context window you woul
16.
▲
by
WhitneyLand
2mo ago
If speed is a metric for you, tokens required to solve a problem affects that metric. All else being equal passing triple the amount of tokens through a model to solve the same problem makes it slower. Doesn’t mean this model is bad, and it
17.
▲
by
WhitneyLand
2mo ago
It’s not outdated at all to use tokens to estimate performance, it’s directly related.
18.
▲
by
WhitneyLand
2mo ago
It’s exciting that a model scoring this high is dirt cheap. It’s also so inefficient, when they release the full performance numbers it’s not going to be good. One example, it takes about 3.6x more tokens to finish the same work as Gemini F
19.
▲
by
WhitneyLand
2mo ago
False dichotomy right? Are Chinese labs impressively innovating? Clearly. However this doesn’t rule out possible gains due to distillation. I don’t know the degree of the latter but both things could certainly be true.
20.
▲
by
WhitneyLand
2mo ago
The problem is what you’re skeptical about, the true cost, is probably the least important part. Did it really cost $1 million instead of the 150k that’s been floating around? If you don’t like the price now just give it some time. The poin
21.
▲
by
WhitneyLand
2mo ago
Cool worlds says no? https://youtube.com/shorts/qreL7htXp98?is=G9zJVbpMT0yZsLQ-
22.
▲
by
WhitneyLand
2mo ago
The silence is deafening. Google watches over the last few months a flat out assault on the Pareto curve from American and Chinese companies. Release after release pushing the boundaries of frontier intelligence and price/performance.
23.
▲
by
WhitneyLand
2mo ago
The world doesn’t need advisors like Moh that are bitter enough to gossip about a family funeral.
24.
▲
by
WhitneyLand
2mo ago
Not convinced just using top n sigma is going to beat a SOTA detector. The fingerprints Pangram uses should in principle be able to detect style above the token selection level. Other papers have tried to beat it with temperature and it did
25.
▲
by
WhitneyLand
2mo ago
I’m personally interested in it as part of the research to improve LLM writing. Detecting “AI voice” is part of understanding what’s wrong with it in the first place and how to improve it. But yeah, in general I think you’re right, the actu
26.
▲
by
WhitneyLand
2mo ago
I’m not sure what you’re saying I’m wrong about. The comment I was responding to, about changing LLM output, referred to prompting, not temp/sampling tricks. I’m not aware of Pangram being beat by clever prompting. There’s some interes
27.
▲
by
WhitneyLand
2mo ago
That’s a different point. I’d want detectors to be as accurate as possible, false positives of 1 in 10000 seems like a good starting point. I believe their results have been independently tested. And as a separate matter, any tool for evalu
28.
▲
by
WhitneyLand
2mo ago
"It's really easy to have a false positive" Not really. The false positives for the SOTA detector are very very low. "It's also very easy to change the pattern of LLM output." Not in a way that can reliabl
29.
▲
by
WhitneyLand
2mo ago
"Text is simply not information dense enough to be able to decode some arbitrary signal of provenance from it...it's a bad fiction to perpetuate that any of this is anything more than tarot card reading." Not true at all. P
30.
▲
by
WhitneyLand
2mo ago
I find it hard to see a language as beautiful that’s grown too complex for a single person to hold a complete mental model of. I used to think that was a personal limitation, until I saw an interview with Bjarne explaining that he used to u
More ›