Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
qeternity
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
qeternity
7d ago
Most people are not even referring to CUDA batching nuances. They think that sampling is an inherent part of Transformers. Even on this site, it is regurgitated with confidence.
2.
▲
by
qeternity
12d ago
It’s only good for them in the sense that it is allowing them additional time and compute to cheat a reward signal. It is not good for them in the sense that they will short circuit the RL path that actually improves general capability.
3.
▲
by
qeternity
14d ago
Has nothing to do with being headless. That's just a natural outcome. If someone wanted to use an 8x B300 as their daily driver...go ahead. It would still be the best way to serve a given model.
4.
▲
by
qeternity
14d ago
*TB/s
5.
▲
by
qeternity
15d ago
https://typebulb.com/u/lab/you-re-relatively-right/full
6.
▲
by
qeternity
16d ago
> Edit: I'm thinking of a headless Mac mini, if you meant running it on the same machine you're using of course you'll need more memory, but LLMs are best served from a headless server so that's what I'd recommen
7.
▲
by
qeternity
16d ago
You're not entitled to this stuff any more than they are. What planet do you live on?
8.
▲
by
qeternity
21d ago
Trolling. GLM is heavily distilled from Gemini.
9.
▲
by
qeternity
24d ago
They are talking about breadth, not depth, multipliers in almost every circumstance. A great X will be able to do far more great X stuff (breadth), and perhaps also be a greater X (depth). But it's most certainly weighted in favor of t
10.
▲
by
qeternity
24d ago
Is it not a bad thing to encourage poor people to continue doing the thing that has apparently made them poor? It's also very not true that farmers are poor in the US (median household income about 30% greater than general population).
11.
▲
by
qeternity
26d ago
College stopped being about getting an education a long time ago. It's just a low pass filter for the job market. Students aren't actually interested in learning anything, they have been condition for a long time to jump through t
12.
▲
by
qeternity
26d ago
It's the other way around: conviction rate includes plea bargains. Convictions at trial are much lower.
13.
▲
by
qeternity
26d ago
You are misunderstanding the conviction rate. Only ~2% of cases actually go to trial, and at trial there is ~80% conviction rate. There are ~4x as many cases dismissed by judges before getting to trial. And the vast majority (90%) of defend
14.
▲
by
qeternity
26d ago
I presume you're a Windows user then. When the M1 was released, I couldn't believe how fast it was.
15.
▲
by
qeternity
27d ago
Ultimately the frontier labs are competitors of Nvidia. There is a fixed amount that the market will pay for tokens. If the frontier lab model premium collapses due open weights models, Nvidia can capture a greater share of aggregate spend.
16.
▲
by
qeternity
28d ago
It is still people, not money, who votes for them over and over again…
17.
▲
by
qeternity
1mo ago
Personal hardware is only sufficient today for some tasks. As data center power efficiency, and large sparse MoE task efficiency increase, personal computing will continue to lose out.
18.
▲
by
qeternity
1mo ago
Computer use will blow through tokens because it's doing image capture for everything. You may have better and more reproducible results using browser controls that aren't image based, or writing tools that completely sidestep bro
19.
▲
by
qeternity
1mo ago
Jevons paradox: large purpose-fit data centers increase efficiency such that you can use AI in more places, and use more tokens for those tasks. The future is not a single chat bot session of bs=1. The future is many agents performing many
20.
▲
by
qeternity
1mo ago
The real breakthrough is going to be thinking in latent space.
21.
▲
by
qeternity
1mo ago
You say this like that isn’t how the vast majority of problems are solved…
22.
▲
by
qeternity
1mo ago
Has nothing to do with Chinese. Frontier labs have already been doing this for a while, verified in smuggled traces from OAT/Ant. Simply a way to reduce tokens.
23.
▲
by
qeternity
1mo ago
35A3 might be more comparable to 10 dense. 27 dense is far more capable than 35A3.
24.
▲
by
qeternity
1mo ago
> more than half its price Less than half its price. More than 50% discount.
25.
▲
by
qeternity
1mo ago
LLMs are, in theory, deterministic. Sampling is not intrinsic to LLMs. Greedy decoding a single batch in most libraries will give you mostly deterministic outputs. Higher batch sizes can increase variance. But all of this is down to CUDA an
26.
▲
by
qeternity
2mo ago
I don't follow this. Clearly all of the frontier labs are doing these things. When OAI released gpt-oss it was released as an mxfp4 checkpoint. OAI, Ant, et al are also obviously employing QAT.
27.
▲
by
qeternity
2mo ago
I cannot be sure what the likes of Cursor have done, but I think it's incredibly unlikely that they have trained a QLoRA for Composer. It's almost certainly full parameter post training of the original model weights.
28.
▲
by
qeternity
2mo ago
I think it’s fairly obvious the Chinese labs are doing mass distillation. I also think the Fable accusation is wrong and it was most likely Opus 4.8 which itself is likely a distillation of Fable.
29.
▲
by
qeternity
2mo ago
But these aren't the right services where the test should be, right? There's another service that says "ok we take the 100 bytes from A, and we take the $17 SKU from B, and this should equal $X". It's the third serv
30.
▲
by
qeternity
2mo ago
It's a defensive tactic to reduce the effectiveness of distillation. Say of that what you will, but it's not because they want to wrest control from users. It's because they don't want Chinese companies to do exactly wha
More ›