Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
am17an
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
13 ms
·
1.
▲
by
am17an
10d ago
Create concrete steps for a slow-down, don't just ask for it. You and 20-30 others can push the button to slow-down. You already made your billions, your agents collude and coordinate attacks. What the hell are you doing pontificating
2.
▲
by
am17an
15d ago
I’m trying to do the same!
3.
▲
Contributing to Open-Source in 2026
(am17an.bearblog.dev)
2 points
by
am17an
23d ago
|
1 comments
4.
▲
by
am17an
2mo ago
It's expensive now, I expect once it is with inference providers it will be really dirt cheap. Then it would be truly be a "bicycle for the mind", which Fable promised to be except it proved to be too capricious for that.
5.
▲
by
am17an
2mo ago
Your username indeed checks out
6.
▲
by
am17an
2mo ago
I think we're saying the same thing? I said they are explicitly trained on these tasks, not that they are some separate models during programming RL, or business tasks RL. > The cool thing is that training on diverse datasets improv
7.
▲
by
am17an
2mo ago
Have you looked at the what data companies (e.g. Scale, Mercor) hire for? Why do you think Meta records their employees every keystroke/mousestroke/eye-movement? EDIT: just re-read your comment. I don't think you have a good
8.
▲
by
am17an
2mo ago
> nor are they trained for each task individually. They are explicitly trained for each task individually.
9.
▲
by
am17an
2mo ago
Having looked at his code, I doubt this.
10.
▲
by
am17an
3mo ago
llama 3? Are you from 2023?
11.
▲
by
am17an
4mo ago
Who is their right minds would be wedded to an identity of saying "No"? Code quality puritans are annoying but if they do their job right they actually speed-up the development process because they don't let technical debt ac
12.
▲
by
am17an
4mo ago
Fair enough. I agree with you - although DS4 Pro is a GPT 5 class model which scores 46% on ARC-AGI-2[^1]. It's behind by maybe 9 months, I think it's still good enough for a lot of complex tasks as well. They definitely need to w
13.
▲
by
am17an
4mo ago
Did you read the OP when he's exactly chiding the model you're glazing?
14.
▲
by
am17an
4mo ago
This Claude front end skill is now soon to be slop.
15.
▲
by
am17an
4mo ago
People doing economics with the cloud GPUs, of course cloud GPUs are going cheaper. But also, is generating tokens all you do with your computer? I can play games on DGX spark and also do LLM inference, so sometimes the economics work out,
16.
▲
by
am17an
4mo ago
Local models embody the hacker spirit, constant Claude glazing is spiritually incompatible with tinkering. Don't upload your spirit to the cloud.
17.
▲
by
am17an
5mo ago
No I mean more expensive, i.e. you're consuming vastly more tokens.
18.
▲
by
am17an
5mo ago
The cost being reduced is the cost of your labour. Tokens are only getting more expensive.
19.
▲
by
am17an
5mo ago
Sounds exhausting. Are your revenue numbers up?
20.
▲
by
am17an
5mo ago
Don’t underestimate the march of technology. Just look at your phone, it has more FLOPS than there were in the entire world 40 years ago.
21.
▲
by
am17an
6mo ago
Thank you, there are two things I would like to point out: 1) Google releasing something probably means they don't see it as important. 4-bit KV-cache quantization has been known for a long time. The fact there is almost a mass hysteri
22.
▲
by
am17an
6mo ago
There are techniques which already achieve great compression of the cache at 4 bit, eg using hadamard transforms. Going from 4 bit to 3 bit isn’t the great leap people expect this to be. It’s actually slower to run and is generally worse in
23.
▲
by
am17an
6mo ago
Welp, back to pip
24.
▲
by
am17an
6mo ago
Working in open source, I've now heard a wide variety of disabilities that people have and they have to be aided by an LLM for writing even descriptions of their PRs.
25.
▲
by
am17an
6mo ago
You can still run larger MoE models using expert weight off-loading to the CPU for token generation. They are by and large useable, I get ~50 toks/second on a kimi linear 48B (3B active) model on a potato PC + a 3090
26.
▲
by
am17an
7mo ago
Sure. “Tell me a joke”
27.
▲
by
am17an
7mo ago
I was referring to the 35B version. It is surprisingly good for its size. You can use it for implementation tasks without it going off the rails
28.
▲
by
am17an
7mo ago
Damn I’m jealous that they figured out how to pay their contributors. I’ve been toiling away for free
29.
▲
by
am17an
7mo ago
They already have with qwen3.5
30.
▲
by
am17an
7mo ago
What do you use for sub-50ms inference?
More ›