Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
cirrusfan
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
cirrusfan
8mo ago
This model is exactly what you’d want for your resources. GPU for prompt processing, ram for model weights and context length, and it being MoE makes it fairly zippy. Q4 is decent; Q5-6 is even better, assuming you can spare the resources.
2.
▲
by
cirrusfan
8mo ago
Anthropic might have the best product for coding but good god the experience is awful. Random limits where you _know_ you shouldn’t hit them yet, the jankiness of their client, the service being down semi-frequently. Feels like the whole in
3.
▲
by
cirrusfan
8mo ago
I think it depends on what you use it for. Coding, where time is money? You probably want the Good Shit, but also want decent open weights models to keep prices sane rather than sama’s 20k/month nonsense. Something like a basic sentime
4.
▲
by
cirrusfan
8mo ago
If it sounds too good to be true…
5.
▲
by
cirrusfan
8mo ago
I find it really surprising that you’re fine with low end models for coding - I went through a lot of open-weights models, local and "local", and I consistently found the results underwhelming. The glm-4.7 was the smallest model I
6.
▲
by
cirrusfan
8mo ago
Huh? What prevents you from installing them "all at once"? The downside is obviously a long stretch of no sun, and for Europe winter being both low solar production and high energy demand due to heating which the soon-to-be-cheap
7.
▲
by
cirrusfan
8mo ago
I get a slow-but-usable ~10tk/s on kimi 2.5 2b-ish quant on a high end gaming slash low end workstation desktop (rtx 4090, 256 gb ram, ryzen 7950). Right now the price of RAM is silly but when I built it it was similar in price to a hi
8.
▲
by
cirrusfan
8mo ago
> but I have to baby sit the process and think whether I want to skip or retry a failed copy Do you import originals or do you have the "most compatible" setting turned on? I always assumed apple simply hated people that use wi