Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
Casteil
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
Casteil
21d ago
Yep.. for 'general purpose' use I found qwen3.8:27b to be disappointing due to overthinking. It's brutal especially considering how slow it is compared to MoE variants. It often overthinks to the magnitude of ~10x the tokens
2.
▲
by
Casteil
22d ago
It's a bit deceptive to state inference speeds without mentioning the additional things you're doing to achieve them
3.
▲
by
Casteil
22d ago
I don't know what black magic you're up to but I see more like 30-35t/s on a 16" M5 Max using 3.8:27b Q4, regardless of whether it's mlx or gguf. qwen3.5:122b-a10b is significantly faster at around 60-65.
4.
▲
by
Casteil
1mo ago
Now? Sure seems like it's been that way for well over a decade.
5.
▲
by
Casteil
1mo ago
Yep.. it's pretty obnoxious for real-world use with the default 'xhigh' thinking. Ridiculous amount of "Wait, actually.." which might help for complex coding tasks but makes it unbearable for general purpose use.
6.
▲
by
Casteil
1mo ago
Given that it apparently defaults to 'xhigh', this is probably the answer. Granted, it's still much lower tokens/s than you'll get out of many MoE models. Edit: Even set to medium or low there's still a lot of
7.
▲
by
Casteil
1mo ago
Yeah, that's probably the answer given that it apparently defaults to 'xhigh'.
8.
▲
by
Casteil
1mo ago
One thing a lot of people don't seem to factor when hyping Qwen is how much models like this tend to 'overthink' with seemingly endless 'second guessing'. 3.8 seems no different from what I've tried thus far. A
9.
▲
by
Casteil
1mo ago
I'm hoping too that they'll put out some MoE variants. Qwen3.5:122b:a10b can run about twice as fast as this 27b dense model. Edit: Like its predecessors, 3.8 seems really inclined to overthinking, and on a 27b dense model that&#x
10.
▲
by
Casteil
1mo ago
I just got this one a few days ago so it's kinda weird I've seen it mentioned or posted about multiple times since. Guerrilla marketing? Anyway, it's definitely better than having nothing, and tests resistance as well. I toss
11.
▲
by
Casteil
2mo ago
They're definitely doing it based on IPs. I just hopped to a different VPN endpoint and it works for me again (same browser). For now.
12.
▲
by
Casteil
2mo ago
Doesn't work for me. Seems like this login gate for old.reddit is being progressively rolled out.
13.
▲
by
Casteil
2mo ago
The difference in noise pollution comes from power density & power generation, if they're doing so on site. The power demand (and resulting waste heat) absolutely eclipses that of data centers of olde. If they're not generatin
14.
▲
by
Casteil
2mo ago
There's more to it than that. They're not just obnoxious to exist in the vicinity of, they also drive up local energy prices.
15.
▲
by
Casteil
2mo ago
>No one ever cared about DCs before now. Fewer people cared, sure.. but it's because until recently they weren't eclipsing power consumption of extremely energy-heavy industries, noticeably driving local residential energy pric
16.
▲
by
Casteil
4mo ago
You can expect around 55-60t/s with Qwen3.5:35b-a3b or gemma4:26b-a4b Q4
17.
▲
by
Casteil
4mo ago
Qwen3.5/3.6 are really prone to looping and 'overthinking'. Gemma4 doesn't seem to have the same problems.
18.
▲
by
Casteil
4mo ago
Why not 35b-a3b? ...or gemma4:26b-a4b? Both will be more capable than 9b and run at roughly similar (perhaps faster) speeds
19.
▲
by
Casteil
6mo ago
It's gotten significantly better with the advent of local/offline MoE models (e.g. qwen3.5:35b-a3b, qwen3:30b-a3b, gpt-oss:20b-3.6b), which offer a good balance of prompt response speed and output quality. 'Dense' models
20.
▲
by
Casteil
7mo ago
OVH is nearly doubling their some of their VPS pricing soon.
21.
▲
by
Casteil
9mo ago
100%. Roku's privacy policy is the most wildly invasive thing I've ever seen - basically everything that used to be just conspiracy theory.
22.
▲
by
Casteil
10mo ago
Nice. Looks like it was peaking around 21:00 - 22:00 local time, got pretty intense for a while.
23.
▲
by
Casteil
11mo ago
Yep. Manufacturers ditching Apple Carplay/Android Auto support will, if not immediately, inevitably pursue rent-seeking behavior in the form of paid subscriptions for services people could otherwise just have for free (and likely bette
24.
▲
by
Casteil
11mo ago
...they're not. This is a release of a 14" with the base M5, alongside the other existing M4 Pro/Max models. The Pro/Max rollout tends to lag behind by about 6 months.
25.
▲
by
Casteil
11mo ago
Good chance they'll introduce it with the upcoming M5 Pro/Max; the non-pro/max devices always tended to be a little lower spec all around.
26.
▲
by
Casteil
11mo ago
Pro/Max rollout tends to lag behind the 'base' by about 6 months It used to be a little less 'weird' when the base M-chips were only available in the Air and 13" MBP.
27.
▲
by
Casteil
11mo ago
Correct. Sometimes much later. Still no M4 Ultra Studio available.
28.
▲
by
Casteil
11mo ago
I wish they'd bring Space Gray back. Not a huge fan of Silver, and the 'Space Black' apparently tends to show smudges more.
29.
▲
by
Casteil
11mo ago
The non-pro/max chipped MBPs have always been a little 'lower spec' in several regards. There used to be a little more separation though, with the non-pro chips available only in the Air & 13" MBP, but back then peop
30.
▲
by
Casteil
11mo ago
Part of me misses my OG base 14" M4 Pro. The battery on that thing was absolutely phenomenal - literal 12-14+ hours of real-world use. Not so much on the 14" M1 Max (64GB) that I upgraded to after about 2 yrs. 'Real-world idl
More ›