Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
trouve_search
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
12 ms
·
1.
▲
by
trouve_search
3d ago
Couldn't you put some sort of faraday cage around the antenna?
2.
▲
by
trouve_search
13d ago
pi.dev + anything else
3.
▲
by
trouve_search
27d ago
That's AMD's fault. RDNA4 is pretty similar to CDNA4, yet over a year after the release of "pro AI" cards like the r9700, they had basic kernels lacking in vllm (like w4a16 int4 kernels) while they were implemented in th
4.
▲
by
trouve_search
27d ago
Yes, it runs in vllm happily. It gets >900TPS output reliably on a single 5090 with the nvfp4 model. It's clearly worse than vanilla 26B-A4B, and lacks some things like structured outputs, and gets some tool calls wrong. So you have
5.
▲
by
trouve_search
29d ago
thanks for posting your setup! I think it's smart to set the reasoning effort default to something saner in the base config. Here's a VLLM command for 3.6 (I'll update to 3.8 today) to test out: ``` PYTORCH_CUDA_ALLOC_CONF=ex
6.
▲
by
trouve_search
29d ago
Yes, I mentioned the setup, but on vllm you can only use TP with speculative decoding or pipeline parallelism without, so there's tradeoff to both. I gave general numbers of what I'm getting above, the performance ratios seemed si
7.
▲
by
trouve_search
29d ago
I think it's a vllm vs llama_cpp performance thing, will pay more into it. One note I had between the two is that gemma has a much higher prefix cache hit rate in general.
8.
▲
by
trouve_search
1mo ago
What configuration are you using? On both vllm and llama-cpp, I get significantly higher speeds from gemma4 than qwen3.6 (with their respective speculative decoding methods). Output TPS in vllm for instance: - Gemma4 26B-A4B: 200-300TPS - Q
9.
▲
by
trouve_search
1mo ago
Using 98.css would still leave you with the AI slop text wording. The core problem is that some people don't even seem to notice / care.
10.
▲
by
trouve_search
1mo ago
Laguna XS is MoE, however.
11.
▲
by
trouve_search
1mo ago
Interesting, what does your tasks & workflow look like with them? I generally found the quality decent (say, similar to other competitors), but the speed of task completion was very slow. I think because they would depend too much on So
12.
▲
by
trouve_search
1mo ago
Does anyone here use Manus actively? I found it worse than alternatives in the similar space (Claude, genspark, kagi research, etc.) and slightly baffling they were being acquired at a $2B valuation in the first place.
13.
▲
by
trouve_search
1mo ago
The value prop really depends on what you're doing. If you're just vibe coding with giant frontier models, yes, the value will be worse. Especially now, where GPU prices have spiked another 20% last month. For some tasks where own
14.
▲
by
trouve_search
2mo ago
I love these, thanks!
15.
▲
by
trouve_search
2mo ago
Everyone has to find an answer to the meaning if life. For the non religious/escapist, you have to stare into the void and find an answer at some point. Existentialism (Sartre) says you have to find your own meaning. Camus says there c
16.
▲
by
trouve_search
2mo ago
The nvidia shield is pretty damn good as well, even if old at this point.
17.
▲
by
trouve_search
2mo ago
From my reading, the sandbox escape came from the JS packages in the harness still having an internet connection (somehow!), the agent having access to the source of those packages, reading it and executing code from them to access the inte
18.
▲
by
trouve_search
2mo ago
The fact that he could only get this story published in the Washington examiner of all places should be a signal that more reputable places don't want to attach their name to this
19.
▲
by
trouve_search
3mo ago
Cerebras is a whole lot of SRAM, basically a ton more L1/L2 cache, hence increasing throughput. They're pretty supply constrained right now though and their production costs seem prohibitive. The interesting players at the moment
20.
▲
by
trouve_search
3mo ago
A lot of benchmarks are setup to not punish false positives (irrelevant answers or extra text) and punish false negatives (missing the snippet being looked for). This leads to answer bloat and/or hallucination if you benchmaxx on those
21.
▲
by
trouve_search
3mo ago
gemma 12B 4bit quant; try something with MTP and an AWQ quant
22.
▲
by
trouve_search
3mo ago
On a 5090, gemma4 26B runs at 350TPS with the command below [1] and gemma4 31B is around 150TPS with a similar command. I'm really surprised how much slower a DGX spark is for the same price. 1. Here's my command. PYTORCH_CUDA_ALL
23.
▲
by
trouve_search
4mo ago
Cerebras are only serving kimi for dedicated endpoint customers; for that you need a >$5m annual deal with them Cerebras also seems to be killing off their regular APIs, they're deprecating models and GLM is still stuck on GLM 4.7,
24.
▲
by
trouve_search
4mo ago
Not sure if the M5 is that massively different but I have a M2 max laptop and the screen is noticeably brighter on the Asus.
25.
▲
by
trouve_search
4mo ago
It hasn't been so bad for me to notice. Compiling rust you'll hear the fan, but you'll also hear it in a MBP. The MBP will compile faster however.
26.
▲
by
trouve_search
4mo ago
The moat is the difference between knowledge and know-how. You can read all the plumbing books, but you need to get your hands dirty a few times, mess it up and fix it, to get mentally comfortable and efficient with the work
27.
▲
by
trouve_search
4mo ago
I've been using an asus zenbook 14 OLED with linux. Compatibility is great. The screen blows apple out completely. It's clearly, obviously better. The fan noise and battery life are worse than Apple. The keyboard feels better to t
28.
▲
by
trouve_search
4mo ago
OK, I'm 100% rooting for both Mistral and task focused small models. But Mistral has fall really far behind since 2025Q3. It seems they can't get good reasoning models working at even medium context sizes, which is necessary to be