Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
searealist
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
searealist
18d ago
Also, even the best speakers in the world can benefit from room correction, because every room behaves differently. The difference is HUGE in the lower frequencies.
2.
▲
by
searealist
22d ago
Don't forget about electricity costs, esp in California. Even with free hardware you may be better off using APIs.
3.
▲
by
searealist
27d ago
Aaron had a history of depression and suicidal ideation that long predates his prosecution. Also: https://www.unqualified-reservations.org/2013/01/noam-chomsk...
4.
▲
by
searealist
27d ago
Swartz was federally charged with wire fraud and violations of the Computer Fraud and Abuse Act based on allegedly unauthorized access, not simply prosecuted for copyright infringement or “downloading articles.” Also, he was offered a plea
5.
▲
by
searealist
1mo ago
> In my experience MTP's speed increase doesn't seem to justify the apparent loss of success at the edge, it would have to be at least 4x faster to meaningfully churn through the first 3 failures in the time it would have taken
6.
▲
by
searealist
1mo ago
Just enable MTP on llama.cpp and you will get the same decode speeds.
7.
▲
by
searealist
1mo ago
If you are not already using MTP, you should be able to get ~2x decode tokens/s with Qwen 3.8 27B.
8.
▲
by
searealist
1mo ago
I expect Pi is mostly used with OpenAI plans, and OpenAI has a dedicated compaction endpoint you should probably be using with their models instead of a compaction prompt.
9.
▲
by
searealist
1mo ago
Sous vide is fool proof, you cant mess it up. All you have to worry about is the sear, and that's easy too.
10.
▲
by
searealist
1mo ago
The internet loves using 137 for Ribeyes for the reason you said. I've had good luck going with 135 for ~4 hours.
11.
▲
by
searealist
1mo ago
What this article wants you to believe: Police are actively using flock to identify people traveling over state borders to buy marijuana and search their cars. What actually happened: Police used flock to locate someone with an active warra
12.
▲
by
searealist
1mo ago
I read that Anglo as in Anglo-Saxon, and I was confused why they would throw that in a article. It turns out Anglo is a diamond company.
13.
▲
by
searealist
2mo ago
NEM 3.0 is more fair. The grid is expensive and just as expensive if it is needed to be fully utilized 1% of the time or 100% of the time. The problem is California forcing new construction to buy overpriced solar from builders that makes l
14.
▲
by
searealist
2mo ago
There is a real tradeoff: - The musl allocator is only slow with multi-threading. - Almost all other allocators have trouble reclaiming memory when using multi-threading. This often results in multiples more RSS than single threaded or musl
15.
▲
by
searealist
2mo ago
Because it's probably still more expensive in electricity costs compared to openrouter. Certainly it's not 98% cheaper.
16.
▲
by
searealist
2mo ago
My answer doesn't change if it is M5s. Where is the math showing a 98% discount over K3 on openrouter. Heck, where is the math showing it is any % cheaper? How much electricity will your M5 sip to hit 1M input and 1M output tokens that
17.
▲
by
searealist
2mo ago
I'm not aware of any service that gives you a 98% discount and is served off of M1s. Did you do the math for what this would cost vs K3 on openrouter?
18.
▲
by
searealist
2mo ago
No one will ever derive any utility from running models at this speed. Please prove me wrong. Give me the number of tokens input and output (and dont forget about reasoning) and acceptable time to wait for it and the use case.
19.
▲
by
searealist
2mo ago
> For even the sparsest MoE open models, having more than a handful of inferences in the batch is enough to make it more likely than not that you'll get some MoE weight reuse within any given layer. This assumes totally random sampl
20.
▲
by
searealist
2mo ago
Fair point. I guess it depends on if you would be using ToU otherwise. It looks like about 50% of Californians use ToU plans, but the number is only 10% nation-wide.
21.
▲
by
searealist
2mo ago
> It's worse than that: a typical MoE model routes a separate set of experts at every layer, not just every token! But in practice, RAM offload (for systems with non-unified VRAM) and even SSD offload still work surprisingly well gi
22.
▲
by
searealist
2mo ago
It's easy to have your EV only charge off-peak, though. It's just a setting.
23.
▲
by
searealist
2mo ago
Thank god this is written in modern Zig and not Zig88.
24.
▲
by
searealist
2mo ago
Holding your cache in VRAM for 5 minutes or 3 hours have very different costs to them. "They already charge me to park my car, why can't I leave it there for a year for the same price as 1 week?"
25.
▲
by
searealist
2mo ago
They could always just extend the cache timeout beyond 5 minutes themselves. They don't do that because it is an expensive resource and there is a trade-off bewteen saving computation and reserving VRAM. Running a tool like this will f
26.
▲
by
searealist
2mo ago
1) To be clear, it costs you 10x more for uncached input tokens _for the next call_, which are still 5x cheaper than output tokens. 2) Now imagine Anthropic or OpenAI now charge your per minute of reserved VRAM time. It would be more fair i
27.
▲
by
searealist
2mo ago
How does an ASIC manage to have 60x the memory bandwidth needed to achieve that speedup?
28.
▲
by
searealist
2mo ago
The models themselves support up to 1M. You are just charged more for all context over 400k. The 272k limit is just a client limitation to ensure you can never go over 400k since the maximum output size of the model is 128k.
29.
▲
by
searealist
2mo ago
Have you ever said "I'm starving"? Do you think that undermines the experience of people actually starving in war or famine?
30.
▲
by
searealist
2mo ago
Please don't act as the hyperbole police. People exaggerate all the time (I'm starving, etc). It's normal, and you are being a jerk to call them out.
More ›