Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
petu
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
petu
3d ago
Sipeed has such device https://wiki.sipeed.com/hardware/en/kvm/NanoKVM_USB/introduc...
2.
▲
by
petu
6d ago
This feature is Pro phones only, not Duo: https://www.apple.com/iphone/compare/ ("Apple Reference Image (Fusion Main)") So only on devices with LiDAR / that can capture depth map.
3.
▲
by
petu
6d ago
Miniature dioramas wouldn't be size appropriate. Apple could detect faces/cars/other common objects of ~known size and verify -- or even just dump depth map for anyone to check.
4.
▲
by
petu
6d ago
It seems to be what Apple is doing, this feature is only available on the 18 Pro's (which have depth sensor on the back), but not Duo.
5.
▲
by
petu
6d ago
Aren't they meant for underfloor heating? Heat pump efficiency drops with delta T increase, but iron radiators need 60-80C supply.
6.
▲
by
petu
6d ago
You're looking at third party providers. V4 Flash prices served by DeepSeek themselves: launch pricing: $0.0028 / $0.14 / $0.28 after Aug 16th: $0.007 / $0.22 / $0.66 during off-peak. after Sep 10th: $0.00
7.
▲
by
petu
7d ago
It's larger than previous V4 Flash. 552B in ~FP4, 306GB. 196B of FP8 Engrams, another 204GB, not necessary to keep in RAM. KV cache sees another 4x size reduction, just 900MB for 1M. So 384GB needed for a chance of ac
8.
▲
by
petu
7d ago
V4 Flash also was released as mostly FP4, but this one is FP8 (?). 160GB vs 510GB. Original Flash good fit for dual Spark / Strix Halo machines. This one would require third party quants and even then 4 machines. Edit: Most of added w
9.
▲
by
petu
7d ago
> Hoje em sites como openrouter o valor é de $0.16 output . You're comparing different providers then. DeepSeek price on OpenRouter is $0.66 output.
10.
▲
by
petu
7d ago
A month ago new V4 Flash 0731 checkpoint was better than existing V4 Pro. They've kept serving Pro, it was updated 13 days later (0813 checkpoint). Now, 4 weeks later new Flash checkpoint (0910?) is again better than existing Pro. Same
11.
▲
by
petu
7d ago
4.1 releases tomorrow, right now you're supposed to be served by same old model
12.
▲
by
petu
13d ago
OpenAI provides API key with ~unlimited use?
13.
▲
by
petu
15d ago
few weeks ago
14.
▲
by
petu
15d ago
Do you think average human cares about "our species is going extinct" stuff? If not, why couldn't they be happy at the same time? You seem to equate birth rate with happiness, but ignore efficacy and availability of modern co
15.
▲
by
petu
17d ago
Before we worry about source code, Microsoft doesn't grant me rights to modify/redistribute/sell copy of Windows I have.
16.
▲
by
petu
18d ago
Yes, but running out of RAM is impractical due to low memory bandwidth. According to the article/Samsung RAM dies inside can support way higher bandwidth than they expose, they're limited by external interface / bus width: &g
17.
▲
by
petu
19d ago
This time they just made FP8 "default", accompanied by "-BF16" model/page (previously "-FP8" was released alongside).
18.
▲
by
petu
19d ago
I assume that's about 5.3 Flash, not full?
19.
▲
by
petu
20d ago
Air temp is measured in the shade, sun hitting windows/interior floors would take you past that. IR blocking film/tint on outside works great if you can't get windows shaded.
20.
▲
by
petu
21d ago
You need VRAM for the whole thing for optimal performance. Activation is chosen "randomly" for each token. PCIe becomes bottleneck, so much that just doing computation on CPU is likely faster. But given it's only 6B, out of w
21.
▲
by
petu
21d ago
It's not 1 bit. It's ~4bit for n-gram and ~2.8bit for the model. Not idea why it's called Q1, but likely it's preliminary quant just for PR testing / very likely to be remade after llama.cpp support is merged.
22.
▲
by
petu
21d ago
This is needs ~80GB of fast memory at 4 bits per weight. Faster memory is better, but probably even something like 3090 + 64GB RAM should work (not fast, but maybe even 20-30 t/s? llama.cpp support pending).
23.
▲
by
petu
21d ago
Haven't tried, would be surprised if it's any different. It's new arch demo for future Qwen 4 family, but (as I understand) training recipe/data is same as any other 3.8 model.
24.
▲
by
petu
22d ago
Mac Mini M4 launched less than 2 years ago at $600 (US, $500 with edu discount; as low as $400 on general discounts). Now same config with M6 is $900.
25.
▲
by
petu
22d ago
It was in description under the countdown initially, but was quickly removed. It also said 51B of n-grams and new attention (IIRC it said "Qwen Sparse Attention"). edit: here's a random screenshot https://x.com
26.
▲
by
petu
22d ago
> This project evaluates local language models running on a single NVIDIA DGX Spark. "Did much better" is a bit misleading w/o that context and 1 hour time limit -- your benchmark design heavily favors V4 Flash. From resul
27.
▲
by
petu
24d ago
What you've linked is very much official -- that's original Qwen 3.0 release, so pretty old, but official.
28.
▲
by
petu
24d ago
Maybe some were fixed, but: 1) Shipping with 2k default context window for the longest time, w/o any warning and being not easy to change (like any other setting). Totally made a lot of people think local LLMs are dumb as rocks. Just c
29.
▲
by
petu
24d ago
There's no official Qwen 3.8 4B (only 27B and 2.4T.. at least for now), so if not a typo you've downloaded some third party model/finetune. Also if you have less than 24GB VRAM, then ollama defaults to 4K context. If that &qu
30.
▲
by
petu
28d ago
They show that CS-4 can't really do batching (or rather it can't properly benefit from it), total throughput barely changes (25%?): https://cdn.sanity.io/images/e4qjo92p/production/6a132331880... Wh
More ›