Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
apitman
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
31.
▲
by
apitman
1mo ago
Aren't things like KV size inherent to the model?
32.
▲
by
apitman
1mo ago
I wonder how this would stack up against 4x RTX 3060, assuming you have the physical room for them.
33.
▲
by
apitman
1mo ago
Interesting. Why don't the unsloth guides ( https://unsloth.ai/docs/models/qwen3.8 ) mention this? Do they already include the fixes in their GGUFs?
34.
▲
by
apitman
1mo ago
Check out OpenCode Go as well. They give some Kimi K3, Qwen3.8 Max, and GLM5.2 (probably 5.3 soon?) usage which may cover your needs for $10/mo
35.
▲
by
apitman
1mo ago
Yeah make sure you're using MTP and potentially tensor parallelism.
36.
▲
by
apitman
1mo ago
This is specifically about unsloth having day-1 GGUF's available.
37.
▲
by
apitman
1mo ago
I used GPT-5.6 Sol high to optimize it, and it claimed it was getting 50. I'm seeing ~40 on my goto smoketest: "Make me a vector add in CUDA". Funny side note. It successfully one shot the program, but it wasn't able to
38.
▲
by
apitman
1mo ago
Running it on 2x3060 now. Works pretty well but VRAM is tight . 4bit quants. 1x128k context, 8bit KV, MTP on.
39.
▲
by
apitman
1mo ago
The Luna price dropped the day before a massive update to Flash (0731 update). GP may be referring to that version.
40.
▲
Unsloth Qwen3.8-27B GGUF files
(huggingface.co)
69 points
by
apitman
1mo ago
|
5 comments
41.
▲
Unsloth - Qwen3.8 - How to Run Locally
(unsloth.ai)
7 points
by
apitman
1mo ago
|
0 comments
42.
▲
by
apitman
1mo ago
Look at OpenRouter. They have lots of providers
43.
▲
by
apitman
1mo ago
But why? Luna Max is almost the same intelligence as Terra xhigh and way way cheaper. And Terra max is almost the same as Sol high. I just don't really see a place for Terra but slower.
44.
▲
by
apitman
1mo ago
Wait people use terra?
45.
▲
by
apitman
1mo ago
Are there any projects that track how much usage of each model translates to how much percentage drop in weekly/5hr windows?
46.
▲
by
apitman
1mo ago
https://artificialanalysis.ai/models/grok-4-6
47.
▲
Qwen3.8 Weights Released
(modelscope.cn)
31 points
by
apitman
1mo ago
|
4 comments
48.
▲
Qwen3.8-27B Countdown
(modelscope.cn)
3 points
by
apitman
1mo ago
|
0 comments
49.
▲
by
apitman
1mo ago
This is what I do. Nice benefit is it lets me passthrough my GPU and share it between multiple containers. Incus is awesome.
50.
▲
by
apitman
1mo ago
Fun fact: QEMU runs natively on windows and supports acceleration with WHP. It works surprisingly well.
51.
▲
by
apitman
1mo ago
> OpenCode Go currently offers $120 for $10 on DeepSeek Flash v4 At DeepSeek's absurdly low rates or market rates?
52.
▲
by
apitman
1mo ago
I'm getting like 25 tok/s on 2x RTX Pro 6000. This is with llama.cpp, but I had GPT tune it for me. I was under the impression vLLM was at most ~2x faster, and usually for highly parallel loads. Any tips on where I should look fir
53.
▲
by
apitman
1mo ago
These are very interesting results, and honestly hard to believe, even as a big 0731 fan. If I'm reading the chart correctly, a couple observations: * deepseek-v4-flash-0731 max is better than kimi-k3 max * glm-5.2 is dumber than a box
54.
▲
by
apitman
1mo ago
I've found it to be pretty good so far.
55.
▲
by
apitman
1mo ago
DeepSeek has far cheaper cache pricing. That's the difference.
56.
▲
by
apitman
1mo ago
These numbers look about right based on my experiences as well. Though for a single user I think 2x DGX Spark (~$10k) runs DSv4 Flash fairly well right?
57.
▲
by
apitman
1mo ago
Welp. That didn't last long
58.
▲
by
apitman
1mo ago
As low as it is, switching between providers on OpenRouter is still lower. That said, it's a fair point. For me, it boils down to things covered here: https://earendil.com/posts/session-portability/ Things li
59.
▲
by
apitman
1mo ago
Looks like coding agent is model+harness. There are far fewer models represented on that page. I believe "agentic index" is still the metric to look at for coding performance. I could be wrong about that though.
60.
▲
by
apitman
1mo ago
For one thing, providers of open models can't arbitrarily increase their prices without facing competition.
More ›