Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
karmakaze
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
karmakaze
4d ago
Personally I'm using Qwen3.8-27B (MXFP4 quant W4A8) locally hosted on a pair of AMD GPUs (with DeepSeek Harness). It starts at 250 tokens/sec down to 120 past 128k context. At work mostly Opus 4.8 (sometimes a GPT or Gemini 3.1 Pr
2.
▲
by
karmakaze
6d ago
Busabase[0] > Languages: TypeScript 98.4%, Other 1.6% Not for me. [0] https://github.com/busabase/busabase
3.
▲
by
karmakaze
6d ago
Seems like a random rant to me. > What is the internet now to me? In many ways it’s something that I have to tolerate to do many things that functioned fine before. I have to use an app to pay for parking. I need to create an online acco
4.
▲
by
karmakaze
7d ago
I think it's largely due to psychology. If 210m is considered the best , then as they approach it they may start to tense up and choke. When the goal and possibility is known as 500m, then there's no point being concerned near 21
5.
▲
by
karmakaze
9d ago
Similar here peak ~250 and down to ~120 as it gets close to 128k (which is where I set DSH compaction) though it can readily do 256k. I just got DeepSeek Harness (DSH) set up with 2x R9700 and it's rather mind blowing that these can do
6.
▲
by
karmakaze
9d ago
The way it does tensor splitting without all-reduce cost over PCIe bus wasn't something I thought was possible. What kind of performance are you getting with 4x R9700s--what do you do with all the VRAM (batching, concurrent requests, e
7.
▲
by
karmakaze
9d ago
Thanks! Didn't expect to see this here. Exactly what I needed to run Qwen3.8-27B-Quark-AWQ-MXFP4-native.gguf as well as other experiments on one or 2x R9700's (I hope).
8.
▲
by
karmakaze
9d ago
> Anything design-y was “too Apple” and unacceptably bourgeois. That was definitely the case of the Unity desktop that really only worked well on netbooks. That's when I lost confidence in Ubuntu for design. Loss in Canonical on the
9.
▲
by
karmakaze
11d ago
The previous 'personal' AI Station I had pictured was the a16z one[0]. Each MI350P[1] in the TR Halo Station has 144GB VRAM and with 4.6 PFLOPs peak MXFP6 performance. Four liquid cooled? Yes please. [0] https://a16z.co
10.
▲
by
karmakaze
15d ago
NasaFTs are back in style!
11.
▲
by
karmakaze
15d ago
It seems we could use a new kind of memory that streams the weight data in, like GDDR in reverse.
12.
▲
by
karmakaze
15d ago
There's clearly a line running through the lake. It should only change name on the lower part.
13.
▲
by
karmakaze
16d ago
AirBnb can only be used if you are willing to accept the risk of having an unacceptable or even no reservation upon arrival. The latter did happen to me and fortunately I was able to make alternate plans--though still not recoup full paymen
14.
▲
by
karmakaze
23d ago
Actually no. I listed all the top stories of the day and this one made the cut. Quality is a low bar these days on HN.
15.
▲
by
karmakaze
23d ago
Why are they tying the client with the AI model instead of using OpenAI or another popular http format? Running llama.cpp locally is about the right level of complexity for most doing local AI.
16.
▲
by
karmakaze
23d ago
I used Q8 kv cache and using 64K context but can go a bit higher. The DFlash2 model takes a few gigs and using Q6_K (rather than a Q6_K_M/Q6_K_XL that unsloth publishes) saves some more. Also using Vulkan that has less VRAM overhead, b
17.
▲
by
karmakaze
23d ago
I did a similar thing running Q6_K model and Q8_0 DFlash2 (draft=7) quants: DavidAU/Qwen3.8-27B-Cold-Fusion-GAIN-V1.1-NM-DAU-NEO-MAX-MTP-GGUF incoai/Qwen3.8-27B-DFlash2-GGUF using llama.cpp PR/commit https:/
18.
▲
by
karmakaze
23d ago
> FluidVoice turns rough, rambling speech into polished, ready-to-send text in any app. Free forever, open source, and 100% on-device.
19.
▲
by
karmakaze
25d ago
Exactly, architecture for its own sake is debt on day 1.
20.
▲
by
karmakaze
1mo ago
I would prefer if it showed a regular 12 hours clock face and show the upcoming sunrise or sunset, or two faces one for each.
21.
▲
by
karmakaze
1mo ago
101 commits. First on August 7th. I don't consider this a serious project. README doesn't mention anything about how transactions work across the shards. I see this https://github.com/schapman1974/briskdb/
22.
▲
by
karmakaze
1mo ago
Funny, I published information in the hopes that humans would benefit from it. If it happens to be through collective intelligence of LLMs I'm ok with that--even more so if through open models.
23.
▲
by
karmakaze
2mo ago
I see varying comments here like use AI, don't use AI, etc and I agree with all of it. Don't always use AI, and even when you do, learn the fundamentals and how to use AI to get the output/outcome you want--that means also le
24.
▲
by
karmakaze
2mo ago
Yes programming can be an art (a la Knuth)--but artifact applications are not, unlike actual art. Exceptions include video games where the author is instilling emotion into their creation for the player to experience--not utilities.
25.
▲
by
karmakaze
2mo ago
Exactly. If one panned out there would be no more.
26.
▲
by
karmakaze
2mo ago
If FIFA was opensource now would be time to fork it.
27.
▲
by
karmakaze
2mo ago
I'll wait and see what they do about factoring all the generations of settings so there's one way to do each thing in a single uniform interface.
28.
▲
by
karmakaze
2mo ago
Or Windows NT 3.51
29.
▲
by
karmakaze
2mo ago
> Note: This article has been created in natural human language to be read easily, devoid of complex technical jargon, so all readers can enjoy. View one of our other past 70 articles for deeper technical dives. I would have preferred it
30.
▲
by
karmakaze
2mo ago
Running the Netscape or Mozilla browsers would use up 100% cpu, so it was fantastic to have a dual-cpu system so that your system was still 100% responsive.
More ›