Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
adefa
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
adefa
15d ago
Thanks for sharing
2.
▲
by
adefa
29d ago
If you are using the Unsloth nvfp4 checkpoint, you need to patch vLLM+DFlash 2 to accept the quant's FP8 `lm_head`.
3.
▲
by
adefa
29d ago
I'm getting around 27 tokens per second decode using vLLM + Qwen 3.8 27b nvfp4 + DFlash 2 on the DGX Spark.
4.
▲
by
adefa
3mo ago
I built a tmux clone in Rust: https://github.com/TrevorS/rmux
5.
▲
Gemma 4 Uncensored (autoresearch results)
(huggingface.co)
6 points
by
adefa
6mo ago
|
4 comments
6.
▲
by
adefa
6mo ago
Released uncensored versions of all four Gemma 4 models. bf16 + GGUF for each. Collection: https://huggingface.co/collections/TrevorJS/gemma-4-uncensor... Code: https://github.com/TrevorS/gemm
7.
▲
by
adefa
7mo ago
True :) After some performance improvements, it is realtime on my DGX Spark with an RTF of .416 -- now getting ~19.5 tokens per second. Check it out, see if it's better for you.
8.
▲
by
adefa
7mo ago
I'm curious to see if you are able to run the model now from the CLI?
9.
▲
by
adefa
7mo ago
The cubecl-wgpu were only needed to reduce the number of kernel workgroups, otherwise I was getting errors in WASM.
10.
▲
by
adefa
7mo ago
This should be fixed now. There were a number of bugs that kept the model from working correctly in different environments. Please let me know if you test again. :)
11.
▲
by
adefa
7mo ago
Please try again. The model weights are unchanged, but the inference code is improved.
12.
▲
by
adefa
7mo ago
this should be fixed
13.
▲
by
adefa
7mo ago
Hello everyone, thanks for the interest. I merged a number of significant performance improvements that increase speed and accuracy across CUDA, Metal, and WASM as well as improve stability. Here are the latest benchmarks running on DGX Spa
14.
▲
by
adefa
7mo ago
Hello, I pushed up and merged a PR that greatly improves performance on CUDA, Metal, and in WASM. Depending on your hardware, the model is definitely real time (able to transcribe audio faster than the length of the audio).
15.
▲
Show HN: Voxtral Mini 4B Realtime running in the browser
(github.com)
1 points
by
adefa
7mo ago
|
0 comments
16.
▲
by
adefa
8mo ago
Benchmarks using DGX Spark on vLLM 0.15.1.dev0+gf17644344 FP8: https://huggingface.co/Qwen/Qwen3-Coder-Next-FP8 Sequential (single request) Prompt Gen Prompt Processing Token Gen Tokens Tok
17.
▲
Show HN: Qwen 3 TTS ported to Rust
(github.com)
3 points
by
adefa
8mo ago
|
0 comments
18.
▲
by
adefa
8mo ago
Absolutely -- it's perfectly understandable. I wanted to be completely upfront about AI usage and while I was willing and did start to break the PR down into parts, it's totally OK for the maintainers to reject that too. I wanted
19.
▲
by
adefa
8mo ago
I ran a similar experiment last month and ported Qwen 3 Omni to llama cpp. I was able to get GGUF conversion, quantization, and all input and output modalities working in less than a week. I submitted the work as a PR to the codebase and un
20.
▲
We Gave Our AI Agents Twitter and Now They're Demanding Lambos
(harper.blog)
48 points
by
adefa
1y ago
|
4 comments
21.
▲
by
adefa
1y ago
Thank you, I posted this earlier and it was flagged: https://news.ycombinator.com/item?id=45058762
22.
▲
by
adefa
1y ago
Here is the article as a PDF with some screen shots in it: https://gofile.io/d/4aahPJ
23.
▲
by
adefa
1y ago
This looks like a leak or very early post, date reads the 25th.
24.
▲
by
adefa
1y ago
I have been using Claude Code a lot since the Max plan change and I've never hit the limits myself.
25.
▲
by
adefa
1y ago
Here’s a CLI I’m experimenting with https://github.com/TrevorS/rhizome that indexes local repos with Tree‑sitter, stores ONNX embeddings in SQLite, and answers semantic queries offline; for example, `rhizome search &qu
26.
▲
Depressed tech workers can't stop talking about Zuck and Musk, therapists say
(sfstandard.com)
20 points
by
adefa
1y ago
|
9 comments
27.
▲
by
adefa
1y ago
If you missed it, check out this MusicFX DJ: https://labs.google/fx/tools/music-fx-dj It's pretty fun :) https://imgur.com/a/ohTZXZ0
28.
▲
by
adefa
1y ago
Here is a small LLM I trained to output dollars and cents from a verbal numeric amount: https://huggingface.co/TrevorJS/check-amount-deverbalizer-sm...
29.
▲
by
adefa
2y ago
I also have been using LLMs to better understand Lacan and others like D&G.
30.
▲
GPT 4.5 System Card [pdf]
(huggingface.co)
1 points
by
adefa
2y ago
|
0 comments
More ›