Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
coder68
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
coder68
4mo ago
This seems pretty useful for AI inference if it can pass Apple approval. I've wanted to use my Nvidia GPUs with a Mac Mini, this would enable it to run CUDA directly. Very cool!
2.
▲
by
coder68
5mo ago
What would be an example of a positive signal?
3.
▲
by
coder68
5mo ago
In fact it is appreciated that Qwen is comparing to a peer. I myself and several eng I know are trying GLM. It's legit. Definitely not the same as Codex or Opus, but cheaper and "good enough". I basically ask GLM to solve a p
4.
▲
by
coder68
5mo ago
AI is much better at front-end than me, it has really enabled me to build visual apps as a normally backend/ML guy.
5.
▲
by
coder68
5mo ago
A bit more context would be helpful, as someone happy to donate -- what is the current situation, why the urgency? Just some more info would be good.
6.
▲
by
coder68
5mo ago
Hmm you might be able to tweak the settings further. Under llama.cpp on one RTX 6000 Pro I get ~215 tok/s generation speed. The key for me was setting min_p greater than 0. My settings: ``` #!/bin/bash llama-server \ -hf ggm
7.
▲
by
coder68
5mo ago
I have not delved into the theory yet but it seems that the smaller open-source models do this already to an extent. They have less parameters, but spend much more time/tokens reasoning, as a way to close the performance gap. If you lo
8.
▲
by
coder68
6mo ago
20 years is quite an optimistic timeline. Of course, we will use agents to solve the problems of agents!
9.
▲
by
coder68
6mo ago
I gave it a whirl but was unenthused. I'll try it again, but so far have not really enjoyed any of the nvidia models, though they are best in class for execution speed.
10.
▲
by
coder68
6mo ago
Are there plans to release a QAT model? Similar to what was done for Gemma 3. That would be nice to see!
11.
▲
by
coder68
6mo ago
120B would be great to have if you have it stashed away somewhere. GPT-OSS-120B still stands as one of the best (and fastest) open-weights models out there. A direct competitor in the same size range would be awesome. The closest recent rel
12.
▲
by
coder68
6mo ago
The good news is local models have significantly improved. If it all goes down today, you can still run e.g. Qwen 3.5 at home, and it's "good enough" for most workloads. With a gaming GPU you can run Qwen3.5-35B-A3B. I use 12
13.
▲
by
coder68
6mo ago
We all probably need to touch grass a bit. Our industry is really out of touch with reality right now, although the looming impact of AI is probably quite real.
14.
▲
by
coder68
6mo ago
Thanks! Super interested in LLMs for translation :D glad to see you folks doing this work.
15.
▲
by
coder68
6mo ago
It does? There is a fast drop followed by a long decay, exponential in fact. The cooling rate is proportional to the temperature difference, so the drop is sharpest at the very beginning when the object is hottest.
16.
▲
by
coder68
6mo ago
Is there interest in benchmarking the proprietary LLMs for translation? Curious as I often use Gemini 3 Flash, but I have no idea how good it is for my language family. I prefer open models (in fact the smaller the better for offline), but
17.
▲
by
coder68
7mo ago
Even working in "tech" but not FAANG this is so true, 10 days is still the norm at many white collar businesses for your first year of employment, sometimes 15 days if they're generous.
18.
▲
by
coder68
7mo ago
The tradeoff with many EU countries would be that they enjoy their leisure time a lot more and sooner than Americans. Americans make more and save more statistically, but they spend it on cars, houses, and medical care, and generally have w
19.
▲
by
coder68
9mo ago
As someone studying Polish, and making excellent progress, I mostly agree with your take. If you want to explore other languages, something like Spanish will get you much more mileage. Polish is difficult and the community of speakers isn&#
20.
▲
by
coder68
1y ago
To chime in about where I'm at -- one problem was solved with a statistical classifier, but to bootstrap another, I ended up using keywords. It took a few hours to get a reasonable solution, and it leans more towards precision than rec
21.
▲
by
coder68
1y ago
To some degree manual labeling has to be done anyway, just to validate that any approach works at all, you'll always need ground truth from somewhere. What I suggested is that zero/few-shotting might not be good enough, depending
22.
▲
by
coder68
1y ago
oh this seems like an interesting idea, what tactics do you use for augmentation? For my own use-case, I think I could reorder semantic chunks, or maybe randomly delete pieces, but curious what tactics you use! I have also considered traini
23.
▲
by
coder68
1y ago
The outputs are working correctly in terms of formatting, but the answers themselves may be inconsistent. I have experimented with varying the prompt and the answers can change dramatically. I could experiment with lowering temperature, but
24.
▲
by
coder68
1y ago
I can confirm that Distillbert has worked well when I have used it for classification, especially on shortish sequences. I'm really interested in trying out ModernBert, or a smaller variant due to the larger context window (8192 tokens
25.
▲
by
coder68
1y ago
I think the idea of backoff by ratcheting up complexity here is a very good idea, thanks for your suggestions.
26.
▲
by
coder68
1y ago
Ah yes this does make sense. We are definitely in agreement on the point of "wildly inefficient and subpar". I'll try out decoder model embeddings soon, e.g. Qwen/Qwen3-Embedding-8B. I'm working with largish amounts
27.
▲
by
coder68
1y ago
By LLMs I meant decoder only, e.g. Gemini, Claude, etc. Can you go into more detail on how you're using the encoder models? I'm curious. Typically I have used them for embedding text or for fine-tuning after attaching a classifier
28.
▲
by
coder68
1y ago
Do you have any sources/links that talk about this? I'm very interested in synthetic data generation, so curious what you've tried or what works / doesn't work, especially with regards to LTR.
29.
▲
by
coder68
1y ago
I have been working on text classification tasks at work, and I have found that for my particular use-case, LLMs are not performing well at all. I have spent a few thousand dollars trying, and I have tried everything from few-shot to asking