Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
curioussquirrel
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
Have we made a unicorn? Continuous SVG-pelican style benchmark
(havewemadeaunicorn.com)
2 points
by
curioussquirrel
3mo ago
|
1 comments
2.
▲
by
curioussquirrel
3mo ago
As usual, take it with a grain of salt.
3.
▲
by
curioussquirrel
4mo ago
My first real programming achievement was building a website for an Ultima Online shard. I wrote some really terrible PHP and HTML, but it worked for 20+ years afterwards. Great memories! I was surprised that there is still an active commun
4.
▲
by
curioussquirrel
5mo ago
V4 is definitely a step-up from V3.2 on our multilingual benchmarks. Two caveats: - when inferring through Openrouter, we've had a lot of issues with very slow speeds (TPS) and an occasional instability. I just checked and it's st
5.
▲
by
curioussquirrel
5mo ago
Came here to say the same. Please add a few screenshots!
6.
▲
by
curioussquirrel
5mo ago
MoE is mostly an optimization of the active parameters and therefore lowering the compute requirements, but it can provide some performance improvements over dense models in some cases. I would not describe reasoning as optimization: In fac
7.
▲
by
curioussquirrel
5mo ago
There are architectural changes (such as reasoning or mixture of experts) that measurably improve how well models perform. So the improvements are definitely not just from data. I can speak for my area of expertise - multilingual capabiliti
8.
▲
by
curioussquirrel
5mo ago
Here you go! https://news.ycombinator.com/item?id=47847282
9.
▲
by
curioussquirrel
5mo ago
Just saw your thinking edit! That's a great question and one I wanted to study in depth, but these days you don't really get access to the raw thinking data. It's usually summarized and you can't even be sure what langua
10.
▲
by
curioussquirrel
5mo ago
I am fairly convinced that there's a certain polyglot snowball effect: once the LLM is fluent in 20 languages, it can pick up on similarities in vocabulary, syntax etc. and learn the 21st language with much less effort (and training da
11.
▲
by
curioussquirrel
5mo ago
One more thing: we're working on a multilingual benchmark that will evaluate core linguistic proficiency in 30 languages. We already have a lot of data internally and I can tell you that: - Gemini 3 Pro is a multilingual monster. - GPT
12.
▲
How well do LLMs work outside English? We tested 8 models in 8 languages [pdf]
(info.rws.com)
3 points
by
curioussquirrel
5mo ago
|
5 comments
13.
▲
by
curioussquirrel
5mo ago
Disclosure: I work at RWS/TrainAI, we did this study. Recently I alluded to it in a comment and was encouraged to share it, so here it is! We focus on multilingual proficiency, which tends to be understudied: most benchmarks are Englis
14.
▲
by
curioussquirrel
5mo ago
Yes, but post training cannot possibly account for all possible use cases. Sane defaults are fine, you can't really do much about sampling parameters in chatbots and coding harnesses anyway. And when making an API call, you have to act
15.
▲
by
curioussquirrel
5mo ago
After Anthropic, Moonshot is another model provider who restricts tweaking of sampling parameters. I do like the idea of the vendor verifier, though.
16.
▲
Claude Opus 4.7 API removes sampling parameters
(platform.claude.com)
5 points
by
curioussquirrel
5mo ago
|
1 comments
17.
▲
by
curioussquirrel
5mo ago
There's been quite a few threads about Opus 4.7 but none of them seems to have discussed some breaking changes on the API side, particularly removal of sampling parameters. From the migration guide: >> Sampling parameters removed
18.
▲
by
curioussquirrel
5mo ago
Will do! Thanks for the encouragement
19.
▲
by
curioussquirrel
5mo ago
Claude's tokenizers have actually been getting less efficient over the years (I think we're at the third iteration at the least since Sonnet 3.5). And if you prompt the LLM in a language other than English, or if your users prompt
20.
▲
by
curioussquirrel
5mo ago
Thanks for sharing! Have been begrudgingly using Darktable since that seems to be your best option on Linux, but the UI/UX never really clicked with me. I wish this was opensource but I will give this a shot (pun intended) for sure.
21.
▲
by
curioussquirrel
5mo ago
Thank you for the transparency and insights! Very helpful. We actually did the same thing re generating charts in brand style to avoid any mishaps, since then I sleep much better
22.
▲
by
curioussquirrel
5mo ago
Absolutely unhinged and very entertaining. Thanks for sharing!
23.
▲
by
curioussquirrel
6mo ago
Give Gemma 31B a shot for translation, it does a very good job at that given its size.
24.
▲
by
curioussquirrel
6mo ago
We're doing multilingual testing and I can confirm what you've observed: Gemma 4 is surprisingly good at multilingual tasks, especially given its size. This is mostly true for the dense 31B model.
25.
▲
by
curioussquirrel
6mo ago
Same, I quickly tested it for code gen and it produced mostly good code for simple problems, but it sometimes hallucinated words in non-English scripts inside the code.
26.
▲
by
curioussquirrel
6mo ago
For anyone interested in multilingual performance, which is not usually well benchmarked or reported: Gemma 4 does really well, especially the dense 31B version. In fact, it outperforms many models with an order of magnitude higher number o
27.
▲
by
curioussquirrel
6mo ago
Thank you. +1. There are obviously differences and things getting lost or slightly misaligned in the latent space, and these do cause degradation in reasoning quality, but the decline is very small in high resource languages.
28.
▲
psmux: Terminal multiplexer for Windows – tmux alternative
(github.com)
2 points
by
curioussquirrel
8mo ago
|
0 comments
29.
▲
by
curioussquirrel
8mo ago
Such a good game and execution. Thank you
30.
▲
Offline Regains Its Value
(blog.avas.space)
1 points
by
curioussquirrel
8mo ago
|
0 comments
More ›