Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
coder543
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
1.
▲
by
coder543
7d ago
Some OpenRouter providers do not implement reasoning levels for these models correctly at all: https://www.reddit.com/r/DeepSeek/comments/1vdqjwr/openroute... If you're going to use OpenRouter to te
2.
▲
by
coder543
10d ago
"Dictionary gains are mostly effective in the first few KB." https://facebook.github.io/zstd/index.html Pretrained dictionaries have never been intended to help with book sized or bigger compression. zstd aut
3.
▲
by
coder543
13d ago
Please run GLM-5.3 and GLM-5.3-Flash. I would love to see how they do. On the smaller end of things, Qwen3.8-27B and Ling-3.0-Flash would also be interesting. In the benchmark, have you considered instructing the models to build their own S
4.
▲
by
coder543
13d ago
FunctionGemma never worked well for me (without fine tuning). Liquid has released 230M and 350M models that work far, far better in my testing: https://huggingface.co/LiquidAI/LFM2.5-230M I really look forward to a hyp
5.
▲
by
coder543
19d ago
Novita does not offer Hy4-preview on either OpenRouter or their own model list. Maybe you confused it with Hy3.
6.
▲
by
coder543
21d ago
I haven't tried it, but this looked promising for that exact task: https://huggingface.co/superwhisper/s1-mini
7.
▲
by
coder543
22d ago
The spark can easily run UD-Q4_K_XL on this model... using IQ1_S doesn't make much sense.
8.
▲
by
coder543
29d ago
No. What else could it reasonably be named? Hard to imagine. The rule has always been intended to cover types that have another word in them but still choose to pointlessly repeat the package name. `uuid.UUIDGenerator` is a hypothetical e
9.
▲
by
coder543
1mo ago
The post does not imply the 5090 is needed, that is just a common reference point. A single six year old RTX 3090 works great: https://www.reddit.com/r/LocalLLaMA/comments/1vkm42m/muse_gl... I fully expe
10.
▲
by
coder543
1mo ago
I think glancing at a random snapshot from today misses all the context. Nemotron 3 is far more significant than you're giving it credit for. At this point, Nemotron 3 is really an 8 month old model series. That's when Nemotron 3
11.
▲
by
coder543
1mo ago
Yes, the article uses AI-assisted text, so if that offends you, no need to accuse: you can just stop reading. I have a good amount of background knowledge on all of this, and I've invested hours into pulling this short series together,
12.
▲
The Next Token – LLMs, from the Beginning
(ceres1.space)
1 points
by
coder543
1mo ago
|
1 comments
13.
▲
by
coder543
2mo ago
A year and a half behind Seedance 2.0? That is a bold claim that needs evidence. According to one user preference leaderboard, MiniMax H3 is already ahead of Seedance 2.0 based on thousands of A/B votes: https://artificialan
14.
▲
by
coder543
2mo ago
The weights were just released a few minutes ago: https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731
15.
▲
by
coder543
2mo ago
But if you compare to Anthropic's models? The cost difference is huge. Anthropic is clearly concerned that people are realizing they are expensive, since the Opus 5 blog post dedicated a lot of time to talking about how cheap the mod
16.
▲
by
coder543
2mo ago
This article seems premature to post. Right now, the price is arbitrarily set by a single provider. Why wouldn't Moonshot collect extra revenue during this exclusivity period when they knew there would be hype? The model weights are su
17.
▲
by
coder543
2mo ago
At this point, I would not recommend ignoring Parakeet TDT 0.6b v2/v3 (english-only versus multilingual). Those models have been out for a year, give or take, and they are both accurate and fast. I would choose Parakeet over Whisper in
18.
▲
by
coder543
3mo ago
It’s not FUD. It is my actual, lived experience. FUD is false, which this is not. I use both vLLM and llama-server. vLLM is very painful, even with the Spark community docker image. It is slow to start, it does not support 3-bit dynamic qua
19.
▲
by
coder543
3mo ago
Unsloth Studio is also very low effort, and a lot better than LM Studio in my opinion. (Performance, compatibility with Gemma 4, actually open source, etc.)
20.
▲
by
coder543
3mo ago
Compared to a dynamic quant like Unsloth's UD-Q4_K_XL, which keeps some important parameters in higher precision, a basic NVFP4 quant seems to do a lot more damage to the model unless it is carefully calibrated. I would recommend using
21.
▲
by
coder543
3mo ago
> The systems but old but I’m seeing 11tks 27b, 15tks 35b MoE If that's accurate, then you must be doing something wrong/weird. On a single RTX 3090, I'm seeing substantially higher performance. Dual GPU won't necessa
22.
▲
by
coder543
3mo ago
> For a MBP I have 48 GB of RAM M5 Pro. It runs at about 12-14 t/s at Q4 Are you running with MTP enabled? I have seen some people on M5 hardware report 20+ t/s on Qwen3.6-27B using MTP... and I think that was a regular M5, not
23.
▲
by
coder543
3mo ago
5.5 Pro is $30 in / $180 out: https://developers.openai.com/api/docs/pricing I think you meant 5.5. I agree it is probably the same size model. It's probably exactly built on top of 5.5, just with more t
24.
▲
by
coder543
3mo ago
EDIT: It's just not even worth arguing this point, so deleting my original, much longer comment. Abstract taxonomies can claim that Taalas is CIM, but this entirely and utterly misses the point, and misses what makes Taalas' appro
25.
▲
by
coder543
3mo ago
CIM does not bake the weights into silicon. The level of optimization that you can do down to the last transistor when the weights are fixed is on an entirely different level than CIM where you still need general purpose ALUs all over the p
26.
▲
by
coder543
3mo ago
Yes, I’m focused on the topic at hand that the person I replied to was also talking about. The person I replied to was acting as if Taalas was ancient history. I was pointing out it has only been a few months.
27.
▲
by
coder543
3mo ago
> It's odd to me that I haven't heard anything about this approach since. It has only been four months since they unveiled their first prototype. I don't understand your confusion. Chip development does not happen overni
28.
▲
by
coder543
3mo ago
Taalas' first chip is for a Llama 3.1 8B quant, not a 3.1B parameter model, to clarify.
29.
▲
by
coder543
3mo ago
Yeah... one of the relevant issues: https://github.com/openai/codex/issues/11940#issuecomment-45... You would think they would support their own GPT-OSS model, but, not really anymore. I wish they would relea
30.
▲
by
coder543
3mo ago
Well, the reason is simple: over the past several months, it has become very difficult to use Codex with non-OpenAI models. They removed the old edit tool that didn't require OpenAI's free form tool calling (that no other LLM host
More ›