Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
re-thc
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
19 ms
·
1.
▲
by
re-thc
3d ago
There were comparisons and Muse Spark is so very similar to Fable / Opus... so...
2.
▲
by
re-thc
4d ago
> You refuse a direct order. I never said. It's not about outright refuse. It's how to deliver better outcomes. It's not a win or lose situation. Don't treat it like that. > If your boss says that he does not care
3.
▲
by
re-thc
5d ago
> Where LLMs excel is in code-level bugs (as opposed to system bugs, design bugs, architecture bugs, integration bugs, etc). Blame the benchmarks game. They're optimizing for that and that's what those things are measuring.
4.
▲
by
re-thc
5d ago
> Work for business people who want fast results. Agentic coding gets you to something presentable much faster at the cost of code quality. I have never seen a customer or business person care about that. That's always false. It
5.
▲
by
re-thc
11d ago
> A high score on benchmarks is not as useful because a model overtrained to always answer will give confidently wrong responses. It's not useful because the benchmarks often measure the wrong thing. They're here yapping about
6.
▲
by
re-thc
11d ago
It did used to use Haiku but that model is now too too far behind…
7.
▲
by
re-thc
13d ago
Hardware definitely has longer lifecycle than AI model releases at this point. You don't see Nvidia and AMD fighting every other month over the latest cards.
8.
▲
by
re-thc
13d ago
Overtaken in cost per token.
9.
▲
by
re-thc
14d ago
> DeepSWE is a very big deal It's clearly been "dealt with" already. When it launched we had interesting gaps and definitely differences. Now every new release is "crushing it".
10.
▲
by
re-thc
14d ago
> 4.0 Flash we will finally get Gemini 3.5 Pro Nah, we'll just get the 4.0 Pro Preview.
11.
▲
by
re-thc
14d ago
> Beginning to think Google is a dark horse in this race Google was so hyped up early Gemini 3 era (only some months ago). And now dark horse? The TPU takeover almost crashed nvidia and everyone else.
12.
▲
by
re-thc
15d ago
> people are still saying the Codex limits are more generous. They're not They are if you follow Tibo on the resets.
13.
▲
by
re-thc
15d ago
With OpenAI you can also apply for the security program, which doesn't require you to be a certified pentester (as per Anthropic).
14.
▲
by
re-thc
15d ago
> IMO, Codex is worse than Claude with Fable. Fable easily trips its safe guards. You can be 95% complete with the plan for it to trip and then lose it all. Anything is better than nothing.
15.
▲
by
re-thc
15d ago
That's load bearing!
16.
▲
by
re-thc
15d ago
The biggest change is the price cut of course.
17.
▲
by
re-thc
18d ago
Agent scale!
18.
▲
by
re-thc
19d ago
> Otherwise the premise for coding using LLMs is essentially untrue. Why? Written by does not mean designed by etc. There's a lot more to it.
19.
▲
by
re-thc
19d ago
Go has API pricing + this weird scaling of how much is it worth. Some models get $60 of usage, some $30 and some $15 etc.
20.
▲
by
re-thc
19d ago
> that would prove rather embarrassing for Anthropic Not really, in that you just work with different constraints. Anthropic and US labs in general has maybe 100s to 1000s of GPUs per person to experiment. Zai and Chinese labs in general
21.
▲
by
re-thc
21d ago
And it could have expanded elsewhere?
22.
▲
by
re-thc
21d ago
> They should've just lead with real, up to date data, because it's good, not the silly old tactics like comparing to Opus 4.8 when 5.0 is out in many of their charts It's what people know. Opus is just the common target.
23.
▲
by
re-thc
21d ago
> The export controls were revoked before Zai is on another "export control" list outside the broader 1. Doesn't help.
24.
▲
by
re-thc
21d ago
Which 9/10 times hasn't been great anyway (stock reaction).
25.
▲
by
re-thc
21d ago
> From a biased source, but would be big if true. I've had great results with GLM 5.2. It's at least close (even if not better) from the Ox Alpha runs. For the price it's definitely great.
26.
▲
by
re-thc
21d ago
2. There was a new checkpoint. Official.
27.
▲
by
re-thc
21d ago
With vision on top
28.
▲
by
re-thc
21d ago
the outperform Fable was a mid (not completed) benchmark run. Real results were lower.
29.
▲
by
re-thc
23d ago
> _China doesn't need to invade Taiwan._ All your points consider what China can do but not how others would respond. Blocking Taiwan when the whole AI economy is running on it will invoke lots of countries to go up in arms. There w
30.
▲
by
re-thc
23d ago
> this model is around 64B and can run on laptops. We're lucky if it'd fit in 1 DGX Spark. Laptops - nah, unless you mean like an M5 Max with 128GB of RAM then maybe. > Thats the reason behind the hype. The hype is imagine D
More ›