Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
easygenes
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
1.
▲
by
easygenes
1mo ago
In the lead-up to the Grok 4.5 release, the CEO of HuggingFace was petitioning Elon Musk to open the weights. It’s odd to me that he didn’t accept, as the model is at a similar competitive level to Muse 1.2. It would have been good PR and e
2.
▲
by
easygenes
1mo ago
Oof, even the post-mortem promising none of it is AI slop is obvious AI slop.
3.
▲
by
easygenes
2mo ago
That’s not what this indicates. This is the biggest and most expensive to serve, and most capable open weights model yet. They’re just pricing it in line with capabilities. Kimi also offers generous subscriptions. Subs aren’t going anywhere
4.
▲
by
easygenes
3mo ago
Was fun to see their developers make nods to Le Chaton Fat in the announcements for this on Twitter. I suspect a true "big new general-purpose" model is around the corner from them, whether or not they were in on Le Chaton Fat for
5.
▲
by
easygenes
3mo ago
I'm a heavy enough user that I have both the OAI and Anth $200 plans. I always use at least 50% of my weekly Opus quota at Extra setting (meaning I use double the limit of the $100 plan, at minimum). Max I rarely touch because it is tw
6.
▲
by
easygenes
3mo ago
There are. If the kernels are nondeterministic (e.g. timing issues) there are minor changes between runs, on a single system, even with eager decode enabled (typically what temperature=0 achieves).
7.
▲
by
easygenes
3mo ago
This is a strange one. We know the hardware capabilities of Cerebras force them to do aggressive REAP pruning to serve Kimi K2.6. Meaning that about 750B parameters is the upper limit of what they can serve economically. Not sure if this me
8.
▲
by
easygenes
3mo ago
M5 Ultra will ship before end of year, likely. Though with current RAM shortage, likely max spec will be 256GB and in short supply. In late 2027 or early 2028, Nvidia will release Vera Rubin DGX Spark, likely with double or better the perfo
9.
▲
by
easygenes
3mo ago
Article reads as though written by someone who doesn't have much experience with deployments like this. Underestimates the memory needed to run with a reasonable amount of context. Misses two other obvious targets: 1) 4x DGX Spark
10.
▲
by
easygenes
3mo ago
The Wired headline reframes the issue in a way that’s misleading. SK Telecom was a previously resolved issue (as in prior to Fable launch). It may have been a contributing factor, but the crux of the shutdown was the industry reporting of F
11.
▲
by
easygenes
3mo ago
This headline is not what I would read from this. The numbers are more favorable than the general tone of rumors, and point towards the expected shape of a fast-growing R&D heavy business.
12.
▲
by
easygenes
3mo ago
Announcement from the founder of Z.ai: “ GLM-5.2 is Fully Open, Frontier Intelligence Belongs to Everyone Today, the sudden restriction of certain frontier models is deeply regrettable. At a time when access to frontier models is abruptly c
13.
▲
by
easygenes
3mo ago
This release was rushed to hang on the coattails of the Mythos drama (“hey, sorry you can’t use Fable, but try us while you wait this weekend!”) I think they planned to release next week, hence benchmarks not all being ready yet.
14.
▲
by
easygenes
3mo ago
That happened a year ago when these shipped as the DGX Spark with only Linux pre installed.
15.
▲
by
easygenes
3mo ago
Mostly a strategy move to protect the CUDA moat… Apple would take over mobile inference in a clean sweep without competition.
16.
▲
by
easygenes
3mo ago
This is the same chip and same memory. Only difference is it is going in a laptop, so will be more thermally limited.
17.
▲
by
easygenes
4mo ago
If I were paying API rates this year, I would have already burned through $20k in tokens. Looking forward to the costs of this level of capability coming down.
18.
▲
by
easygenes
4mo ago
I have now also tried it on this scatter plot: https://3215535692-files.gitbook.io/~/files/v0/b/gitbook-x-p... Similarly, the 26B A4B Gemma 4 and the 35B A3B Qwen 3.6 identify it clearly, give me the tit
19.
▲
by
easygenes
4mo ago
They haven't made one for this new model, but Unsloth has a comprehensive quant KLD map of Gemma 4 26B A4B here: https://3215535692-files.gitbook.io/~/files/v0/b/gitbook-x-p...
20.
▲
by
easygenes
4mo ago
I want to like the vision capabilities of the model. However, when I gave it an image which Gemma 26B A4B and Qwen 3.6 35B A3B has no problem correctly describing in detail, including identifying the Taj Mahal in the background it utterly f
21.
▲
by
easygenes
4mo ago
Have you run it through DeepSWE? I understand that's probably a high ask for this class of model, but would be interesting to see regardless. Even if it can't fully pass much, there are so many tests against most of the scenarios
22.
▲
by
easygenes
4mo ago
While I agree directionally, I'll caveat that "cost per token" != "cost per task". In the case of Qwen3.6 it tends to think 1.6x more than Haiku, so the cost of Haiku on the same tasks tends to only be about double.
23.
▲
by
easygenes
4mo ago
Speaking as someone who has had a DGX Spark all year and been active developing at the driver and kernel level for it and other ARM64 Linux devices the last couple of years, it's not bad now and certainly doesn't have any issues t
24.
▲
by
easygenes
4mo ago
Looks like RTX Spark desktop is the DGX Spark desktop, minus the expensive 200GbE Connect-X NIC. Only since the DGX Spark released, memory and nand prices have jumped, so it will likely retail for the same amount as the DGX Spark did on rel
25.
▲
by
easygenes
4mo ago
Claude Opus 4.7 defaults to exactly this design language for a lot of "just make me a rich html presentation page" requests without further specification.
26.
▲
by
easygenes
4mo ago
While I endorse the message of TFA (though do find the framing a bit on the overly blunt side), I believe it's unfair to reduce to "losing the person". The person is still willing to engage with you and still had to use their
27.
▲
by
easygenes
4mo ago
Their methods are only calibrated on open models (of course) and they admit very broad confidence bounds. You can also just see from comparing their estimates of the same models at different reasoning levels that there are major confounders
28.
▲
by
easygenes
4mo ago
This is the reality of the premiums available from being in the lead by ~8 months on model building technicals.
29.
▲
by
easygenes
4mo ago
We know from NVIDIA's public Vera Rubin inference engine marketing materials that the frontier lab models are ~1-2T total. Mythos is an exception that's larger.
30.
▲
by
easygenes
4mo ago
For those who would like to know the total and active parameter count of this model: even though Google doesn't disclose the model technicals, we can infer them within relatively tight margins based on what we do know. We know they ser
More ›