Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
GodelNumbering
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
GoBench: Evaluating LLMs on 9×9 Go using KataGo opponents as Elo anchors
(rolandgao.com)
2 points
by
GodelNumbering
2h ago
|
1 comments
2.
▲
by
GodelNumbering
5d ago
Some months ago I was evaluating command output compressors to integrate into Dirac[1] as that seemed like an easy win that would compliment and compound with Dirac's other mechanisms. I tested rtk among these and it was actually a net
3.
▲
by
GodelNumbering
6d ago
Tangential to the subject, but this is a bluesky post, containing a screenshot of an X post, which itself starts with "in a detailed Mastodon post"...
4.
▲
by
GodelNumbering
13d ago
> That's quite common with many models Such as? I can't think of any. Diminishing returns, yes. Occasionally flat, yes. Downright regression, no.
5.
▲
by
GodelNumbering
13d ago
The most interesting part, even more than ARC 3 score, to me is that this is the first model I recall seeing that scores lower on Max than High reasoning effort on some coding benchmarks: Terminal-Bench 4.0: High (57.9%), Max (56.7%) DeepSW
6.
▲
by
GodelNumbering
13d ago
I decided to front run and added support for it in Dirac (coding agent) a couple of hours ago, using best guess pricing: input/output/cache: $10/$50/$1.
7.
▲
by
GodelNumbering
14d ago
Do you have instructions on how to run a custom harness against this? There are none on the linked page. I want to run Dirac ( https://github.com/dirac-run/dirac ).
8.
▲
by
GodelNumbering
15d ago
Interesting, even if we were to ignore the cache-hits, reads and output, the reasoning cost (aka test time compute) per task should remain a fully comparable metric - it went from $1.25 (Fable5) to $1.48 (+18.4%) for an improvement signific
9.
▲
by
GodelNumbering
15d ago
The price reduction comes from the cache read pricing falling from $1/M to $0.25/M, which means that Fable 5.1 now costs half of Opus's cache read costs ($0.5/M). This gives a lot of credit to the theory that Anthropic d
10.
▲
by
GodelNumbering
22d ago
1.2 TB/s bandwidth of M5 Ultra comes from two dies of M5 Max (each 614 GB/s) connected together using 4.4 TB/s inter-die fabric. For a non-quantized Deepseek V4 flash on an ultra, I would estimate about 1000+ tokens per secon
11.
▲
by
GodelNumbering
24d ago
This is a plug, but relevant. I recently added a 'build native tools on the fly' functionality to Dirac ( https://github.com/dirac-run/dirac ) that works like: 1. You can use the '/new-tool' and
12.
▲
by
GodelNumbering
1mo ago
I agree in principle that the buck has to stop somewhere, but taking full responsibility for the actions of a fundamentally statistical system that you didn't build and have no interpretability of is uncomfortably risky for most partie
13.
▲
If your agent commits a crime, who is responsible?
(signalbloom.ai)
37 points
by
GodelNumbering
1mo ago
|
104 comments
14.
▲
by
GodelNumbering
1mo ago
I think introductory here means more or less permanent but they can't publicly admit there are no takers at a higher price. Anthropic for instance announced a couple of days ago that they are making Sonnet's 'introductory pri
15.
▲
by
GodelNumbering
1mo ago
The corresponding OpenAI post https://openai.com/index/previewing-ultrafast/ There is no pricing info, which could mean it's "if you have to ask..." territory or they are simply gauging interest bef
16.
▲
by
GodelNumbering
1mo ago
> introductory price They should call it 'face saving pricing after we realized just how terribly did we mis-price the flash 3.5' > since google betrayed us with those price hike, people already spent their time making their
17.
▲
by
GodelNumbering
1mo ago
https://xcancel.com/finkd/status/2086755195535413696 "... Soon we'll also release the weights for Muse Spark 1.2, our latest foundation model..." This is bigger news - good for self hosting enthusia
18.
▲
by
GodelNumbering
1mo ago
Am I missing something or the evals do not compare it to the baseline deepseek-v4-flash? Without a baseline comparison, it is hard to tell what works well and what doesn't
19.
▲
by
GodelNumbering
1mo ago
Apologies, I had no idea. I lookup host image and that site came up. I don't think the site itself is of NSFW nature.
20.
▲
by
GodelNumbering
1mo ago
If that was it, wouldn't Claude go directly to the canonical source?
21.
▲
by
GodelNumbering
1mo ago
I just checked Cloudflare for SignalBloom ( https://www.signalbloom.ai , which I own and operate). Over the last 72 hours, Claude-searchbot [1] alone fetched ~205,000 pages. Sent exactly 1 referral. There is a lot of free financia
22.
▲
by
GodelNumbering
1mo ago
You should play Blood on the clocktower
23.
▲
by
GodelNumbering
1mo ago
I think that's a fair offering tbh
24.
▲
by
GodelNumbering
1mo ago
If there were, do you believe it would be in their interest to answer this publicly?
25.
▲
by
GodelNumbering
1mo ago
Correct, but it was preview release. I was referring to GA in my comment.
26.
▲
by
GodelNumbering
1mo ago
This is one of the interesting aspects the 'AI job loss' community doesn't account for. As the technology unlocks things, more startups are created. And even at a lower nominal engineer-to-work ratio, overall demand for talen
27.
▲
by
GodelNumbering
1mo ago
So, in last several months, all the prominent names Google lost: Demis Hassabis (technically still with google but these things are usually presented with a spin), Jeff Dean, Sanjay Ghemawat, Oriol Vinyals, Quoc Le, Noam Shazeer, John Jumpe
28.
▲
by
GodelNumbering
1mo ago
I disagree to the maximum extent possible. I have lived in 6 countries. Americans to me are by far most friendly and supportive of new ideas. When you are visiting a place, you see it with rosy eyes. You are not forced to deal with the day
29.
▲
by
GodelNumbering
1mo ago
I have been writing a 'paper' [1] on an adjacent topic for months now. At some point, I decided to make it an empirical paper vs position paper. I am still chasing the experiments (when I get some free time waiting for agentic loo
30.
▲
by
GodelNumbering
1mo ago
Thanks, I will reach out. I have also posted for contributors on localllama https://www.reddit.com/r/LocalLLaMA/comments/1vg40w8/anyone_... > We can do a mix of general use (as in user stories) plus a
More ›