Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
KaoruAoiShiho
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
20 ms
·
1.
▲
by
KaoruAoiShiho
2mo ago
Appears to be benchmaxxing https://x.com/quietnning/status/2080786711861407883
2.
▲
by
KaoruAoiShiho
2mo ago
You're probably just responding to the headline but this person is an AI bull and isn't claiming it's a big deal, she's going into it and explaining it.
3.
▲
by
KaoruAoiShiho
2mo ago
Click through to the link, the answer is no it uses the latest gpt models now.
4.
▲
by
KaoruAoiShiho
3mo ago
I think the parent's point is that if you are genuinely open to losing, the arguments can be productive because you can learn something instead... So stopping arguments is just another way of closing yourself off.
5.
▲
by
KaoruAoiShiho
3mo ago
And before you know it, you invented some openrouter provider from first principles...
6.
▲
by
KaoruAoiShiho
3mo ago
Are you sure fireworks is unquant? It's not listing precision on openrouter like everyone else.
7.
▲
by
KaoruAoiShiho
3mo ago
Terrible zero value article, I am extremely surprised it is upvoted. That being said Artificial Analysis just came out with a brand new benchmark where it scored between opus 4.8 and gpt-5.5 and well behind fable-5 so it's definitely f
8.
▲
by
KaoruAoiShiho
3mo ago
This is really held back by one bench (omniscience accuracy) where it's really very far behind otherwise i think it's got at least a couple of points higher.
9.
▲
by
KaoruAoiShiho
3mo ago
Fable largely fixed the annoying chatterness so sucks that it's gone now.
10.
▲
by
KaoruAoiShiho
4mo ago
Then that's just a video game might as well as play a video game why limit yourself to still confined to the rules of chess.
11.
▲
by
KaoruAoiShiho
4mo ago
Kimi 2.5 has the best long context. For raw coding benchmark scores you can just post train on top of it with more specialized data. 2.5 is kinda old, 2.6 is the current release which is exactly just that and catches up to the frontier in m
12.
▲
by
KaoruAoiShiho
4mo ago
Do people like "gacha"? I thought people played games for the game experience, story, etc, and the gacha is just the monetization mechanic. It's like making a big deal out of paying $20 bucks a month, or buying loads of DLCs
13.
▲
by
KaoruAoiShiho
5mo ago
No I think the best agent with hundreds of millions in ARR should be worth more than the 15th best model company with tiny revenue. ur the joke.
14.
▲
by
KaoruAoiShiho
5mo ago
Manus is saved, 2 billion is such an undervaluation considering much worse companies like minimax is valued at 30 billion.
15.
▲
by
KaoruAoiShiho
5mo ago
SOTA MRCR (or would've been a few hours earlier... beaten by 5.5), I've long thought of this as the most important non-agentic benchmark, so this is especially impressive. Beats Opus 4.7 here
16.
▲
by
KaoruAoiShiho
5mo ago
Huh, that's not a thing?
17.
▲
by
KaoruAoiShiho
5mo ago
Might be sticking with 4.6 it's only been 20 minutes of using 4.7 and there are annoyances I didn't face with 4.6 what the heck. Huge downgrade on MRCR too.... 256K: - Opus 4.6: 91.9% - Opus 4.7: 59.2% 1M: - Opus 4.6: 78.3% - Opus
18.
▲
by
KaoruAoiShiho
5mo ago
Talking nonsense.
19.
▲
by
KaoruAoiShiho
5mo ago
Well they're not public yet so you'll have to put up with rumors. But the numbers are available for companies like DeepSeek say they have an 80% profit margin, so it stands to reason OAI etc would do similar numbers considering th
20.
▲
by
KaoruAoiShiho
5mo ago
After googling https://www.reddit.com/r/singularity/comments/1psesym/openai...
21.
▲
by
KaoruAoiShiho
5mo ago
TLDR: Writer hasn't heard of agents.
22.
▲
by
KaoruAoiShiho
5mo ago
I feel like netflix is definitely very cheap, with OpenClaw or whatever your favorite agent is, it's trivial to subscribe to watch one show and then have it cancel immediately.
23.
▲
by
KaoruAoiShiho
5mo ago
Blog post is new but the model is about 2 weeks in public.
24.
▲
by
KaoruAoiShiho
5mo ago
The non-awesome context window is the sad part, but I think a better harness can deal with this.
25.
▲
by
KaoruAoiShiho
5mo ago
Can you paste the relevant section in your soul please?
26.
▲
by
KaoruAoiShiho
7mo ago
Buy more GPUs.
27.
▲
by
KaoruAoiShiho
7mo ago
Sam Altman gave millions to Andrew Yang for pushign UBI, so they are trying to forewarn and experiment with finding the right solution. Most of the world prefers to shove their heads in the sand though and call them grifters, so of course w
28.
▲
by
KaoruAoiShiho
8mo ago
Something like this: Character Name: Marcus Cole Voice Profile: A bright, agile male voice with a natural upward lift, delivering lines at a brisk, energetic pace. Pitch leans high with spark, volume projects clearly—near-shouting at peaks—
29.
▲
by
KaoruAoiShiho
8mo ago
Have you tried specifying the emotion? There's an option to do so and if it's left empty it wouldn't surprise me if it defaulted to rng instead of bland.
30.
▲
by
KaoruAoiShiho
8mo ago
What's the best claude code terminal? I'm not sure if ghostty is it, which one can sync to iphone / android tablet for remote use of the same session?
More ›