Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
osti
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
osti
26d ago
You mean when they said they are doing safety evaluation and hardening? I'm having the same fear as you.
2.
▲
by
osti
1mo ago
It's an axiom that the modern Western mind is built upon.
3.
▲
by
osti
2mo ago
Do you have any numbers on the solve quality? Exploitability numbers etc.
4.
▲
by
osti
2mo ago
You can use geekbench 5 in that case. But given that they deprecated that, it might be harder to compare to others.
5.
▲
by
osti
2mo ago
I agree. For me personally I mostly only care about single thread geekbench variant, I believe it's an excellent proxy for general performance of a CPU. Multi thread geekbench (or other benchmarks) for most purposes and for most people
6.
▲
by
osti
2mo ago
In computer science, one definition of algorithm is basically any program that runs on a turing machine. By that definition, any LLM is an algorithm.
7.
▲
by
osti
2mo ago
Lol yet I've used Apple and Android phones extensively and would choose Android every single time.
8.
▲
by
osti
2mo ago
He's talking about the plans, you are talking about API prices.
9.
▲
by
osti
2mo ago
No idea lol, didn't even know those exist..
10.
▲
by
osti
2mo ago
It is complicated, but paying for the cheaper usd plans really don't get you much usage.
11.
▲
by
osti
2mo ago
Nah that won't work. I don't know tbh, I just used someone else's number.
12.
▲
by
osti
2mo ago
Absolutely do not pay for the kimi plans thinking they will be cheaper. If you sign up with a Chinese phone number, you can get the same plan for 200 yuan instead of 200 usd, it also only accepts Chinese payment methods iirc. So the plans a
13.
▲
by
osti
2mo ago
GPT should be better at these optimization problems given that they won the recent atcoder heuristics competition against top humans. And Anthropic is less focused on these types of things.
14.
▲
by
osti
2mo ago
This was extremely impressive to me. AtCoder has the hardest problems these days, usually the human onsite final round contestants can't solve more than 2 or 3 problems. This year the problem setter sets the round in a way that maximiz
15.
▲
by
osti
2mo ago
Mythos probably wouldn't, otherwise they'd have included it in their release. Next version of Mythos probably will though. And yeah.. Reality has not been kind to LeCun.
16.
▲
by
osti
2mo ago
SWE-bench series just aren't that great by today's standard, even Anthropic previously stated Claude had memorized solutions for the non Pro version of the benchmark, I suspect the recent increase in the score for the Pro version
17.
▲
by
osti
2mo ago
Huh so that's why it's hard to find. They probably haven't properly optimized their caching, or they are just trying to make more money from there.
18.
▲
by
osti
2mo ago
And they'd be right, it's an almost saturated benchmark where even some subpar open source models score very well on. And most models are clustered within a small range so it really doesn't tell you much.
19.
▲
by
osti
2mo ago
SWE-Bench pro is pretty much useless now even though many ppl still look at it. OpenAI published a report yesterday saying so as well. Only look at DeepSWE and FrontierCode right now for coding imo.
20.
▲
by
osti
2mo ago
GPT usually performs better on DeepSWE while Claude does better on FrontierCode. These two coding benchmarks are pretty much the only ones right now that's still worth taking a look at imo.
21.
▲
by
osti
2mo ago
Most programming is that, but most music and literature are probably uninspired junks as well. But there are many beautiful algorithms (such as the ones in Knuths books) that are more beautiful than any music for me personally.
22.
▲
by
osti
3mo ago
I think at the current stage of LLM, it just doesn't make sense to ever have an annual sub. Things change too quickly that you really don't want to be stuck with one model.
23.
▲
by
osti
3mo ago
If anything that'll be more obnoxious because they have to show the government that it's safe.
24.
▲
by
osti
3mo ago
This announcement also mentioned that they will release the next version (official non preview version) of v4 in mid July.
25.
▲
by
osti
3mo ago
Is it only the Indian government? I don't think that's in any way unique to India, I've seen many poor government websites.
26.
▲
by
osti
3mo ago
Ah got it. Reread previous comment and that makes sense.
27.
▲
by
osti
3mo ago
Sol? Looks like openai is jealous of anthropics good model naming ability and wants to emulate it.
28.
▲
by
osti
3mo ago
Hmm your last sentence seems to exactly agree that it's a class of algos that parallelize well? What does sped up arbitrarily mean? It's still polynomial speed up right?
29.
▲
by
osti
3mo ago
So is it a class of problems that can be parallelized well?
30.
▲
by
osti
3mo ago
Does that support modern gaming?
More ›