Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
redox99
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
31.
▲
by
redox99
8d ago
The stochastic parrots have done it again!
32.
▲
by
redox99
8d ago
Mistral models are way worse than Chinese models in the real world. It's not benchmarks.
33.
▲
by
redox99
8d ago
Yes and no. It means the system is pushed away from a macroscopic theory into one where molecular effects matter. So it's not that you'd get infinite velocities in the real world, but you might get significant real world behavior
34.
▲
by
redox99
9d ago
Kinda surprised they didn't get a single sale for some of those. I think it probably looked too much like AI slop so customers refrained.
35.
▲
by
redox99
9d ago
I use the $100 which currently doesn't have the 5h limit. The 5h limit hurts the most on the $20 plan because the limit is already very small (5h is ~15% of your weekly).
36.
▲
by
redox99
9d ago
5h limits are awful. It means it is literally unable to complete a large task. A better approach if they want to balance the load is having higher usage or lower usage consumption at different times of day.
37.
▲
by
redox99
12d ago
It is (uses way less tokens)
38.
▲
by
redox99
12d ago
The problem is not photorealism. The SVG is outright dumb, its on the wrong side of the table, the table has fucked up geometry (its tilted) and many more minor flaws.
39.
▲
by
redox99
12d ago
Images from their X account
40.
▲
by
redox99
12d ago
They realized Astra having the same score as Sol was silly so they rushed to update the index so it fits what people expect. The old index was clearly bad (Astra is way better than Sol) but it's also unscientific to tweak it like this.
41.
▲
by
redox99
12d ago
Those two you mentioned completely demolish opus 4.5. It's not even close. I'd say they are between opus 4.8 and opus 5. And better in some tasks.
42.
▲
by
redox99
12d ago
4.5 became useful for one shotting large features. LLMs were useful for coding ever since GPT3 (copilot), and sonnet 3.5 for agentic coding.
43.
▲
by
redox99
13d ago
Yeah 5 was very underwhelming.
44.
▲
by
redox99
13d ago
The jump from 3.5 to 4 felt gigantic to me back then. GPT 5.0 did feel underwhelming though.
45.
▲
by
redox99
13d ago
I hate the term "AGI" but IMO Fable, 5.6 Sol, et al. were already AGI.
46.
▲
by
redox99
14d ago
Definitely not Google, countless horror stories and infamous for killing stuff. OpenAI is still serving GPT 3.5 turbo as far as I remember.
47.
▲
by
redox99
14d ago
You can ask 100 people and they'll all give you a different list. It's subjective. I think a less personal ranking would be, as a business owner, which of those providers is more dependable? As in, you don't care about evil,
48.
▲
by
redox99
14d ago
Google, Zuck, Sama, Elon, Amodei (in no particular order). They all suck. Pick your poison.
49.
▲
by
redox99
14d ago
No because the frontier keeps advancing very fast.
50.
▲
by
redox99
14d ago
Yeah for sure. Easily 10x that. And obviously we're talking just output tokens. I was mostly just pointing out a theoretical lower bound.
51.
▲
by
redox99
14d ago
> I think it's roughly possible for any piece of software. Still not possible for games (many people are attempting it on X). It's definitely getting better with every new model and it's just a matter of time before they c
52.
▲
by
redox99
14d ago
Obviously the agent would need to build a very extensive test setup for Paint.NET, a lot of it with computer use or something similar. Opening the same files, performing the same clicks, etc should output the same on paint.net and the clone
53.
▲
by
redox99
14d ago
700k LoC of human written code is probably like 2M LoC of LLM "slop". That's probably around 8M tokens. Assuming a single fable agent that one shots the new source code at 30 tok/s, that's 74 hours. Of course it nee
54.
▲
by
redox99
14d ago
How much in tokens/Claude subs would it cost to make a paint.net open source clone? Seems like the kind of software that with enough tokens an agent would be able to create from scratch autonomously (because you have a clear goal and r
55.
▲
by
redox99
14d ago
Qwen 3.8 27B absolutely demolishes Mistral best offerings at coding, and you only need a 5090 or 2x3090 to run it.
56.
▲
by
redox99
14d ago
Just wait and see
57.
▲
by
redox99
15d ago
It's just a leak don't take it too seriously, the model releases later this week.
58.
▲
by
redox99
15d ago
Am I 100% sure the author is telling the truth? No Does it matter? No
59.
▲
by
redox99
15d ago
It's not my image
60.
▲
by
redox99
15d ago
Yes
More ›