Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
vb-8448
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
vb-8448
3d ago
Everyone agreeing on something is not even possible on small local scales, just imagine the situation on a global scale. It's not a challenge at all because it's not possible, it's just empty rhetoric to push their own agenda
2.
▲
by
vb-8448
3d ago
My experience up to now is that the models can do both "less LOC" and "clean code" at the same time, you just have to keep reminding it to them. So the capabilities are definitively there.
3.
▲
by
vb-8448
3d ago
> The models do not have a fear of future regret. I so much feel this specific point. All models up to now (including astra, fable) are too much trained to "get the job done" and pass the benchmark that its doesn't care at
4.
▲
by
vb-8448
4d ago
Debt cost is high since some time, going public probably is just cheaper. Anyway, any evidence for "prices are moving inference to highly profitable" because they are still selling $20 for $1 with their subscriptions.
5.
▲
by
vb-8448
4d ago
If the alignment is such a problem ... why not using self-improving capabilities of the latest models to solve it?
6.
▲
by
vb-8448
4d ago
If there were really concerned about humanity future they'd donate everything to public and/or to not profits ... but the not profit turned in a for profit and the other one is seeking for the biggest IPO in history. Maybe they ar
7.
▲
by
vb-8448
5d ago
The only difference with the past that you need less time to dig through the codebase or documentation, the agent can do it for you and provide only meaningful info, but without a "Mental model"(or knowing what's going under
8.
▲
by
vb-8448
6d ago
The future is basically something between: AGI/ASI will kill us all and a privacy nightmare.
9.
▲
by
vb-8448
6d ago
So basically I record what the browser sends to the server when I click the toggle and the server response? I wonder how I can attach a timestamp that cannot be faked.
10.
▲
by
vb-8448
6d ago
Out of curiosity, how one is supposed to "document it properly"?
11.
▲
by
vb-8448
7d ago
Tower Bridge suits you? But hurry up, there are a lot of pretenders.
12.
▲
by
vb-8448
7d ago
If you trust a "toggle" I have a bridge to sell. Users data is just too precious to ignore.
13.
▲
by
vb-8448
8d ago
> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models AKA everything you send to them (and I bet it's the same for any other lab) will be used, no matter what
14.
▲
by
vb-8448
11d ago
Same experience here, but I have some strange feeling. Sol I'm used to working a month ago doesn't feel the same I'm using today, slower and less accurate. My gut feeling is that they quantize previous models to prioritize ne
15.
▲
by
vb-8448
12d ago
I find curios that Astra's pelicans are basically the same (yellow sun top right corner, green bike, same bike shape, same legs style, very similar background) while in other there is more randomness.
16.
▲
by
vb-8448
12d ago
Played in codex app a couple of hours today: it feels much faster than SOL, even if the TPS is half of it.
17.
▲
by
vb-8448
12d ago
But someone could probably build a harness what will be able to do play the game.
18.
▲
by
vb-8448
13d ago
But scored less on V2 and V1 ... too much overfitting?
19.
▲
by
vb-8448
13d ago
It's not a criticism, I was really looking forward to trying out such a powerful model at this speed. But I burn my 5$ allowance in 10 minutes ... and only because I was hitting rate limits, without it would probably be less than a min
20.
▲
by
vb-8448
13d ago
Imagine Luna at 10x tps and 1/100 of current cost. At that point you will be able to "brute force" basically everything. IMO also a lot of problems with memory and context rot will be solved too.
21.
▲
by
vb-8448
13d ago
At that speed it's too pricey for agentinc tasks.
22.
▲
by
vb-8448
13d ago
Am I the only that thinks that anything similar to AGI will come not from raw model capacity but from model speed and efficiency? In my experience the harness is more important than the model, and anything able to run at 700tps will be the
23.
▲
by
vb-8448
14d ago
The author's job description is literally "Staff Software Engineer at OpenTeams. Dask maintainer."!
24.
▲
by
vb-8448
14d ago
I agree on the past, when you have limited resource and have to print something on paper that you cannot recall to fix you have to carefully choose the layout. But that era is gone since decades, nowadays, given how easy it is, it's a
25.
▲
by
vb-8448
14d ago
> Only if you aren't schooled in reading graphs ... So we can say the same about the authors "AA’s plot is misleading" claim, he is "not schooled in reading graphs"? > How exactly would you zoom into a section
26.
▲
by
vb-8448
14d ago
It gives you a wrong perspective, especially if you are distracted, on model capabilities: Fable 5.1 is not 30% better than Sol, but is the very first impression you get when you look at the first graph. If I'm not wrong OAI tried a si
27.
▲
by
vb-8448
14d ago
Y-axis is between 0 and 100. But even if it was between 0 and Inf+, it still gives you a wrong perspective, especially if you are not paying attention, on model capabilities.
28.
▲
by
vb-8448
14d ago
Tried one of simonwillison's pelicans: We didn’t find any signs this file was processed by Claude!
29.
▲
by
vb-8448
14d ago
IMO in this case is mandatory to start from 0 because it alters the visual perception. Just look at the first chart: the distance between Fable 5.1 and Sol is <5%, but it looks like 25 or 30%.
30.
▲
by
vb-8448
14d ago
Complaining about "bad charting" and posting a chart with y-axis that doesn't start at 0 is kinda weird.
More ›