Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
kadoban
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
10 ms
·
1.
▲
by
kadoban
3d ago
> Anthropic paid an average of $3000 per work they scanned based on their settlement Not sure you get to count breaking the law and getting in trouble in your cost-of-doing-business. That's a little too on the nose. You're basi
2.
▲
by
kadoban
4d ago
> That seemed silly at the time. Still does. It's an idiotic idea that only goes to show either how stupid Musk is or how stupid he thinks we are.
3.
▲
by
kadoban
4d ago
I think you know that's basically nothing like this? The model cards have every incentive to be biased, this doesn't necessarily. But even so, pretty much yes: companies that actually have reliable and accurate info in their relea
4.
▲
by
kadoban
4d ago
If it builds up history and perceived reliability, this type of thing can be valuable. You're giving up transparency for it being harder to game.
5.
▲
by
kadoban
5d ago
That's pretty good vfx, that'll be a useful tool. It has a couple of janky frames, the cup position is wrong on the first hard pan upwards, probably pretty easy to fix though. I believe there's already tools for essentially t
6.
▲
by
kadoban
5d ago
It almost would feel like a fair graph, but it's also below that one ~outlier in the middle. So it's ignoring all data but two points and purporting to fit an exponential to it. That's within spitting distance of meaningless,
7.
▲
by
kadoban
5d ago
It was created by Nixon, so yeah I'm sure the early days were a little weird (and it hasn't gotten any better since).
8.
▲
by
kadoban
5d ago
> The situation you're imagining is one of a global all encompassing depression caused by a single government. Doesn't even have to go that far (though it pretty easily can, the US is a big enough market). If the US causes a re
9.
▲
by
kadoban
5d ago
Can these big companies that want to distill not just get the models? Like how many machines at how many different providers are running Opus 5? Nobody is just throwing the model on a thumb drive and taking it home, or grabbing an old drive
10.
▲
by
kadoban
5d ago
> "Don't interrupt your opponent when they're making a mistake" Yeah, that's the obvious answer, but I suspect there's a point where if the US fucks things up too quickly and too messily, there's fewer
11.
▲
by
kadoban
5d ago
Is it me or is that not what that graph shows? That line has ~no relation to the datapoints it's purporting to be a fit for.
12.
▲
by
kadoban
5d ago
I think it's not even that sophisticated here, the providers are just lying.
13.
▲
by
kadoban
5d ago
The context of this is which party China would prefer to win the midterms. The dems winning would distract Trump from being able to fuck up the rest of the world a bit. My bet would be that China doesn't care much either way, they win
14.
▲
by
kadoban
5d ago
Yeah something must be actively broken or just an awful extension for it to have that much effect. Most things you can mess with the big effect is like, oh a thousand tokens ended up in the ~system prompt, or 10% extra or fewer work based o
15.
▲
by
kadoban
5d ago
> the triple digits IQ of Hegseth, Trump and Bessent. One digit each.
16.
▲
by
kadoban
5d ago
At some point they might prefer stability and sanity. I would bet that at this point they don't care much either way. China is playing long, most of the ~unchanging things are in their favor and the US already proved that they're
17.
▲
by
kadoban
5d ago
You know you can define your own provider filters and orderings, right? The filters and such are not _that_ advanced, but it might do what you need if you haven't already tried that.
18.
▲
by
kadoban
5d ago
You'd have to look where the extra tokens are coming from. If it's using extra turns or doing extra work because a tool it ~wants is missing, then adding extra things will help. Otherwise, it won't. If it's missing guida
19.
▲
by
kadoban
6d ago
Pi is just a nice base and it has defined extension protocols and such. You might as well start there, it's just easier and going from nothing to working to adding whatever functionality is like 2 minutes.
20.
▲
by
kadoban
6d ago
> "this is a benchmark" is such an easy category to determine I mean, it's not _that_ hard to determine most likely, and/or it's hard to be sure you didn't get found out by llm-assisted analysis on your traf
21.
▲
by
kadoban
6d ago
> AlphaZero (which later became Leela? I'm not sure) Pretty much. DeepMind never released their code or models, just the main ideas of their techniques and training. Leela, which came out of the go-ai world, implemented the paper(s)
22.
▲
by
kadoban
6d ago
Yes. The ways to improve are: faster search (which lets you search more deeply) or better evaluation of a position or better heuristics.
23.
▲
by
kadoban
7d ago
Doing this in a way that doesn't get you noticed or fucked with is potentially going to be quite difficult. If they have a "hey we're being benchmarked" mode, which is not hard to imagine, avoiding tripping it is going t
24.
▲
by
kadoban
7d ago
Way back when I worked in a restaurant, there was that but then also if a customer complains they also go clean. They're never going to be like "Oh, someone shat on the floor? We'll deal with that at the next scheduled cleani
25.
▲
by
kadoban
9d ago
> The fact that they didn't understand this when they proposed the idea shows you how little they knew of the Linux ecosystem, and the fact that they're still bringing up so many years later shows that they haven't learnt
26.
▲
by
kadoban
9d ago
> It's surprising it's that high (8%). I would imagine it takes as long to mop a floor that has pee droplets or mop a floor that doesn’t. You might have to mop more often from complaints or inspection, maybe?
27.
▲
by
kadoban
10d ago
> AI slop nothing burger. The “exploit” has nothing to do with coding agents. It's pretty small potatoes, but it is a harness ~bug that they treat this so poorly. It getting triggered before some of them even ask you if you trust th
28.
▲
by
kadoban
11d ago
Yeah, thanks, I forgot that existed. That definitely weakens my point quite a bit. It's still not _quite_ the same thing because an advantage in playouts is not a great model for how a stronger/weaker player dynamic actually works
29.
▲
by
kadoban
11d ago
That's the basic idea, yeah. When playing white in handicap games, you want to make your opponent uncomfortable. Play moves where the simple/safe/obvious move is just a little bit bad. Force them to choose between complex fig
30.
▲
by
kadoban
13d ago
> That humans can make sufficiently strong calculators has never been a dispute in my mind. Within our lifetimes (unless you're quite young) it was doubtful if a go ai would ever beat a decent human. Same was true for chess a genera
More ›