Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
lukev
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
lukev
15d ago
Begun, the AI SEO wars have.
2.
▲
by
lukev
17d ago
From the report: > Because there were over a thousand transcripts and most were extremely long, we had to heavily delegate our analysis to AI agents; these agents had significantly worse judgment and reliability than human researchers, a
3.
▲
by
lukev
17d ago
The elephant in the room here is that the METR report itself was researched and compiled almost entirely by AI, with only very limited human "spot checks." So I'm really not sure how much of it can be believed, especially sin
4.
▲
by
lukev
2mo ago
Well, it’s mentioned as a limitation of the analysis, very much not ruled out (or in.) That simonw is causing labs to do extra fine-tuning runs for this seems highly probable :)
5.
▲
by
lukev
2mo ago
What if they’re not pelicanmaxxing, but svgmaxxxing in general? Because otherwise using a LLM to generate complex svgs is pretty niche and what I thought made this a good benchmark when it was new - generalized programming and spatial knowl
6.
▲
by
lukev
4mo ago
You get calls about a new service or promotion, and it's the diction of the caller that makes you not wish to engage...?!
7.
▲
by
lukev
5mo ago
Counterpoint: the standardized surface area of a browser is already enormous, and while these components seem simple, there are a billion different options, variables or alternative implementations to consider. At some point, functionality
8.
▲
by
lukev
5mo ago
What was the “well defined” definition? I’m not aware of any other than “this particular thing a human can do that I expect would be difficult for a computer.”
9.
▲
by
lukev
5mo ago
"intelligence" is not well defined. LLMs are throwing this into high relief with how "spiky" their capability curve is. Yes, they can solve some crazy hard problems with enough compute and thinking tokens. Yes, they also
10.
▲
by
lukev
5mo ago
I would bet that Canva's bet is that companies will always want a "last mile" of manual control, even if only for the Queen's Duck effect. If Canva is the default, zero friction path for that, great for them. The alterna
11.
▲
by
lukev
5mo ago
Well if you are talking about environmental stuff (like leaded gasoline), sure. If you’re talking about trying to improve the genetics of populations at scale… yikes.
12.
▲
by
lukev
5mo ago
This is a must-read series of articles, and I think Kyle is very much correct. The comparison to the adoption of automobiles is apt, and something I've thought about before as well. Just because a technology can be useful doesn't
13.
▲
by
lukev
5mo ago
To be clear: most people who are keen on making such an argument, or who are identifying racial genetic differences as the primary takeaway of studies like this, are doing so to justify racism, either implicitly or explicitly. But that'
14.
▲
by
lukev
5mo ago
Did you read the article? There's a whole section on "this is already happening."
15.
▲
by
lukev
5mo ago
I think we should all consider the possibility that part of the reason Anthropic hasn't immediately released Mythos is that it would be slightly disappointing relative to the benchmark scores.
16.
▲
by
lukev
5mo ago
This is a really interesting point though -- it's really scaffold-dependent. Because for the same price, you could point the small model at each function, one by one, N times each, across N prompts instructing it to look for a specific
17.
▲
by
lukev
5mo ago
Disagree, I don't particularly want to up the level at which I'm building the core. Core is where I want to prioritize quality over speed, and (at least with today's models) what I build by hand is much, much higher quality
18.
▲
by
lukev
5mo ago
I like this framing, but it does seem to imply that a whole dev shop, or a whole product, can or should be built at the same level. The fact is, I think the art of building well with AI (and I'm not saying it's easy) is to have a
19.
▲
by
lukev
6mo ago
Yes. That is the point I was making. Calculators provide a deterministic solution to a well-defined task. LLMs don't.
20.
▲
by
lukev
6mo ago
That's not what I mean. If I use a calculator to find a logarithm, and I know what a logarithm is, then the answer the calculator gives me is perfectly useful and 100% substitutable for what I would have found if I'd calculated th
21.
▲
by
lukev
6mo ago
I think that's too easy an analogy, though. Calculators are deterministically correct given the right input. It does not require expert judgement on whether an answer they gave is reasonable or not. As someone who uses LLMs all day f
22.
▲
by
lukev
6mo ago
So, isn't this a rather longwinded way to say that a signature only extends to the scope of the message it contains? It doesn't matter if I sign the word "yes", if you don't know what question is being asked. The si
23.
▲
Vibe Code vs. Trad Code
(dolthub.com)
6 points
by
lukev
6mo ago
|
5 comments
24.
▲
by
lukev
6mo ago
I could not possibly enumerate all the possible things that have been enclosed. Human beings obviously being the most morally egregious.
25.
▲
by
lukev
6mo ago
Also, the defining feature of capitalism is that it encloses what was previously common . Land used not to be owned (feudal lordship was functionally different than private ownership.) Then, society shifted, land became private, and that
26.
▲
by
lukev
6mo ago
Rent is charging money for access to an asset or property . The P in IP is Property.
27.
▲
by
lukev
6mo ago
I'm not sure how this relates to AGI. This measures the ability of a LLM to succeed in a certain class of games. Sure, that could be a valuable metric on how powerful (or even generally powerful) a LLM is. Humans may or may not be good
28.
▲
by
lukev
6mo ago
This is bad in tech. But at least we are (relatively) well equipped to deal with it. My partner teaches at a small college. These people are absolutely lost , with administration totally sold on the idea that "AI is the future" w
29.
▲
by
lukev
6mo ago
That's right -- the best way to succeed within a system is to hustle as hard as you can, and definitely don't stop to question the system itself.
30.
▲
by
lukev
6mo ago
I'm not speaking of burdens of proof about unfalsifiable statements. I'm saying that I think this is an important enough question that I think we should seek real evidence in either direction, especially since apparently everyon
More ›