Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
timabdulla
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
1.
▲
by
timabdulla
6mo ago
It's a neat idea, but the internet is not very tolerant to things that don't appear to be human traffic. That's why browsers used by web automation infrastructure are often Chromium-derived. Using a browser like this would al
2.
▲
by
timabdulla
7mo ago
Google tends to trumpet preview models that aren't actually production-grade. For instance, both 3 Pro and Flash suffer from looping and tool-calling issues. I would love for them to eliminate these issues because just touting benchmar
3.
▲
by
timabdulla
8mo ago
I'd be curious to see screenshots or a video! I only have a Mac at my disposal, unfortunately.
4.
▲
by
timabdulla
8mo ago
This seems cool, but beware that Fly's other products are not exactly models of stability and polish. API downtime is a semi-frequent occurrence, as are transient API errors and slowness. I've also had a ticket open with support f
5.
▲
by
timabdulla
1y ago
> We conducted three runs per experiment and selected the run with the highest final accuracy for inclusion in the chart (though illustrative examples and anecdotes may be drawn from any of the runs). Can you comment on the variance? It&
6.
▲
by
timabdulla
1y ago
I mean, the fact that OpenAI, at the bleeding edge of it all, has decided to buy an IDE is a rather strong hint that the future of agents handling entire engineering tickets might be further out than many believe. If autonomous agents were
7.
▲
by
timabdulla
1y ago
What were the human PhDs able to do after more than 48 hours of effort? Presumably given that these are top-level PhDs, the replication success rate would be close to 100%?
8.
▲
by
timabdulla
1y ago
How does it perform on e.g. WebVoyager, WebArena, or OSWorld? These seem to be the oft-cited benchmarks when comparing computer-use agents.
9.
▲
by
timabdulla
1y ago
This is the most interesting aspect to me. I had Claude generate a guide to all the gyms in Pokemon Red and instructions for how to quickly execute a play through [0]. It obviously knows the game through and through. Yet even with encyclope
10.
▲
by
timabdulla
2y ago
There is no 3.6. There is 3.5 and 3.5 (New), both of which remain available.
11.
▲
by
timabdulla
2y ago
There was never a Sonnet 3.6. They released what is commonly known as 3.6 as "Sonnet 3.5 (New)". Then, because so many folks ended up referring to it as 3.6, they decided to call this new model 3.7, as the mental territory for 3.6
12.
▲
by
timabdulla
2y ago
My feeling (totally unproven) is that in the drive to make Sonnet 3.7 more "agentic", they've lost some of its ability to actually just stick to what you asked it to do. It seems that it "wants" (I know, it's n
13.
▲
by
timabdulla
2y ago
It's copy on write.
14.
▲
by
timabdulla
2y ago
I didn't actually give it a goal of writing any particular length, but I do think that perhaps given my not-so-large online footprint, it may have felt "pressured" to generate content that simply isn't there. It didn
15.
▲
by
timabdulla
2y ago
I tried it on a few things I was familiar with just to assess its reliability. The first was on a topic with which I am deeply familiar -- myself -- and it made three factual errors in a 500-word report: https://news.ycombinator.
16.
▲
by
timabdulla
2y ago
I just gave it a whirl. Pretty neat, but definitely watch out for hallucinations. For instance, I asked it to compile a report on myself (vain, I know.) In this 500-word report (ok, I'm not that important, I guess), it made at least th
17.
▲
by
timabdulla
2y ago
I think I hit all those points in my previous post, except for the fact that it's two different models, as you've noted. That said, neither of them seem to report scores for the other benchmark in each particular case.
18.
▲
by
timabdulla
2y ago
Those numbers are not the full story. Note that GP specifically says: "Big jumps in benchmarks from _Claude's Computer Use_ though." Claude Computer Use was not SOTA for browser tasks at the time of its release (and is still
19.
▲
by
timabdulla
2y ago
OpenAI is merely matching SOTA in browser tasks as compared to existing browser-use agents. It is a big improvement over Claude Computer Use, but it is more of the same in the specific domain of browser tasks when comparing against browser-
20.
▲
by
timabdulla
2y ago
I'm not too sure about HTMX in particular, but my Rails app's FE is just HTML and Stimulus/Turbo. I'm not sure why you think it simply "doesn't work". To me, it's a lot simpler. I use forms and links
21.
▲
by
timabdulla
2y ago
That's certainly possible. I'm not convinced AGI is just around the corner either, but I can't say with a high degree of certainty that it definitely won't arrive in the next few years.
22.
▲
by
timabdulla
2y ago
I don't think the conclusion of this article is controversial if you accept the premise: If a horizontal AI model is able to serve as a "drop-in remote worker" and all you need to do to get it going is give it access to a com
23.
▲
by
timabdulla
2y ago
Your take seems much more positive than theirs. What do you think the key differences are between your experience and the one here?
24.
▲
by
timabdulla
2y ago
Based on the author's company that be founded, I assume he believes this technology is just years away. I think with a lot of AI folk in San Francisco, this is a tacit assumption when having these sorts of conversations.
25.
▲
by
timabdulla
2y ago
I think one thing ignored here is the value of UX. If a general AI model is a "drop-in remote worker", then UX matters not at all, of course. I would interact with such a system in the same way I would one of my colleagues and I w
26.
▲
by
timabdulla
2y ago
So what percentage would you say falls to simple inability versus the other two factors you've mentioned?
27.
▲
by
timabdulla
2y ago
Right, but the branching factor increases exponentially with the scope of the work. I think it's obvious that they've cracked the formula for solving well-defined, small-in-scope problems at a superhuman level. That's an amaz
28.
▲
by
timabdulla
2y ago
What's your explanation for why it can only get ~70% on SWE-bench Verified? I believe about 90% of the tasks were estimated by humans to take less than one hour to solve, so we aren't talking about very complex problems, and to bo
29.
▲
by
timabdulla
2y ago
There's no rule against minification, which I assume is what you're referring to when you say it would make using React or Vue impossible. There's a difference between minification and obfuscation, but again, I'm not sur
30.
▲
by
timabdulla
2y ago
You can unpack and view the code of any extension after you've installed it. There's even a rule against obfuscation, though I'm not sure how enforced that is. A Chrome extension is basically a zip archive with a bunch of Jav
More ›