Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
Snuggly73
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
Snuggly73
7mo ago
Just to nitpick the math. If you are going to fire 50% of the company, the AI tools should actually make the remaining people 100% more efficient, not 50% :)
2.
▲
by
Snuggly73
7mo ago
it has been pretty much a benchmark for memorization for a while. there is a paper on the subject somewhere. swe bench pro public is newer, but its not live, so it will get slowly memorized as well. the private dataset is more interesti
3.
▲
by
Snuggly73
8mo ago
I mean, its right there in their blog - https://cursor.com/blog/scaling-agents "We've deployed trillions of tokens across these agents toward a single goal. The system isn't perfectly efficient, but it
4.
▲
by
Snuggly73
8mo ago
if it’s too hard for you to write, it’s too hard for you to understand and comprehend. how are you going to take responsibility for that code and maintain it if needed?
5.
▲
by
Snuggly73
8mo ago
Well, could it be because it was instructed to kinda "study" Servo? https://github.com/wilsonzlin/fastrender/blob/3e5bc78b075645...
6.
▲
by
Snuggly73
8mo ago
I've watched them today work in the new repo - https://github.com/wilson-anysphere/fastrender/tree/main , adding another 50k lines trying to optimize scroll/rendering performance (spoiler: not reall
7.
▲
by
Snuggly73
8mo ago
Not sure - if it works, then who needs Cursor (and all other IDEs). You just ask for a browser and it comes out of the thin air.
8.
▲
by
Snuggly73
8mo ago
This is from the "official" build - https://imgur.com/fqGLjSA The "in progress" build has a slightly different rendering but the same result
9.
▲
by
Snuggly73
8mo ago
Noticed that as well - I think it was “manual”
10.
▲
by
Snuggly73
8mo ago
The latest commit now builds and runs (at least on my Mac). It’s tragically broken and the code is…dunno…something. 3m lines of something. I couldn’t make it render the apple page that was on the Cursor promo. Maybe they’ve used some othe
11.
▲
by
Snuggly73
8mo ago
And there is the thing about the cost. The blog post says that they've spent trillions (plural!) of tokens on that experiment. Looking at OAI API pricing, 5.2 Codex is $14 per 1 million output tokens. Which makes cool $14m for 1 trilli
12.
▲
by
Snuggly73
8mo ago
The only thing that I got to actually run on WSL2 was the "Excel" (couldnt get anything actually to compile on Mac or Windows). It a broken mess that probably implements 0.00001% of Excel. And its 1.2m locs. With codebases develop
13.
▲
by
Snuggly73
8mo ago
error: could not compile `fastrender` (lib) due to 34 previous errors; 94 warnings emitted I guess probably at some point, something compiled, but cba to try to find that commit. I guess they should've left it in a better state before
14.
▲
by
Snuggly73
8mo ago
I mean...the naive approach for a prime number check is o(n) which is linear. Probably u've meant constant time?
15.
▲
by
Snuggly73
8mo ago
Well, for some reason it doesnt let me respond to the child comments :( The problem (which should be obvious) is that with a/b real you cant construct an exhaustive input/output set. The test case can just prove the presence of a
16.
▲
by
Snuggly73
8mo ago
I mean "have been bad" doesnt exclude "getting worse" right :)
17.
▲
by
Snuggly73
8mo ago
Test cases are great, but not a total solution. Can you write a test case for the add_numbers(a, b) function?
18.
▲
by
Snuggly73
8mo ago
Yes, and for some cases no. The models are gotten very good, but I rather have an obviously broken pile of crap that I can spot immediately, than something that is deep fried with RL to always succeed, but has subtle problems that someone w
19.
▲
by
Snuggly73
8mo ago
Thanx. More of a "faster keyboard" so far then? And yeah - if I had a crystal ball, I would be on my private island instead of hanging on HN :)
20.
▲
by
Snuggly73
8mo ago
The article is arguing that it will basically replace devs. Do you think it can replace you basically one-shotting features/bugs in Zed? And also - doesn’t that make Zed (and other editors) pointless?
21.
▲
by
Snuggly73
8mo ago
Ok, if its almighty, then why is not the benchmarks at 100%? If you look at the individual issues, those are somewhat small and trivial changes in existing codebases. https://swe-rebench.com/ (note that if you look at indiv
22.
▲
by
Snuggly73
9mo ago
This type of comment implies that it’s going to stop with “them” and somehow “us that adopted the LLM” will be the winners. The goal is full automation, there is no “adapt or be left behind”.
23.
▲
by
Snuggly73
9mo ago
I am going to prefix this with that I could be completely wrong. Simon - you are an outlier in the sense that basically your job is to play with LLMs. You don't have stakeholders with requirements that they themselves don't unders
24.
▲
by
Snuggly73
10mo ago
i thought it might be something like this (still a weird overkill), but if you are effectively replacing the parser with new peg and replacing the backend with something new - then there is nothing left - just start from scratch :)
25.
▲
by
Snuggly73
10mo ago
looking at the "att" branches (excuse my unhealthy curiosity) I can only say - "jesus fucking christ". from the old parser ast -> to json -> to new ast representation (that is basically again copy of the old one) -
26.
▲
by
Snuggly73
10mo ago
reflection seems slightly wrong as well
27.
▲
by
Snuggly73
10mo ago
Neither :( LCB Pro are leet code style questions and SWE bench verified is heavily benchmaxxed very old python tasks.
28.
▲
by
Snuggly73
1y ago
Ignoring the tests, the first change was adding a single parent id column and the second "more complex" refactoring added few more hash columns to the table (after you've specified that you wanted them, i.e. not an open-ended
29.
▲
by
Snuggly73
1y ago
If you like CC - I'll just leave this here - https://github.com/sst/opencode
30.
▲
by
Snuggly73
1y ago
Not sure how it went in their tests - I've tried Opus and GPT5 and it was few lines of react + tests, so I guess 'no'
More ›