Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
leerob
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
leerob
1mo ago
We plan to eventually have another model at that weight class, but right now trying to train the best possible model.
2.
▲
by
leerob
1mo ago
Probably can't advance frontier math yet, yeah. But please let us know other places you want to see Grok improve for future models!
3.
▲
by
leerob
1mo ago
We are and will continue to.
4.
▲
by
leerob
1mo ago
(I work on Grok) This isn't allowed. CSAM / deepfakes are against our acceptable use policy.
5.
▲
by
leerob
1mo ago
(I work on Grok) We've been working on teaching the model how to reason about great visual design principles. Obviously this is hard and somewhat subjective, but through a combination of writing down these principles (e.g. how to think
6.
▲
by
leerob
1mo ago
It works with the highest tier Cursor/Grok accounts during the beta.
7.
▲
by
leerob
1mo ago
It's quite a bit different, namely that ChatGPT Work has both local conversations and cloud agents. But for each cloud agent, you are spinning up and tearing down a new VM each time. This is an always-on Linux box, which stays logged i
8.
▲
by
leerob
3mo ago
You can read our full technical report here: https://cursor.com/blog/composer-2-technical-report
9.
▲
by
leerob
3mo ago
(I work at Cursor) We score well on Terminal-Bench and SWE-bench Multilingual. DeepSWE, not so great yet, as it's more for very long-horizon tasks. We're planning to include more public benchmarks in our next model release.
10.
▲
by
leerob
3mo ago
(I work at Cursor) When Composer 2.5 launched, we initially scored very competitively on AA's composite benchmark. I believe 3rd place overall. They have recently updated to use DeepSWE, which has more of a focus on very long-horizon t
11.
▲
by
leerob
3mo ago
(I work at Cursor) CursorBench includes many evals from actual engineering tasks from the Cursor team, which include our private codebase. This codebase is held-out from training so models haven't seen it, including Composer.
12.
▲
by
leerob
3mo ago
Will do. On it.
13.
▲
by
leerob
3mo ago
(I work at Cursor) Sorry about this, we should have made this more clear. The new privacy mode is needed because we have to store some state to enable running agents in the cloud. If you don't want to use cloud agents, you can continue
14.
▲
by
leerob
6mo ago
Glass was a codename while the UI was in early alpha with testers. It redirects to download now because there is no special link anymore. It's just part of Cursor 3 itself.
15.
▲
by
leerob
6mo ago
I'm an engineer at Cursor, can try to clarify questions here. > I wish they'd keep the old philosophy of letting the developer drive and the agent assist. Even when I'm using AI agents to write code, I still find myself sp
16.
▲
by
leerob
6mo ago
We used a Kimi base, with midtraining and RL on top. Going forward, we'll include the base used in our blog posts, that was a miss. Also, the license is through Fireworks: https://x.com/Kimi_Moonshot/status/20
17.
▲
by
leerob
6mo ago
Are there other coding benchmarks we should include next time? We included Teminal-Bench 2.0 and SWE-bench Mulitilingual. We don't plan on reporting SWE-bench Verified, for similar reasons to OpenAI: https://openai.com/
18.
▲
by
leerob
6mo ago
You can disable this if you want, it's under "Inline Diffs" in the Cursor settings.
19.
▲
Cursor Automations
(cursor.com)
7 points
by
leerob
7mo ago
|
0 comments
20.
▲
Cursor agents can now control their own computers
(cursor.com)
12 points
by
leerob
7mo ago
|
0 comments
21.
▲
by
leerob
7mo ago
We've found it to be a strong mix of speed and intelligence. It scores higher than Sonnet 4.5 on Terminal-Bench 2, maybe we will post more on this later.
22.
▲
Cursor Composer 1.5
(cursor.com)
20 points
by
leerob
7mo ago
|
9 comments
23.
▲
by
leerob
7mo ago
There's a setting to turn this off if you prefer (Cursor Settings > Agent > Attribution).
24.
▲
Towards Self-Driving Codebases
(cursor.com)
8 points
by
leerob
7mo ago
|
0 comments
25.
▲
by
leerob
8mo ago
> in particular the extra layer of diff-review of AI changes (red/green) which is not integrated into git We're making this better very soon! In the coming weeks hopefully.
26.
▲
by
leerob
8mo ago
(I work at Cursor) We have all these! Plan mode with a GUI + ability to edit plans inline. Todos. A tool for asking the user questions, which will be automatically called or you can manually ask for it. Hooks. And you can use Opus or any ot
27.
▲
Cursor 2.4
(cursor.com)
4 points
by
leerob
8mo ago
|
0 comments
28.
▲
by
leerob
8mo ago
> The JS engine used a custom JS VM being developed in vendor/ecma-rs as part of the browser, which is a copy of my personal JS parser project vendored to make it easier to commit to. https://news.ycombinator.com/ite
29.
▲
by
leerob
8mo ago
Should compile now: https://news.ycombinator.com/item?id=46650998
30.
▲
Cursor: Dynamic Context Discovery
(cursor.com)
2 points
by
leerob
8mo ago
|
0 comments
More ›