Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
wujerry2000
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
1.
▲
Sims with verifiable rewards for web agent benchmarking and RL
(halluminate.ai)
1 points
by
wujerry2000
10mo ago
|
1 comments
2.
▲
by
wujerry2000
10mo ago
Hi all! Sharing some of our recent work around building RL envs and sims for agent training. There are a lot more technical details on building the benchmark in the post. If you are interested in more RL/Post-Training, I'd highly
3.
▲
by
wujerry2000
1y ago
This is a really important question. I definitely think as companies begin optimizing for an "Agent first" economy, they will start figuring out how to optimize their sites for agent traffic. They definitely could do this themselv
4.
▲
by
wujerry2000
1y ago
Lets do it! There's a cal link on our website if you wanna chat more
5.
▲
by
wujerry2000
1y ago
Self driving cars are a really good place to derive intuitions. Robotics as well! Both those spaces are still optimizing on the last mile performance gains that get exponentially harder. The good thing about computer use is building softwar
6.
▲
by
wujerry2000
1y ago
OpenAI agent is very impressive! That being said, there are still a lot of use cases its not good at, and also looking at long trajectory tasks, enterprise work tasks, etc. I imagine those are all still very nascent. I think we are still ve
7.
▲
by
wujerry2000
1y ago
I think this is totally going to be the case! AI vibe coding tools already prefer some solutions over others, probably because of training data distribution/post training preferences. This is leading to massive revenue differences and
8.
▲
by
wujerry2000
1y ago
Theses are really good questions! we share the public/consumer simulators, but we also build bespoke environments on a per customer basis (think enterprise sites or even full VMs loaded with applications and data). environment creation
9.
▲
by
wujerry2000
1y ago
Computer use agents are starting to perform well on websites/apps that are in their training distribution, but still struggle a lot when dealing with tasks outside their distribution. A big reason why is because many more niche/en
10.
▲
by
wujerry2000
1y ago
A few common ones we've heard Engineering: QA automation is huge, closes the loop on "fully automated" software engineering if another computer use system is able to click around and help identify bugs in software Deep Resear
11.
▲
by
wujerry2000
1y ago
UI refreshes knocking down simulator realism is a real issue that we're still trying to solve. I think this will probably be a mixture of automated QA/engineering and scale. Another interesting path is actually partnering directly
12.
▲
by
wujerry2000
1y ago
We agree that as a demo flight booking is probably overused. However, in talking with my AI Labs, their perspective on flight booking is a little different. "Solving" flight booking requires the AI agent to solve a LOT of hard pro
13.
▲
by
wujerry2000
1y ago
Yea haha ... early idea was illuminate + hallucinations. Naming isn't our strength :)
14.
▲
Launch HN: Halluminate (YC S25) – Simulating the internet to train computer use
70 points
by
wujerry2000
1y ago
|
46 comments
15.
▲
Most liked comment is AI
(drive.google.com)
2 points
by
wujerry2000
1y ago
|
2 comments
16.
▲
by
wujerry2000
1y ago
Running a test with browser agents and asked it to leave a comment on a Medium article. Came back 1 month later and found it had risen to be the most liked comment on the thread. Is there research around large scale Turing Tests like this?
17.
▲
by
wujerry2000
2y ago
For fun, I calculated how this stacks up against other humanity-scale mega projects. Mega Project Rankings (USD Inflation Adjusted) The New Deal: $1T, Interstate Highway System: $618B, OpenAI Stargate: $500B, The Apollo Project: $278B, Inte
18.
▲
by
wujerry2000
2y ago
My takeaways (1) Companies will probably increasingly invest in building their own evals for their use cases because its becoming clear public/allegedly private benchmarks have misaligned incentives with labs sponsoring/cheating (
19.
▲
FrontierMath was funded by OpenAI
(lesswrong.com)
483 points
by
wujerry2000
2y ago
|
199 comments
20.
▲
How are generative AI companies monitoring their systems in production?
18 points
by
wujerry2000
3y ago
|
17 comments