Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
nkov47
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
Show HN: Coarena - A community-driven arena to benchmark models for computer-use
(coarena.ai)
2 points
by
nkov47
1mo ago
|
0 comments
2.
▲
Show HN: Open-Cowork – an open-source, model-agnostic computer-use agent
(github.com)
2 points
by
nkov47
2mo ago
|
0 comments
3.
▲
by
nkov47
2mo ago
That’s exactly the failure mode we’re most worried about. A UI saying “success” is weak evidence that the underlying work is correct. We’re using a layered approach rather than relying on the computer-use model alone. The execution agent ha
4.
▲
by
nkov47
2mo ago
Exactly. We’ve found that for real production workflows, “did the task finish?” isn’t enough. You need to know what the agent saw, why it acted, what changed, and where a human approved something consequential. The checkpoints and event log
5.
▲
by
nkov47
2mo ago
I mean we've spent a year of our life on this(ironing out infra for more than half of the time), so i wouldn't say super low quality but again YC accepted us based on how our business is growing! Thanks for the feedback though!
6.
▲
by
nkov47
2mo ago
This is a great read of it, and honestly we probably did undersell that part. The handoff works through checkpoints. The deterministic workflow defines the state it expects before resuming, and the agent’s recovery job is to get the UI back
7.
▲
by
nkov47
2mo ago
Yep, this is a real failure mode, and screenshot-after-action alone does not solve it. A field can look populated while the application never commits the underlying value. We try to define verification around the actual outcome of the task
8.
▲
by
nkov47
2mo ago
Thanks!!! completely agree that this is where trust is won or lost. We support explicit human approval gates in workflows, so you can have the agent prepare everything, then pause before steps like submitting a form, sending a message, or c
9.
▲
by
nkov47
2mo ago
We've hit 82.8% on OSWorld Verfied(on the official page) and with our latest internal testing we're hitting 85.6% which we've posted ( https://github.com/coasty-ai/coasty-osworld ).
10.
▲
by
nkov47
2mo ago
We're SOC2 and HIPAA compliant and have zero data retention policies through our enterprise platform where it's all covered and we don't keep any data that you don't want us to!
11.
▲
by
nkov47
2mo ago
Yes, you have access to all your logs and screenshots and playback!
12.
▲
by
nkov47
2mo ago
With our API, you can also create a very customizable blend of deterministic workflows and AI calls for CUA, so for example, Coasty for recovery can kick in when something unexpected like a dialog box or popup happens during your determinis
13.
▲
by
nkov47
2mo ago
We don't just provide end-to-end API, we're a SOTA harness where you can bring in your own OpenAI or Claude keys and run it on and also we're the only modular API in the market - we give you the ability to control any part of
14.
▲
Launch HN: Coasty (YC S26) – An API for computer-use agents
(coasty.ai)
44 points
by
nkov47
2mo ago
|
26 comments