Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
jing09928
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
jing09928
6d ago
模拟测试把语音 agent 的回归从“人工反复打电话”变成可重复的基础设施,这个方向很实用。实际落地时,你们更看重哪些生产指标来触发回放或新增场景:转写错误、工具调用失败,还是业务结果偏差?
2.
▲
by
jing09928
16d ago
Separating declarative facts from active execution skills definitely helps prevent the prompt-drift loop. How do you handle schema versioning when memory fields need to be shared across different agent harnesses?
3.
▲
by
jing09928
2mo ago
Useful direction, but the hard part seems to be measuring novelty after each fix. Are they reporting whether later red-team cases are genuinely distinct, or mostly variants of the same failure mode?
4.
▲
by
jing09928
3mo ago
The Vercel-for-MCP framing is useful; the hard part seems like permissions and audit trails once tools cross org boundaries. Are policies enforced per server/app, or at each tool call?
5.
▲
by
jing09928
3mo ago
The interesting bit is making cloud cost a first-class constraint for the agent loop, not just a post-hoc report. I'd be curious how you handle confidence/uncertainty in estimates, since a wrong cheap-looking recommendation can be