Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
draismaa
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
draismaa
6mo ago
Every company I work with evals is on the agenda, but often takes so long to get started or it got removed again, just spinned up the first skills in the article, if this works well it’s gonna hit home, thnx for the finding and awesome work
2.
▲
by
draismaa
6mo ago
We had so many successfull stories with the LangWatch MCP server, an MCP integration that brings agent evaluation infrastructure directly into Claude Code, Cursor, and any MCP-compatible environment. That i had to share some of the successe
3.
▲
by
draismaa
7mo ago
Try: https://langwatch.ai/scenario/ Pretty amazing into running simulations at scale
4.
▲
Added OTEL Observability to OpenClaw agents full GenAI spec support
(github.com)
6 points
by
draismaa
7mo ago
|
2 comments
5.
▲
by
draismaa
7mo ago
Clawdbot has been exploding in usage over the past weeks, so we ran a hackday around it at LangWatch and quickly hit a familiar problem: great agent behavior, zero visibility into what was actually happening. This weekend contributors from
6.
▲
by
draismaa
1y ago
We open-sourced Scenario, a tiny framework to simulate and test AI agents by using another AI agent — much like how self-driving cars are tested in controlled environments before real-world deployment. The idea: if you wouldn’t deploy an au
7.
▲
by
draismaa
1y ago
Wow, they were really one of the first companies I noticed in this space, always heard very great feedback on the founder when speaking to prospects. All the best in the upcoming journey. If there are any users/customers around here, w
8.
▲
Agent simulations = unit testing for AI?
2 points
by
draismaa
1y ago
|
2 comments
9.
▲
We hit a wall testing AI agents, agents simulations works better
2 points
by
draismaa
1y ago
|
1 comments
10.
▲
Evaluations are crucial, but what should you eval on?
(github.com)
2 points
by
draismaa
2y ago
|
1 comments
11.
▲
by
draismaa
2y ago
LLM evaluations are tricky. You can measure accuracy, latency, cost, hallucinations, bias... but what really matters for your app? Instead of relying on generic benchmarks, build your own evals --> focused on your use case, and then, bri
12.
▲
Show HN: Experiment with DSPy optimzers, track performance of LLM-features
(github.com)
2 points
by
draismaa
2y ago
|
2 comments
13.
▲
by
draismaa
2y ago
Excited to introduce LangWatch, the tool designed for developers working with LLMs. It allows you to experiment with DSPy optimizers in a simple way and monitor the performance of LLM features in your projects. Key features include: DSPy op
14.
▲
by
draismaa
2y ago
Awesome to see more opensource tools in this space. In transparency we'r building the oss tool https://github.com/langwatch/langwatch which is tool for tracing and monitoring your LLM features and open telemetry i