Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
antonap
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
Show HN: Agent Contracts – a new way to trust agents
(github.com)
2 points
by
antonap
2y ago
|
0 comments
2.
▲
Building the same AI agents with 3 frameworks: LangGraph vs. CrewAI vs. Swarm
(relari.ai)
3 points
by
antonap
2y ago
|
0 comments
3.
▲
by
antonap
2y ago
Thanks for the feedback, love your article diving deep into DSPy! Here's how our platform is different: 1. You are absolutely right, the dataset is a big hurdle for using DSPy. That's why we offer a synthetic dataset generation pi
4.
▲
by
antonap
2y ago
Thanks for bringing this up! The best thing is to see how we can make the enterprise plan work for you, feel free to reach out to us (founders@relari.ai).
5.
▲
Show HN: Relari – Auto Prompt Optimizer as Lightweight Alternative to Finetuning
32 points
by
antonap
2y ago
|
4 comments
6.
▲
by
antonap
3y ago
That's correct, but let me dig a little deeper. Continuous-eval provides two types of metrics, reference-based and reference-free metrics. In the case of reference-based metrics, you provide a dataset with the input/expected outpu
7.
▲
by
antonap
3y ago
Originally, we were going to do a Show HN for the modular evaluation and another Show HN for the synthetic data, because our understanding is that the Show HNs are for individual projects. But then we realized that it's the combination
8.
▲
by
antonap
3y ago
Great suggestion, thanks!
9.
▲
by
antonap
3y ago
Arize is a great tool for observability, and their open source product, Phoenix, offers many great features for LLM evaluation as well. Some key unique advantages we offer: - Component-level evaluation, not just observability: Many great to
10.
▲
by
antonap
3y ago
Thank you for the feedback! That’s a great suggestion. We do want to make the demo into a separate page, and also add a live evaluation demo using the synthetic data generated on the fly.
11.
▲
by
antonap
3y ago
Thank you for catching that! Looking into it now.
12.
▲
by
antonap
3y ago
Indeed - decomposition improves reliability but also makes the testing more challenging. That’s why we made the framework modular! Let us know of any feedback as you try it out!
13.
▲
Launch HN: Relari (YC W24) – Identify the root cause of problems in LLM apps
106 points
by
antonap
3y ago
|
15 comments
14.
▲
by
antonap
3y ago
Yup that’s what drove us to work on a module-level framework since our app is made up of many non-deterministic components. Give it a spin and let us know what you think!
15.
▲
Show HN: Continuous-eval – Granular evaluation of GenAI pipelines
(github.com)
10 points
by
antonap
3y ago
|
2 comments