Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
jeffreyip
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
1.
▲
DeepEval Open-Sourced for TypeScript
(deepeval.com)
2 points
by
jeffreyip
1mo ago
|
0 comments
2.
▲
Show HN: Vibe code your agents without vibe coding your agent
(deepeval.com)
6 points
by
jeffreyip
4mo ago
|
0 comments
3.
▲
The Complete LLM Evaluation Playbook: How To Run LLM Evals That Matter
(confident-ai.com)
2 points
by
jeffreyip
1y ago
|
0 comments
4.
▲
DeepTeam: Penetration Testing for LLMs
2 points
by
jeffreyip
1y ago
|
0 comments
5.
▲
DeepTeam: Open-Source Pennetration Testing for LLMs
1 points
by
jeffreyip
1y ago
|
0 comments
6.
▲
Show HN: DeepTeam – Penetration Testing for LLMs
(github.com)
3 points
by
jeffreyip
1y ago
|
0 comments
7.
▲
YC helped us raise our seed round in 5 days
(confident-ai.com)
4 points
by
jeffreyip
1y ago
|
0 comments
8.
▲
by
jeffreyip
2y ago
You sure can! A few lines of code is all it takes, and a few simple rules to follow as shown here: https://docs.confident-ai.com/guides/guides-building-custom-... If you're using DSPy, you can also include it dire
9.
▲
by
jeffreyip
2y ago
Definitely, feel free to join our discord for any questions on it: https://discord.com/invite/a3K9c8GRGt
10.
▲
by
jeffreyip
2y ago
Do check it out, the early feedback has been great: https://docs.confident-ai.com/docs/metrics-dag
11.
▲
by
jeffreyip
2y ago
Hey yes would definitely love to, my contact info is in my bio, please drop me an email :)
12.
▲
by
jeffreyip
2y ago
Interesting, how are you remixing the order of questions? If we're talking about an academic benchmark like MMLU, the questions are independent of one another. Unless you're generating multiple answers in one go? Do do synthetic d
13.
▲
by
jeffreyip
2y ago
It's actually langfuse.com! Our quickstart walks you through the whole process: https://docs.confident-ai.com/confident-ai/confident-ai-intr...
14.
▲
by
jeffreyip
2y ago
That's great! Hope you enjoyed it :)
15.
▲
by
jeffreyip
2y ago
Thanks and great question! There's a ton of eval tools out there but there are only a few that actually focuses on evals. The quality of LLM evaluation depends on the quality of dataset and the quality of metrics, and so tools that are
16.
▲
by
jeffreyip
2y ago
I see, although most users come to us for evaluating LLM applications, you're correct that the academic benchmarking of foundational models is also offered in DeepEval, which I'm assuming what you're talking about. We actuall
17.
▲
Launch HN: Confident AI (YC W25) – Open-source evaluation framework for LLM apps
117 points
by
jeffreyip
2y ago
|
27 comments