Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
krawfy
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
krawfy
3y ago
How is this different from other solutions like Open Interpreter?
2.
▲
by
krawfy
3y ago
Good catch! We're looking to add function calling support very soon, and have an open issue for it on our GitHub. If you want to raise a PR and add it, we'll help you land it and get it merged
3.
▲
by
krawfy
3y ago
Thanks Neel! We totally agree that automated evals will become an essential part of production LLM systems.
4.
▲
by
krawfy
3y ago
Awesome! Let us know if there's anything from that tool that you think we should add to PromptTools
5.
▲
by
krawfy
3y ago
This is really cool! When we were trying to launch the GSPMD feature for PyTorch/XLA at Google, one of our biggest bottlenecks was network overhead, but we didn't really have any robust tools to dig into it and perform root cause
6.
▲
by
krawfy
3y ago
We've actually been in contact with the qdrant team about adding it to our roadmap! Andre (CEO) was asking for an integration. If you want to work on the PR, we'd be happy to work with you and get that merged in
7.
▲
by
krawfy
3y ago
Great question, chainforge looks interesting! We offer auto-evals as one tool in the toolbox. We also consider structured output validations, semantic similarity to an expected result, and manual feedback gathering. If anything, I've s
8.
▲
by
krawfy
3y ago
Glad you think so, we agree! If you end up trying it out, we'd love to hear what you think, and what other features you'd like to see.
9.
▲
by
krawfy
3y ago
For now, we just aggregate those across the models / prompts / templates you're evaluating so that you can get an aggregate score. You can export to CSV, JSON, MongoDB, or Markdown files, and we're working on more persis
10.
▲
Show HN: PromptTools – open-source tools for evaluating LLMs and vector DBs
(github.com)
211 points
by
krawfy
3y ago
|
24 comments