Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
ofermend
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
1.
▲
Dags are the wrong abstraction for multi-agent systems
(band.ai)
10 points
by
ofermend
5mo ago
|
0 comments
2.
▲
Ask HN: How do you handle PR density (and slop) in open source
3 points
by
ofermend
6mo ago
|
1 comments
3.
▲
by
ofermend
9mo ago
Gemini-3-flash is now on Vectara hallucination leaderboard, and rated at 13.5% grounded hallucination rate. https://github.com/vectara/hallucination-leaderboard
4.
▲
by
ofermend
9mo ago
We just evaluated Nemotron-3 for Vectara's hallucination leaderboard. It scores at 9.6% hallucination rate, similar to qwen3-next-80b-a3b-thinking (9.3%) but of course it is much smaller. https://github.com/vectara/
5.
▲
by
ofermend
9mo ago
GPT-5.2 just added to Vectara Hallucination Leaderboard. Definitely an improvement over GPT-5.1 - congrats to the team https://github.com/vectara/hallucination-leaderboard
6.
▲
by
ofermend
10mo ago
Can't wait to try Opus 4.5 We just evaluated it for Vectara's grounded hallucination leaderboard: it scores at 10.9% hallucination rate, better than Gemini-3, GPT-5.1-high or Grok-4. https://github.com/vectara/
7.
▲
by
ofermend
1y ago
If you have built AI agents in the last 6-12 months you know they fail a lot. I built this repository to be a community-curated list of failure modes, techniques to mitigate, and other resources, so that we can all learn from each other and
8.
▲
by
ofermend
1y ago
Enterprise Deep Research is like "consumer" deep research just pointed at your private data, and I think may become the "killer app" of Agentic AI for business. Lots of valuable use-cases: compliance monitoring, sales en
9.
▲
by
ofermend
1y ago
One of the biggest challenges in RAG Evaluation is the assumption that you somehow can get the "source of truth" generated, specifically the set of "golden answers" (or golden chunks/documents). In practice that is
10.
▲
by
ofermend
1y ago
Well, we expect AI to become AGI sometime in the future. Some say it's here, others say it's in 5 years or 50 years or whatever. So imagine AGI is here already (for sake of argument), and really has superintelligence, and will be
11.
▲
Trust in AI
1 points
by
ofermend
1y ago
|
3 comments
12.
▲
Shadow AI
1 points
by
ofermend
1y ago
|
0 comments
13.
▲
by
ofermend
1y ago
RAG Evaluation is difficult, primarily because it's hard to come up with "golden answers" (or golden chunks). We made Open-RAG-Eval to solve this - RAG Eval that only requires the question, yet provides great metrics for retr
14.
▲
by
ofermend
1y ago
A great day for open source, and so glad to see llama4 out. However, I'm a bit disappointed that the hallucination rates of Llama4 are not as low as I would have liked (TL;DR slightly higher than Llama3). Check the numbers on the hallu
15.
▲
by
ofermend
1y ago
This model is quite impressive. Not just useful for math/research with great reasoning, it also maintained a very low hallucination rate of 1.1% on Vectara Hallucination Leaderboard: https://github.com/vectara/hall
16.
▲
by
ofermend
1y ago
It is common these days to see in large companies multiple teams developing isolated RAG applications. This is similar to the problem of "Shadow IT" back in the early cloud era - causes a big headache to IT teams. I work at Vecta
17.
▲
by
ofermend
2y ago
DeepSeek-R1 is an amazing reasoning LLM, but it seems to hallucinate more than we might expect.
18.
▲
by
ofermend
2y ago
Gemini-2.0-Flash does extremely well on the Hallucination Evaluation Leaderboard, at 1.3% hallucination rate https://github.com/vectara/hallucination-leaderboard
19.
▲
by
ofermend
2y ago
We've done a study (see link) that shows that - unlike common belief - semantic chunking is not always the best approach. Curious to hear from the YC community - anyone else did systemic testing and if so what did you find?
20.
▲
by
ofermend
2y ago
Check out Granite 3.0 on the hallucination leaderboard: https://github.com/vectara/hallucination-leaderboard
21.
▲
by
ofermend
2y ago
We recently launched UDF reranking as part of the RAG stack, and we think this supports a lot of interesting use-cases to go beyond simple relevance. For example, it supports ranking by distance (geo-location), by recency, and more. I wante
22.
▲
by
ofermend
2y ago
I remember the Magnus/Niemann controversy from 2023 - that was quite a drama... https://en.wikipedia.org/wiki/Carlsen%E2%80%93Niemann_contro...
23.
▲
by
ofermend
2y ago
Great release. Models just added to Hallucination Leaderboard: https://github.com/vectara/hallucination-leaderboard . TL;DR: * 90B-Vision: 4.3% hallucination rate * 11B-Vision: 5.5% hallucination rate
24.
▲
by
ofermend
2y ago
About a year ago we launched in partnership with the Airbyte team the Vectara Destination, to help developers accelerate Generative AI applications - congrats on the Airbyte team on this great launch and looking forward to 2.0
25.
▲
Show HN: Vetara Agentic
(pypi.org)
2 points
by
ofermend
2y ago
|
0 comments
26.
▲
by
ofermend
2y ago
Oh, and there's a demo of an AI assistant for hacker news here: https://huggingface.co/spaces/vectara/hacker-news-chat
27.
▲
Show HN: Vectara-Agentic
(pypi.org)
1 points
by
ofermend
2y ago
|
1 comments
28.
▲
Show HN: Short course about embedding models
(learn.deeplearning.ai)
1 points
by
ofermend
2y ago
|
0 comments
29.
▲
by
ofermend
2y ago
I'm excited to try it with RAG and see how it performs (the 405B model)
30.
▲
by
ofermend
2y ago
Congrats. Very exciting to see continued innovation around smaller models, that can perform much better than larger models. This enables faster inference and makes them more ubiquitous.
More ›