Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
gergelycsegzi
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
Show HN: A search for meaning: retrieval methods visualized
(parsewise.ai)
3 points
by
gergelycsegzi
14d ago
|
1 comments
2.
▲
by
gergelycsegzi
14d ago
We’ve been building our startup at the intersection of retrieval and document processing. Having pitched it to customers, team members and investors, I realized that I had priors that I had not shared. Thus, I set up visualizations for the
3.
▲
by
gergelycsegzi
3mo ago
Re 1 - that is a very kind offer! Our current public template library is very limited, so let me come back to you on this. 2. We see exactly the same thing. There is a trade-off in correctness vs token burning. However, some tokens (models)
4.
▲
by
gergelycsegzi
3mo ago
Yes, we do it by having multiple stages to the pipeline. First we would extract the independent data points (from say both page 4 and 40) and a second pass step establishes relationship (we call this resolution). On the scale aspect, becaus
5.
▲
by
gergelycsegzi
3mo ago
This does indeed look really interesting. We have deterministic validations (and some deterministic excel transformations) but using more deterministic transformations for text based on traditional NLP would be a nice complement.
6.
▲
by
gergelycsegzi
3mo ago
If Claude is good enough for your use case then for sure. If you need scale, persistent structure and verifiability we can help:)
7.
▲
by
gergelycsegzi
3mo ago
Haha thanks, the reader can try and guess which is which;) We actually don't use embeddings or vector similarity, since those tend not to work well in specialist domains (e.g. for the OfficeQA benchmark where we have 90k pages talking
8.
▲
by
gergelycsegzi
3mo ago
I can see why, it's tempting to go for full automation. The reason we go for fine grained sourcing is so that people can build their awareness quickly. Plus many of our customers work in regulated industries where full automation is pr
9.
▲
by
gergelycsegzi
3mo ago
Potentially, but at that scale cost and latency may actually become an issue, so probably better to consider some sort of indexing or keyword searching.
10.
▲
by
gergelycsegzi
3mo ago
100% the really hard challenge is that the intermediate representation (ie the parquet equivalent) will be dependent on the given use case. So what we do with the platform is have the users configure the intermediate layer that serves most
11.
▲
by
gergelycsegzi
3mo ago
I'll need to check it out! We had the same observation in that the possible space is almost endless, and for example even for the same file type there may be different kind of processing required (e.g. an excel can be database style, v
12.
▲
by
gergelycsegzi
3mo ago
We were also surprised at first. The reason the models don't do so well is that they need to find information across 90k pages. When they are pointed to the right location they tend to do much better. And with these treasury documents
13.
▲
by
gergelycsegzi
3mo ago
Hey, that's exactly it!
14.
▲
by
gergelycsegzi
3mo ago
Fully agree, that's why we quite like the Databricks OfficeQA benchmark.. it made us experts on historical US treasuries haha Some screenshots in here: https://www.parsewise.ai/officeqa-sota
15.
▲
by
gergelycsegzi
3mo ago
Similar to my other comment, we assume that llamaparse and others can provide the individual page OCR. But once you have that the way that you can integrate it into your workflows often requires additional complexity around combining result
16.
▲
by
gergelycsegzi
3mo ago
Hey, good point about structure for integrated workflows:) Fully agree, for enterprises we need to guarantee types, flag discrepancies and provide underlying sources so they can integrate it downstream (whether that's Databricks, n8n e
17.
▲
by
gergelycsegzi
3mo ago
In practice we find that each domain (and even each organisation) ends up having highly customized definitions. At first, fairly generic templated definitions sort of work, but what we've seen is that over time data comes up that is ou
18.
▲
by
gergelycsegzi
3mo ago
Haha no appreciate it! That's on me for not calling it out explicitly (was trying to make the video as short as possible), but the demo UIs were literally vibe coded to show the ease of integration https://youtu.be/F1cS
19.
▲
by
gergelycsegzi
3mo ago
Planning to serve good things for sure, and appreciate your note. Ofc I didn't agree with everything Palantir was doing (also to the extent that we even knew about them at the time). I was working on vaccine distribution and cancer res
20.
▲
by
gergelycsegzi
3mo ago
Great question! 1. We are working with the assumption that OCR is (or soon will be) solved at super low prices. So if we have the extracted data, what can we do with it? Where we see Parsewise making a difference is for use cases that span
21.
▲
by
gergelycsegzi
3mo ago
I learnt a lot at Palantir, though always worked in commercial so no ties to security state (for the better or worse). (Also side-note, we are working towards enabling frontier performance with smaller open models that allows our customers
22.
▲
by
gergelycsegzi
3mo ago
"That is a great catch!"
23.
▲
by
gergelycsegzi
3mo ago
Ah probably should add a link to our website: https://www.parsewise.ai/api
24.
▲
Launch HN: Parsewise (YC P25) – Reason Across Documents with an API
56 points
by
gergelycsegzi
3mo ago
|
54 comments