Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
vov_or
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
1.
▲
Show HN: UForm v2 – tiny CLIP-like embeddings in 21 languages and Graphcore API
(github.com)
16 points
by
vov_or
3y ago
|
1 comments
2.
▲
From Dating to Vector Search – “Stable Marriages” on a Global Scale
(ashvardanian.com)
36 points
by
vov_or
3y ago
|
31 comments
3.
▲
A Theory on Adam Instability in Large-Scale Machine Learning
(arxiv.org)
140 points
by
vov_or
3y ago
|
51 comments
4.
▲
Abusing vector search for texts, maps, and chess
(ashvardanian.com)
110 points
by
vov_or
3y ago
|
23 comments
5.
▲
MultiModal-GPT: A Vision and Language Model for Dialogue with Humans
(github.com)
4 points
by
vov_or
3y ago
|
1 comments
6.
▲
by
vov_or
3y ago
Guys trained a multi-modal chatbot with visual and language instructions based on the open-source multi-modal model OpenFlamingo! Paper link: https://arxiv.org/abs/2305.04790
7.
▲
by
vov_or
4y ago
Hi! MSCOCO and Flickr datasets are the main datasets for Image retrieval. The results published in most papers (including CLIP) are based on them. So we used exactly these datasets for evaluation.
8.
▲
by
vov_or
4y ago
Hi! You are right that we had to clarify that "100 times better at retrieval". Btw, we have plans to tune models, evaluate, and publish results in different tasks (zero-shot ImageNet classification, etc)
9.
▲
by
vov_or
4y ago
More efficient - for sure!
10.
▲
by
vov_or
4y ago
The datasets we used are pretty clean themselves if we compare them with LAION. But we also filtered out images with captions on them and by CLIP's scores. Btw, huge thanks for Laion and Open_clip projects! It inspires us a lot.
11.
▲
by
vov_or
4y ago
It will take some time, but yes, we have this in our plans.
12.
▲
by
vov_or
4y ago
Hi! Just added Apache2.0 to HF models card. Thanks!
13.
▲
by
vov_or
4y ago
Yes, we are training text embedding models right now. And also have plans to open-source some of them! In addition, we train encoders for different modalities with retrieval purposes. For example, video data.
14.
▲
by
vov_or
4y ago
Thanks Seems like a typo. It will be fixed soon
15.
▲
by
vov_or
4y ago
143M - English 206M - Multilingual
16.
▲
by
vov_or
4y ago
There is not only a difference in the data source but pre-trained tasks as well. But you are right, a fine-tuned models on human-annotated data are way better than zero-shot (just pre-trained) on Image retrieval. And it is correct for CLIP,
17.
▲
by
vov_or
4y ago
Yes, it is possible. Approaches, on which our model is based, are capable to solve VQA and other similar tasks showing SOTA results.
18.
▲
by
vov_or
4y ago
Hi! I am one of the contributors! We were focused on image retrieval only. Almost all semantic search engines for images are based on CLIP today. We are also building a semantic multimodal search engine as a DBMS component. That is why Imag
19.
▲
by
vov_or
4y ago
There are also dataset sizes for Albef and ViCHA.
20.
▲
by
vov_or
4y ago
The results are reported with the multimodal model.
21.
▲
Beating OpenAI CLIP with 100x less data and compute
(unum.cloud)
342 points
by
vov_or
4y ago
|
57 comments
22.
▲
UKV: Modular Transactional NoSQL DBMS Bringing Zero-Copy Semantics to Storage
(github.com)
9 points
by
vov_or
4y ago
|
0 comments