7 ms·
Beating OpenAI CLIP with 100x less data and compute
- juxtaposicion 4y agoIt is exciting that you could train a CLIP-style model from scratch with only 4M datapoints. But if you’ve got that data, why not fine tune a pretrained model with your 4M points? It seems likely to outperform the from-scratch method.
- vov_or 4y agoThere is not only a difference in the data source but pre-trained tasks as well. But you are right, a fine-tuned models on human-annotated data are way better than zero-shot (just pre-trained) on Image retrieval. And it is correct for CLIP, ALBEF, VICHA, and UFORM.
- riku_iki 4y agoperhaps this approach can lead to better training of foundational models?..
- vov_or 4y agoMore efficient - for sure!
- kimihailv 4y agoDid the author report metrics of the unimodal model or of the multimodal model with re-ranking?
- vov_or 4y agoThe results are reported with the multimodal model.
- swyx 4y ago> The original CLIP was trained on 500x A100 Nvidia GPUs. The latest Open_CLIP trained on 1024x GPUs. > We trained on the setup of 3x workstations, with 4x RTX 3090 consumer-grade GPUs in each, connected over 200 GBit InfiniBand HDR. ok so 85x improvement on the GPU count (i suspect even better once you take into account the differences in consumer grade GPU) but i must still be missing something - where does it say it uses 100x less data?
- alexandargyurov 4y agoAm I the only one who is very confused what this is?
- jasonjmcghee 4y agoThis is a good introduction to OpenAI CLIP, which should help provide context. https://openai.com/research/clip https://openai.com/research/clip
- pizzaknife 4y agothank you for this primer!
- margorczynski 4y agoFrom what I understand the basis for their model are these two described in these papers: https://arxiv.org/abs/2107.07651 https://arxiv.org/abs/2107.07651 https://arxiv.org/abs/2208.13628 https://arxiv.org/abs/2208.13628 Lot of tricks put together for a great final result it seems
- ashvardanian 4y agoThank you! Founder here :) You are right, those are the base papers, but we have extended the set of objectives quite significantly, tapping into modalities that haven’t been publicly CLIP-ed :) It is probably worth writing a paper about, but we are just too busy building tons of open-source stuff. Check out the GitHub org here: https://github.com/unum-cloud https://github.com/unum-cloud It is not just about the tranformers, but also about databases, networking, and improving the modern data stack for very large scale retrieval-based AI. A lot of the pieces may be pre-production, but I believe the amazing HN community may still enjoy the ways we use io_uring, SIMD, and a few other less then popular technologies.
- cosmojg 4y agoAre the pretraining and training pipelines available anywhere under a FOSS license? I'd love to take a swing at training a mid-fusion model on data other than text and images (e.g., sound, neuron spike trains, etc.)
- ashvardanian 4y agoNot yet, but you can ping our team on Discord or Twitter. They are soft like marshmallows, a couple of compliments and they will be leaking scripts left and right :)
- debdut 4y agoman I just looked at ukv, it looks to good to be true, 30x RocksDB, wtf! Hoping it's true
- 4y ago
- varispeed 4y agoI read a lot about training models and so on, but very little about inference. Let's say you came up with the custom model that gives good results, how do you transfer that model so it can be used in an API?
- alex_sf 4y agoThere's no one answer to that since different models are.. different. Beyond just modalities (text input and image output? image input and video output?), there are different common underlying tools used to build them. And then, of course, what do you mean by API? How do you want to interact with it? As a general thing, you'd take a request that would require an inference step, which would then invoke the model with some parameters and input, and return the output. Beyond that, you'd need more detail.
- deleted 4y ago[deleted]
- binarymax 4y agoI specialize in this area and build a product for self hosted inference. The challenge to support a new model architecture is about coding the preprocessing for inputs (like tokenization or image resizing and color feature extraction) and post processing the outputs (for example entity recognition needs to lookup the entities and align the text). Once an architecture is coded for the pre/post processing, then serving a new model for inference with that architecture is easy!
- sashank_1509 4y agoThey seem to be only testing for the image retrieval task, but I don’t think CLIP is actually used for image retrieval. Most cases, I see CLIP being used for semantic segmentation, detection etc. Do these guys have similar results on these tasks?
- vov_or 4y agoHi! I am one of the contributors! We were focused on image retrieval only. Almost all semantic search engines for images are based on CLIP today. We are also building a semantic multimodal search engine as a DBMS component. That is why Image retrieval is so crucial for us as well as inference perf. Also, for semantic segmentation and detection, you probably use only the image encoder part of the CLIP.
- sashank_1509 4y agoI think it’s fine you’re focused on retrieval but you should add that as a caveat to your results, 100 times better at retrieval. As an ML researcher in grad school here’s what >80% use case of clip I’ve seen: 1. Take a random image, take a random set of text (can just be categories separated by commas). CLIP will find the text that’s the best match to your image. CLIP is also incredibly robust at this, you can literally take an image with your phone and it will give you reasonable results. If you speed such a model up by 100X in inference or training, that would be a huge deal to the entire ML research community and you can expect some best paper awards (maybe even VC capital looking at stable diffusion) to come your way
- vov_or 4y agoHi! You are right that we had to clarify that "100 times better at retrieval". Btw, we have plans to tune models, evaluate, and publish results in different tasks (zero-shot ImageNet classification, etc)
- ShamelessC 4y agoIn practice CLIP can be used for many things. Originally however, the primary focus was/is indeed retrieval. This is obvious from the contrastive loss used where they minimize errors with regard to a single hard positive from a batch of thousands of known hard negatives. This is also informed by existing computer science objectives surrounding indexing, clustering of data and efficient search over data features.
- ilaksh 4y agoThis may be a dumb question, but would it be possible to apply these techniques to something like text completion and/or visual question answering? If you went ahead and used the optimizations but still scaled the model up?
- vov_or 4y agoYes, it is possible. Approaches, on which our model is based, are capable to solve VQA and other similar tasks showing SOTA results.
- freediver 4y agoDo you have/plan to have a text embeddings model?
- vov_or 4y agoYes, we are training text embedding models right now. And also have plans to open-source some of them! In addition, we train encoders for different modalities with retrieval purposes. For example, video data.
- ilaksh 4y agoDo you know anyone working on a large text completion model based on it?
- deleted 4y ago[deleted]
- sva_ 4y agoNot sure if I'm blind, but what is the number of parameters?
- vov_or 4y ago143M - English 206M - Multilingual
- fabbari 4y agoThe sample code has an error in it, it uses `model` before initializing it.
- vov_or 4y agoThanks Seems like a typo. It will be fixed soon
- bilater 4y agoFor me - the biggest thing I am looking for is a serverless vector data store. Competitors like Pinecone work just fine but they go from 0-70 as soon as you upgrade to a pod. If you can figure out pricing primarily based on usage you can capture a whole segment of this market.
- ashvardanian 4y agoGreat point! I would be happy to get more input and brain-storm a good pricing model together, one that is fair both for developers and for users. We have an source project UKV, that partly overlaps with vector-search: https://github.com/unum-cloud/ukv https://github.com/unum-cloud/ukv Another one - UNSW, is a placeholder for now: https://github.com/unum-cloud/unsw https://github.com/unum-cloud/unsw Both will be soon available on cloud marketplaces, but server-less options are a bit harder to cook. Our Discord is the best place to continue conversation: https://discord.gg/Bbh2bjNhvz https://discord.gg/Bbh2bjNhvz Thank you for advice!
- tomhamer 4y agoUKV looks really cool. At Marqo.ai we are achieving end-to-end semantic search building on top of existing embedding search functionality, keen to take a closer look at what you are doing. Seeing UKV come out on the cloud serverless will be super interesting - it's something thats taken us a while for us to work out.
- mahnerak 4y agoI could now find license in the huggingface repo, but it seems like the codebase is Apache 2.0. Are the pretrained weights / checkpoints also covered under this (or other permissive) license? In other words, can we use it for commercial purposes for free?
- vov_or 4y agoHi! Just added Apache2.0 to HF models card. Thanks!
- cosmojg 4y agoAre the pretraining and training pipelines available anywhere under a FOSS license? I'd love to take a swing at training a mid-fusion model on data other than text and images (e.g., sound, neuron spike trains, etc.)
- grammers 4y agoGood question, was about to ask the same!
- OkayPhysicist 4y agoAre weights even copyrightable under US law? It seems like they'd be the output of an automatic process (the training program) the same way the art/text produced by AI models is, which to my understanding makes them not copyrightable material.
- skybrian 4y agoCompression, even lossy compression, doesn't remove copyright. Whether this is more like compression or a more "transformative use" is something the courts will have to decide someday. It might be a good time to reread What Color Are Your Bits: https://ansuz.sooke.bc.ca/entry/23 https://ansuz.sooke.bc.ca/entry/23
- renonce 4y agoThere is a lot of manual process involved such as writing training scripts, scarping and processing training data, choosing the best weights among several runs, and spending lots of costly computation. Maybe these should make it copyrightable.
- ipsum2 4y agoHow did you deal with data contamination?
- vov_or 4y agoThe datasets we used are pretty clean themselves if we compare them with LAION. But we also filtered out images with captions on them and by CLIP's scores. Btw, huge thanks for Laion and Open_clip projects! It inspires us a lot.
- nl 4y agoThis looks interesting for image retrieval. I don't love the way their tables[1] report performance though. My understanding is that the "Dataset" column in the table represents the size of the training dataset, not the size of the dataset they are evaluating on. Note that this undersells their performance though, so it isn't like they are trying to hide something here! Also I'd love to see someone do a similar benchmark for the OpenAI CPT-3 embeddings. I'm pretty unclear how well they compare to something like FLAN-T5, because they don't seem to be evaluated anywhere in the retrieval setting (unless I've missed it?) [1] See "Zero-Shot Image Retrieval, English-only" in https://www.unum.cloud/blog/2023-02-20-efficient-multimodality https://www.unum.cloud/blog/2023-02-20-efficient-multimodali...
- vov_or 4y agoHi! MSCOCO and Flickr datasets are the main datasets for Image retrieval. The results published in most papers (including CLIP) are based on them. So we used exactly these datasets for evaluation.