7 ms·
MLLMs are surprisingly bad at this out of the box and to some extent even with fine tuning. https://jina.ai/news/the-what-and-why-of-text-image-modality-gap-in-
by patricklef 2y ago
MLLMs are surprisingly bad at this out of the box and to some extent even with fine tuning. https://jina.ai/news/the-what-and-why-of-text-image-modality-gap-in-clip-models/ https://jina.ai/news/the-what-and-why-of-text-image-modality...