5 ms·
My benchmark is tiny compared to WorldVQA or FG-BMK, which are available, so I'd point you in that direction if you're interested in a VLM benchmark. My use-cas
by jerkstate 27d ago
My benchmark is tiny compared to WorldVQA or FG-BMK, which are available, so I'd point you in that direction if you're interested in a VLM benchmark. My use-case isn't exactly captioning as in "what is in this image?" -> caption, I am using the VLM to validate captions, as in "is this an image of [supposed subject]?" - my ranking of models I've benchmarked is gemini-3.7-flash > seed-2.1-turbo > gpt-5.6-luna > qwen-3.7-plus > qwen-3.7-flash. Gemini is almost perfect on my test dataset, only failing on some esoteric pop-culture minor celebrities and being over-specific in some cases (i.e. Q: is this [common name of fruit]? A: that's a [latin species name of fruit], not a [common name of fruit]; false). However, gemini-3.7-flash is only in my test list because openrouter has it on 75% introductory discount; otherwise it would be about 4x more expensive than seed.