6 ms·
video captioning without human labels or large teacher models, improving accuracy and efficiency by self-generating and refining captions automatically.
by badmonster 10mo ago
video captioning without human labels or large teacher models, improving accuracy and efficiency by self-generating and refining captions automatically.