5 ms·
That’s a fair question, but it seems that it’s not yet necessary. See here https://dylancastillo.co/posts/pelicanmaxxing.html https://dylancastillo.co/posts/pe
by ziofill 18d ago
That’s a fair question, but it seems that it’s not yet necessary. See here
https://dylancastillo.co/posts/pelicanmaxxing.html https://dylancastillo.co/posts/pelicanmaxxing.html
https://simonwillison.net/2026/Jul/22/ https://simonwillison.net/2026/Jul/22/
- stymaar 17d agoI don't think this argument is a good one though, as it would be quite natural for a lab rhat want to macimize the performance of their model on the pelican bench to train it for “text-to-svg simple image generation” rather than just “pelicans on bicycle”.
- Anon1096 17d agoThat is the point though, if labs are maximizing svg image generation capabilities it is a very good thing. That's a general skill that is useful. So assuming they aren't specifically maximizing pelican bicycle svgs (and it doesn't look like they are) then incentives are aligned that the "benchmark" is measuring a general desirable capability.
- flexagoon 17d agohttps://xkcd.com/810/ https://xkcd.com/810/
- simonw 17d agoGemini have done exactly that. (I doubt it's because of my stupid benchmark, though!)