5 ms·
Btw, they had 3D pelican-on-the-bicycle easter egg in one of the promo videos: https://youtu.be/bOC3DisEOfg?t=117 https://youtu.be/bOC3DisEOfg?t=117 so I'm pret
by dector 13d ago
Btw, they had 3D pelican-on-the-bicycle easter egg in one of the promo videos: https://youtu.be/bOC3DisEOfg?t=117 https://youtu.be/bOC3DisEOfg?t=117 so I'm pretty sure that they spent some small amount of resources to train the model to produce good svg version as well. :D
- theturtletalks 13d agoThis is one of those benchmarks I don’t mind if they benchmax because it would mean models are actually good at creating svg graphics.
- theturtletalks 13d agoLook at how well it made this Xbox controller SVG: https://www.svgviewer.dev/s/i8t1VXzQ https://www.svgviewer.dev/s/i8t1VXzQ
- alexgoodhart 12d agoI’m not sure that this follows. And I don’t understand why the pelican bro is not purposefully demonstrating variety of svg tasks to begin with
- anthonyrstevens 9d agoBecause sequential comparison is part of the point of doing them?
- SkiFire13 12d agoNo, it would mean models are good at pelicans svgs, not svgs in general.
- theturtletalks 11d agoThat’s my point. I think if they are trying to benchmax creating a pelican riding on a bicycle SVG. In that process, they’re probably making SVG creation better and easier as a whole, even though they might just be training for that one specific SVG.
- stymaar 13d agoWhen you see how good the output of the Luna model without reasoning is compared to SotA just a year and a half ago, it's pretty clear that it's been trained on explicitly.
- maleldil 13d agoI don't see how this is proof, and not just that the model got the better. You're comparing models a year and a half apart; this is a lifetime in LLM development.
- brookst 13d agoYeah, suggestive but hardly proof
- WarmWash 13d agoMany people still believe that LLMs can only reproduce what is in their training set.
- stymaar 12d agoThat'd be too restrictive, but to get improvement in a particular domain you definitely need to train specifically for it. Just cramming more random internet text in a bigger model has stopped being a effective way of scaling since at least mid 2023.
- maleldil 12d agoNo, you don't. There have been technological advances beyond just "more random Internet text", and they lead to more powerful models with more refined emergent behaviours.
- stymaar 12d agoIt's a lifetime on things that are explicitly being trained on! But small models like Luna didn't magically become more powerful than SotA models on stuff that they weren't explicitly trained on with a dedicated RL-pipeline.
- kyorochan 12d agoThe 3D version is interesting, because it's more detailed which means more things to get wrong. The mudguards are symmetrical for some reason which you would never see in real life, and there are three brake cables but no brakes! I also think there's an extra level that I would hope an AI would nail which maybe an amateur artist would also fail at, such as thinking about what position a pelican would actually ride a bike in (maybe angling the beak down for aerodynamics etc.), but we are far away from this.