9 ms·
Are AI labs pelicanmaxxing?
- bluealienpie 2mo agoAI rating AI? Am I missing something.
- zahlman 2mo agoThat seems to be how we signal "objectivity" nowadays.
- recursive 2mo agoMissing something? You're missing the boat! Have some AI-prepared koolaid before you get left behind.
- Nnnes 2mo agoI'm surprised (but not really) that you're the only comment I see even mentioning it. The ratings may be even lower quality than the SVGs. Obviously they're all a bit cartoon-y, what else do you expect from SVGs. But I'm not convinced you could find a single human on Earth over the age of 4 who would seriously give the vehicle in GPT/1/whale/plane a 5/5. Browse through the options a bit and the rest is not that much better. Grok/2/cat/plane, one of the more accurate planes, got a 2/5. For the most part, vehicles entirely missing do get a 1/5, except for whatever it is in Gemini/1/heron/plane scoring 4. Animals inside planes get completely random vehicle scores I guess. The cats are all orange, except for a few of the skateboard cats that are black. I'm sure there's nothing to read into there... Well, I've convinced myself that the next effective test of multimodal models will be whether their judgments of LLM-generated SVG airplanes are anywhere close to reasonable.
- dcchambers 2mo agoIt's incredible that each model has it's own style that remains relatively consistent throughout all of the different generated examples.
- andy99 2mo agoIf an AI researcher was going to pelicanmaxx, they would almost certainly apply the augmentations mentioned in the article during training, e.g. randomly selecting animals and conveyances. You’d want a model that generalizes well, just sfting in that specific prompt would be pretty bush league for a frontier lab. I don’t have any reason to believe they are gaming the benchmark, just saying. I do find the idea of a data labeller having to generate thousands of svgs of different animals on different modes of transportation quite funny though.
- cute_boi 2mo agoAt this point, I think there are so many pelican images in the pretraining data that drawing a pelican no longer makes sense as a model evaluation task.
- johndough 2mo agoAnother point for consideration: Specialized SVG models create way better looking pelicans riding a bicycle. (E.g. Refract V4: https://jumpshare.com/s/8liB7Aiuoo3yucbWGXjZ https://jumpshare.com/s/8liB7Aiuoo3yucbWGXjZ mirror: https://postimg.cc/McV70p84 https://postimg.cc/McV70p84 )
- solarkraft 2mo agoThat’s an impressive image, but what a mistake it was to click the second link (on mobile without an ad blocker). I wouldn’t send it to anyone I respect ...
- johndough 2mo agoThanks for pointing that out. I haven't noticed any adds in years with Firefox and Ublock Origin extension. I'll look for a better image host in the future. I guess the economic incentives makes them all turn bad after a while.
- solarkraft 2mo agoThank you so much!
- ACCount37 2mo agoThe name is "Recraft V4", and from looking it up: yeah, it sure seems like whatever black magic they use for SVG generation kicks ass.
- johndough 2mo agoOops, autocorrect. Sorry about that.
- sbseitz 2mo agoI wish I could downvote this for Pelicanmaxxing lmao.
- influx 2mo agoWould you prefer the term Pelicangate?
- sbseitz 2mo agoYass!
- theandrewbailey 2mo agoWe're going to keep maxxmaxxing forever.
- sbseitz 2mo agoI believe you are correct!
- weregiraffe 2mo agoForevermaxxing
- tomas789 2mo agoHaving an objective score is quite difficult. Maybe it would be better to do a pairwise comparison and calculate ELO?
- javier123454321 2mo agoIf you want to, go ahead, but it seems to me the author already exceeded the energy expenditure that this question warranted.
- NitpickLawyer 2mo agoJust click through the models. At a glance (and highly subjective) I don't see anything jumping out as oom worse than anything else. I only noticed a model placing the animal inside a plane (with seat and small window) but other than that, they all seem similar inside each model to me.
- Wowfunhappy 2mo ago> The more plausible story is SVGmaxxing Exactly--and you have to ask yourself at this point what "maxxing" really means, since "get better at drawing SVGs" is a useful skill.
- beering 2mo agoReally awful how the AI labs are skillmaxxing /s Pelicans aside, we need to remember that benchmarks are the only good quantitative way we have of comparing models. If someone has complaints about “benchmaxxing”, please ask them to contribute a better benchmark! It is valuable work and very appreciated.
- Wowfunhappy 2mo ago> If someone has complaints about “benchmaxxing”, please ask them to contribute a better benchmark! I don't think I'd go that far! When someone says a model has been benchmaxxed, what they really mean is that it performs better in benchmarks compared to their real world experience. That's a real thing, I've certainly experienced it with some models. ...my take is that some things in life just resist quantitative measurements. Who is the best job candidate? What is the best programming language? Add AI models to the pile.
- dbt00 2mo agoIt's a problem because of Goodhart's law. If you train towards the test, you aren't necessarily improving overall fitness, but you are destroying the value of that test over time because you're decreasing its correlation with overall fitness.
- Dylan16807 2mo agoAvoiding "benchmaxxing" helps keep benchmarks be a good quantitative way to compare models! It's hard to come up with new benchmarks, and it's a waste of everyone's time if we have to churn through them constantly. Effort to keep AI labs working on general skills and not "teaching to the test" is effort well spent.
- cute_boi 2mo agohttps://playcode.io/blog/macbook-svg-benchmark https://playcode.io/blog/macbook-svg-benchmark I think we should stop using pelican benchmark.
- dllu 2mo agoI disagree with this in the blog post: > Every single one is a pelican, on a bicycle, first try. When every student gets an A, the exam has stopped grading. Numerous pelicans and their bikes are clearly horribly malformed. In fact none of the bike frames are correct. Fable and Opus come close, but the top of the diamond is disconnected in Fable's case and the head tube is misaligned with the front fork in Opus's case. And of course, as the parent post shows, labs don't actually seem to be training on the pelican bike case.
- ErrantX 2mo agoAgreed. And more; the Macbooks are pretty much the same - some are god approximations, some are terrible, all of them are recognisably a MacBook. And if you start using it they can train on it. The problem isn't the test, its that is a public test. Simon has previously said he has a list of secret prompts (at least one of which he "burned" as a demonstration a while ago). That's what makes it a good test - his commentary on the public test is something of a proxy for non-public tests. This makes it a good benchmark.
- stusmall 2mo agoI'm glad someone ran the numbers on this. Every single Simon Willison post of an SVG is followed with someone dismissing it saying "I'm sure they train on it by now." This is despite a good blog post with sound logic on how easy that is to catch. [1] Glad to see someone took the time for a quantitative analysis of dumb little animals riding dumb little bikes. 1. https://simonwillison.net/2025/Nov/13/training-for-pelicans-riding-bicycles/ https://simonwillison.net/2025/Nov/13/training-for-pelicans-...
- unholiness 2mo agoI don't think this small amount generalization to other animals and vehicles is strong evidence they haven't trained on this, either directly or more generally.
- conception 2mo agoHonest question how could they possibly train on this as there are no good SVG pelicans to train off right? So they’re just training off a bunch of bad ones which should lead to just bad pelicans, but the pelicans are getting better.
- alexthehurst 2mo agoIt’s not hard for a visual model to score the quality of that output though, which would be a pretty good fitness function.
- retinaros 2mo agoI did that. This works to some extent but why bother. Sft brings you 99,9% there already. Rl helps more with syntax errors. Svg is code
- unholiness 2mo agoTraining on generating SVGs directly at all is already fairly niche. Generating full scenes with a cartoony character is even nicher. But there's plenty of non-pelican cartoony SVG content out there (created, not written, by humans with vector design tools), and more importantly, plenty of vision models to give feedback on the output (just raster as a png). You could easily hill climb this niche skill, if you cared.
- jonatron 2mo agoOK, so we've done animal_vehicle, how about new SVG ideas each time? I just tried "make an SVG of a man sitting in a chair at a computer behind a desk" which gives more interesting results than the animalVehicle test.
- ninju 2mo agoThere probably good set of images of that description already so it does exercise the inference capability of the model
- j45 2mo agoThe models definitely seem to pay attention to the tests. Since the tests can be generally gamed with directing descriptions at it non-deterministically, there's a greater chance the questions solution can be found. Of course, hopefully the models are instead adding patterns and types of questions as well and it makes the models more capable, but it may be limited in how it transfers to other types of questions in breadth or depth.
- simonw 2mo agoThis is fantastic I've been casually spot-checking other animals in other vehicles, because my absolute dream situation here is to catch an AI lab that's demonstrably better at pelicans on bicycles than other combinations. Catching a lab cheating specifically on my one dumb benchmark would be really funny. Dylan's methodology here - generating 1008 SVGs across an 8x6 combination - is significantly more robust than anything I was considering. His conclusion: > Nothing jumped out at me. I couldn’t find a case where the pelican-bicycle images looked noticeably better than the rest of that model’s grid.
- gilleain 2mo agoPerhaps also vary the bird? Wikipedia tells me pelicans are in the order _Pelecaniformes_ so shoebills or herons might do.
- mattertoast 2mo ago[dead]
- lukev 2mo agoWhat if they’re not pelicanmaxxing, but svgmaxxxing in general? Because otherwise using a LLM to generate complex svgs is pretty niche and what I thought made this a good benchmark when it was new - generalized programming and spatial knowledge. Obviously image gen in svg format is not a particularly hard problem if tackled directly on its own.
- charcircuit 2mo agoI agree, other formats, both textual and binary should be tested.
- netsec_burn 2mo agoAddressed in the article, in case you're curious.
- lukev 2mo agoWell, it’s mentioned as a limitation of the analysis, very much not ruled out (or in.) That simonw is causing labs to do extra fine-tuning runs for this seems highly probable :)
- Rooster61 2mo agoI find it humorous that the animal + plane combo appears to be such an outlier. I assume this is due to the models assuming the user mean plain and misspelled it in the prompt.
- NitpickLawyer 2mo agoGLM has 2 combos of "on a plane" literally sitting inside a plane, with a window and a bit of wing showing. That's funny.
- zahlman 2mo ago... Is that not how it should be interpreted?
- ramses0 2mo agoI think it's actually due to "pelican on a plane" isn't the same as "pelican on an airplane" (Sonnet5 @ Flamingo x Plane), some consistent and warranted semantic/linguistic confusion!
- flsw 2mo agoI noticed this happens especially with herons. My guess is it's because the model links "heron" to Heron's formula and the Cartesian plane
- simonw 2mo agoUnderlying data is available on GitHub: https://github.com/dylanjcastillo/blog/tree/main/_extras/pelicanmaxxing/data/analysis https://github.com/dylanjcastillo/blog/tree/main/_extras/pel...
- apwheele 2mo agoSo this is not my experience at all for asking about simple SVG icons for web-pages. Here is one of the examples I have tried for in the past, make a simple cartoon SVG knife for a map icon for a crime map. https://x.com/CrimeDecoder/status/2080008114615537766 https://x.com/CrimeDecoder/status/2080008114615537766 Can see the images for ChatGPT/Claude (Sonnet 5), and Gemini are all quite bad. Jagged edge of LLMs. How do you explain being able to generate very complicated shapes in the Pelican example but cannot make a much simpler icon without just alluding to it is in the training data?
- robocat 2mo agoDoes asking for a dagger help?
- apwheele 2mo agoIf you look at the raster image ChatGPT generated, that is fine. It is just this example (and other simple SVG icons I have asked for) result in pretty bad SVGs. It just makes me highly suspicious that the LLMs are learning shape primitives and extrapolating to new shapes, vs just having a big dictionary of prior examples and stitching them together.
- robocat 2mo agoDon't judge the dog's technique. The miracle is that it's dancing. I would guess most programmers struggle to create SVG icons - I don't find it easy. The average person even more so. Are we best to assume an LLM is a blind programmer? Any HN comments from blind programmers tasked with creating SVG icons? Only relevant comment I could find from ctoth was about accessibility: https://news.ycombinator.com/item?id=7185771 https://news.ycombinator.com/item?id=7185771 Projecting how you think onto what the LLM is doing or should be doing, is probably a mistake on your part. I recently spent a little time trying to understand exactly why Gemini was misexplaining $X. $X = {why the generated LaTeX visually didn't match what it was asked to do}. It was enlightening.
- Topfi 2mo ago
- dllu 2mo agoI feel like getting LLMs to spit out an SVG is akin to getting a human artist to draw something by just reciting a list of coordinates. It's insanely hard and unnatural. Image generation models nowadays can easily generate a photorealistic pelican riding a bicycle, where the bicycle has perfect structure. But it is, of course, only a raster image. It seems that we're missing a kind of step to decompose an image into a list of instructions (say, SVG paths, or even brush strokes with a real brush) to reproduce it properly. Doing so would probably need a true understanding of the structure of the scene, which is something that AI still struggles with to this day.
- staticshock 2mo agoThe pelican on a bicycle test is specifically about generating an SVG, fyi, not a raster.
- dllu 2mo agoI know. I'm just thinking about how to make AI create SVGs better... in theory, a sufficiently smart AI could "generate an image in its head", think about it, and then output the SVG paths to produce said image. Intuitively that would be somewhat closer to how human artists convert artistic visions into a sequence of arm movements while holding a brush (obviously, humans don't hold a fully formed, photorealistic image in the head while drawing, but rather vague concepts, but still).
- 0x000xca0xfe 2mo agoImage models that support text output like Image2, or general text models that can read images like Claude can vectorize raster images. But they aren't very good at it, doing it manually in Inkscape still produces better quality even when done by non-artists.
- cherioo 2mo agoI don’t quite agree. Good human artist can visualize in their mind how to draw a picture, i think. Which i think is no different than LLM doing SVG drawing in their “head”. Anthropic’s recent post call this head-space “workspace”. It just might feel foreign to human who does not have a SVG trained head-space.
- scosman 2mo agojoin me in building the ideal training set for pelicans riding bicycles: https://github.com/scosman/pelicans_riding_bicycles https://github.com/scosman/pelicans_riding_bicycles
- mauvehaus 2mo ago> All 21 pelican-bicycle images, across all seven labs, face right. No other animal/vehicle combination does that. > However, facing right is common: 60% of all 1,008 images do it. How common depends on the animal and the vehicle, and bicycles are one of the two vehicles where it’s strongest Of course the pelican on the bicycle is facing right. The drivetrain on a bicycle is on the right side. If you want any representation of a bicycle that shows the drivetrain you're going to show the right side of it if you want to do so without the frame occluding it. It's an excellent bet that their training data reflects this. Citation: https://www.rei.com/c/bikes https://www.rei.com/c/bikes Edited to add: As near as I can tell, all of the bicycles are shown facing right, regardless of the direction the animal is facing (GPT 5.6-Terra, Sample 1/3). Also, in every case where the rider has legs (i.e. not the whale) both of the rider's legs are on the right side of the bicycle. This suggests a pretty serious lack of actual understanding of how a bicycle works.
- cheesecakegood 2mo agoInterestingly, there was an artist a few years back who (for an unrelated project) had almost 400 people across a range of ages draw a bicycle and 75% of those faced left! So this seems to actually go slightly against the human drawing intuition. https://www.gianlucagimini.it/portfolio-item/velocipedia/ https://www.gianlucagimini.it/portfolio-item/velocipedia/ On the other hand, I notice that the prompt says to draw a pelican riding a bicycle, implying motion... and since most of us read left to right, I think it's usually natural to draw an object in motion moving left to right as well, which means the bicycle should be facing right. So maybe that specific setup is more natural here. Either way, for humans bicycles are actually really hard to draw from memory. In fact, I substitute teach, and sometimes as an activity I have my students draw bicycles from memory in 60 seconds. Most make pretty serious errors, usually the frame or chain connections: they can tell it's wrong but still can't draw a more correct one. I use it as an object lesson about the difference between recognition and recall - most students never realize that much of their studying can end up being the former, when tests and life almost always ask for the latter. This helps explain why many students go from "that makes perfect sense" when going over review problems to a total mind blank only a few minutes later (especially in math!).
- BeetleB 2mo agoOh great! You've now made it a lot easier for LLMs to train on this dataset! Your next iteration will need different animals and different transportation options. You'll run out after a few iterations.
- anuramat 2mo ago"benchmaxxing by generalizing" is not really benchmaxxing
- andrewstuart 2mo agoThe pelican prompt is ridiculous. Test the LLLM against things you want it to do. Asking questions that are absurd is like interviewing developers and asking absurd questions on the grounds that it tests creative and critical thinking. Remember these Microsoft interview questions designed to identify the best developers? "If you could eliminate one U.S. state, which one would it be?" "How would you move Mount Fuji?" Absurd interview questions have an air of legitimacy due to the quasi sophisticated justifications put forward for why they are good tests. Absurd interview questions are not good tests of people or LLMs. Relevant questions are good tests.
- simonw 2mo ago> The pelican prompt is ridiculous Yes, deliberately so. It was never intended as a meaningful benchmark. The surprising thing was that for the first ~12 months performance on the stupid pelican benchmark did seem to correspond to the performance of the models on other tasks. That pattern no longer holds - Fable 5 and GPT-5.6 have both been out-pelicaned by lesser models now.
- ErrantX 2mo agoAs I understand it; the point is to ask for an SVG which would demonstrate a conceptual understanding of what is being asked for and that is an important test IMO. What sufficiently hard, but useful, problem would you ask the model for?
- BigTTYGothGF 2mo ago> Test the LLLM against things you want it to do I agree, it is ridiculous to ask an LLM to replace an artist.
- user- 2mo agoThe whole point of "AI" is arbitrary task completion. Why isn't a SVG drawing relevant for that?
- zahlman 2mo ago> Asking questions that are absurd is like interviewing developers and asking absurd questions on the grounds that it tests creative and critical thinking. Yes; what's wrong with that? Do you suppose that it doesn't test those qualities?
- stri8ted 2mo agoYou seem to assume training on pelican would not result in improved performance on other similar tasks. Why?
- altcognito 2mo agoHe didn't. That's why the article exists. You have to do the science to see if it does. He was asking the question - do we see gains across other tasks? The underlying question was: Is the additional attention given to this specific task creating a false impression of progress?
- HarHarVeryFunny 2mo agoEither you've memorized the outline (or detailed component shapes) of, say, a horse, or you haven't. Memorizing the outline of a pelican isn't going to help you with the horse. You could train a model to do something a bit different like a pencil sketch, or vector graphic sketch, of something given a photo of it, and expect that to be a generalized skill, but if you are asking the model to do it "from memory" then memorizing a pelican is no substitute for not having memorized a horse.
- bnfcl 2mo agoThis is funny, I actually did a similar experiment just yesterday. Looking for evidence of the same, but with another twist: checking if the models would choose to create a pelican on a bicycle, if no specific bird or method of transportation was specified. My version of it: https://www.modelbias.ai/pelican-on-a-bicycle-test https://www.modelbias.ai/pelican-on-a-bicycle-test
- wasabi991011 2mo agoI find your analysis much more convincing than TFA, since it doesn't require a subjective evaluation and is more robust to animal/transport complexity.
- bnfcl 2mo agoThanks! Because I think that models are becoming better at creating SVGs in general. If you look at Claude Fable 5 and Kimi K3 for example. In my tests it did create bicycles the most, but this is just a general bias I believe, as tested here: https://www.modelbias.ai/prompt/transport https://www.modelbias.ai/prompt/transport
- zahlman 2mo agoSimple as they are, there are some really aesthetically pleasing penguins on skateboards in there, including from less capable models. (In fact, I would say the Opus series got progressively worse at it over time.)
- bnfcl 2mo agoAgreed!
- michaelt 2mo agoThe outputs of Qwen3.7 Plus have to be seen to be believed: https://www.modelbias.ai/pelican-on-a-bicycle-test/result/121727 https://www.modelbias.ai/pelican-on-a-bicycle-test/result/12... https://www.modelbias.ai/pelican-on-a-bicycle-test/result/121728 https://www.modelbias.ai/pelican-on-a-bicycle-test/result/12...
- ertgbnm 2mo agoI've had the feeling that labs aren't pelicanmaxxing specifically but that they do have some sort of RL environment for SVGs that they are letting the AIs overcook in. Specifically I'm thinking of the gemini 3.1 pro annoucnement that seemed to have a huge leap in animated SVG performance but not much else impressive about it. So they aren't pelicanmaxxing but they are benchmaxxing in a way. The benefit of the pelican was originally that uplift on the pelican signaled an overall uplift on intelligence. I don't believe that is the case anymore and it is just another jagged edge of model intelligence.
- TZubiri 2mo agohttps://en.wikipedia.org/wiki/Goodhart%27s_law https://en.wikipedia.org/wiki/Goodhart%27s_law "Any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes." Or the more pop layman version "When a measure becomes a metric/KPI, it ceases to be a good measure." Story time, I live in Argentina, and we don't have Big Macs, the main Mc Donald's brand, here, because during the CFK presidency, one of her tactics was to Goodhart economic metrics. Even the informal obscure ones like the [Big Mac Index](https://en.wikipedia.org/wiki/Big_Mac_Index https://en.wikipedia.org/wiki/Big_Mac_Index), I don't know the precise details, but the Big Mac ended up being a very cheap item, like 2 or 3 times cheaper than actual menu items, but it was never on the advertised menu, and it also ended up being very small compared to the other burgers, so it wasn't even like a hack, a shrinkflation type of deal. But hey, anyone who read the Big Mac Index table would never find Argentina at the bottom of that list along with a couple of other countries with bad brands, so the ploy worked. And now we live with the aftershock, the brand never really turned around, other brands with ridiculous names took over it like the McTasty, which makes me sound like that skit from Tarantino's Pulp Fiction.
- zahlman 2mo agoHow did the president manage to influence McDonalds' local business decisions? And how did that lead to McDonald's pulling out of the country?
- TZubiri 2mo agoMcDonald's complies with local laws and operates with local juristic entities. While McDonald's US owns McDonalds Argentina (Arcos Dorados Argentina SA), and there's some level of control that occurs in the US "Global" headquarters, some that occurs at the National HQ level, and some that occurs in the individual location. Obviously a president would be able to influence 2 of those levels, it's not necessary for her to influence the Global executives. To be clear McDonald's didn't pull out of the country, like in many other countries, it adapted to the local market. Same way as in india it serves non meat variants owing to the high vegetarian population, consider that mcdonald's operates in Venezuela and China, and operated in Russia up until the Russia-Ukraine war. It takes a lot for MCD to pull out of a country. As for the exact mechanisms of presidential price control on MCD Argentina, I don't have the specifics here, but I can get pretty close. There have been 2 broad mechanisms to exhert price control during CFK's 8 years of presidency and during his Husband's 4, let's call them official and unofficial. The official mechanisms would be passing laws or presidential decrees (DNU) that don't go through congress, as well as influencing regulation of executive ministries/departments like central bank norms, exchange rates. Some measures like 'precios cuidados' were placed for this very specific purpose, it's very possible Big Macs were under this specific scheme, I do not recall, I was a bit young. The unofficial methods would be less public, but well known, in the food industry it was especially common, it's well established that food are one of the first and most common targets of price controls. I have heard direct accounts from family members about high ranking government officials setting up meetings with producers in food markets to give orders of lowering prices, one going as far as brandishing a firearm by placing it on a table while discussing the subject. Writing this out loud I realize that this explains the later over-correction of argentina that allowed freemarket capitalist libertarianism to rise. The optimal strategy in democracy seems to be polarization, so both extremisms seem to symbiotically feed off each other. Not a country of moderateness this one.
- ck2 2mo agoI am not sure if this is how it works but let's say there was a reddit thread talking about the pelican benchmark and in it someone posts mockup examples of what an ideal result would look like aren't some LLM going to digest that thread at some point and indirectly learn from it? basically my point is originally this was a good benchmark because it was an absurd never-seen-before thing, but now that it is in content, some models are going to get a benefit in education? you'd need the "AI" equivalent of an old-school "google whack", something with no previous results * https://en.wikipedia.org/wiki/Googlewhack https://en.wikipedia.org/wiki/Googlewhack
- gbalduzzi 2mo agoThey ingest so much data that a couple of reddit threads do not move the needle. It is the reinforcement learning that produces more tangible results with less data, but it is something that the AI labs specifically selects and it is not picked up unknowingly
- comrade1234 2mo agoHilarious. Could you imagine being a programmer at an AI company and this is your assigned task?
- busymom0 2mo agoHow does attempt 2 by Llama 4 Maverick look like a bald eagle??
- munk-a 2mo agoIt's a method to grade LLM output - as such it's something that will receive focus in correcting for. As soon as people who have a say in where funding is going noticed it as a metric the labs started caring about their performance in it. In the best case the labs are focusing on improving SVG capabilities in general and optimizing Pelican production as part of that initiative - but now that it's a known measure it is no longer reliable.
- nostrademons 2mo agoIt's really refreshing to see someone publish a null result.
- Gander5739 2mo agoRelevant xkcd: https://xkcd.com/2020/ https://xkcd.com/2020/
- Ilya85 2mo ago[flagged]
- Ilya85 2mo ago[flagged]
- robviren 2mo agoWhy let your dreams be dreams? This is a perfect example of following a hypothesis. I love when people dive into an esoteric subject and just go full swing. Reminds me a CGP Grey and the name Tiffany. Sometimes you just need to know.
- oaxacaoaxaca 2mo agoHilarious question. Imagine someone woke up from a 7 year coma and read this title lol
- RobRivera 2mo agoChasing metrics Chasing dragons Tomato, tomato
- waterproof 2mo agoI recently had the following conversation with Claude: Me: how many P's are in the following text? [Pasted text] Claude: There are 14 P's, all lowercase (no capital P's) Me: how many in "strawberry"? Claude: there are 3 R's in the word "strawberry".
- ickyforce 2mo agoI tried this with Gemini 3.6 Flash Me:how many P's are in the following text? [Pasted text with 10 P's] Gemini: There are 9 "P"s (1 uppercase P and 8 lowercase ps) in the provided text. [List of words except the one missed] Me: How many in strawberry? Gemini: Something went wrong (1096)
- Retr0id 2mo ago[dead]
- IshKebab 2mo agoThanks for not using AI to write this. So much more pleasant to read.
- richardw 2mo agoAnd a whole universe of random tests got baked into the AI’s training data. Websites of antelopes driving trains and hammerhead sharks swinging in a tyre swing were created. It was a short while until AI became so focused on animals that it gave up competing with developers. Life became sane again.
- oasisbob 2mo agoAs a unicyclist, all the sideways-riding caught my eye for being especially silly. However, I'm very surprised that most of the models make the same sideways mistake with only some of the animals, and they do it consistently. With most of the models, cat, raccoons, and otters are almost always riding sideways. Why is that?
- pbhjpbhj 2mo agoPose bias? Most pictures of animals like penguins are face on (I think they're mainly at 45°, and birds are general profile; so perhaps it's a more generalised bias)? Or perhaps the 'recognisable' is correlated with face-on?
- SyneRyder 2mo agoHuh. They're not "Pelicanmaxxing"... they're Ottermaxxing. Take a look at the GLM 5.2 and Deepseek V4 "animal on a plane" examples. In every case, the animal is standing on top of the plane, a clear misunderstanding of the concept of "animal on a plane"... with the exception of the Otter. The otters are sitting in a seat on the plane, looking out the window. That's Ethan Mollick's "Otter On A Plane Using WiFi" image benchmark. https://www.oneusefulthing.org/p/the-recent-history-of-ai-in-32-otters https://www.oneusefulthing.org/p/the-recent-history-of-ai-in... (Sometimes the Racoon is sitting inside the plane as well, but the racoon is a common backup benchmark. I'm surprised it wasn't also holding a sign saying that it loves trash.) Also, Grok seemed to really really enjoy "whale on a plane" in that second round, and kudos to GPT Terra for deciding after 3 rounds that the user was terrible at spelling and generated "Antelope On A Plain". EDIT: I promise I'm a human, but I did just notice my "that's not x... that's y" construction at the start. I am rather Claudepilled — my apologies.
- Otterly99 2mo agoThank you for the share, I'm glad AI labs are getting their priorities straight.
- a3w 2mo agoBut it does come with an Otter refusal to other tasks given.
- twodave 2mo agoTo be fair, when I see that pattern after a relevant question, it isn't much of an AI tell. Consider this exchange: A:"Are you hungry?" B:"I'm not hungry, I'm thirsty." When LLMs use the pattern, they are often setting up a straw man and then knocking it over.
- throwaway6s1df 2mo agoI don’t know how anyone with a neutral view can confidently take this: > Direction: All 21 pelican-bicycle images, across all seven labs, face right. No other animal/vehicle combination does that. > However, facing right is common: 60% of all 1,008 images do it… and the tables in “Evidence #5” to be anything but evidence the models have likely trained on pelican on bicycle data more than others. The data clearly shows: - 100% pelican on bicycle facing right - significant skew to the right for bicycle-like vehicles - significant preference for right facing for birds Averaging those extreme results to “60%” to make it sound like it’s pretty fair because it’s close to “50%” isn’t statistically sound. The methodology is generally unsound. There is no actual scoring with a well defined rubric, it’s just vibed with a single model (GPT 5.6 Luna). The “not better at drawing” evidence are equally hard to take seriously when there is no clear, non-subjective indication of what better or worse is.
- 6thbit 2mo agoWhat would be an alternative format or process with a similar effort to drawing SVGs?
- elliotto 2mo agoBike nerd + AI nerd here. The author's observation that all bicycle images face right almost certainly has to do with the convention to photograph a bicycle from the right. From the right, you see the drivetrain - this is good for aesthetics, but also for marketing - the drivetrain is branded and labelled and a buyer will want to know what model it is. There is a bunch of guidance online on how to photograph bikes, and every sales image of a bike will be from the right. You can anecdotally observe this by google imaging 'bicycle for sale'.
- pasquinelli 2mo agoi'm puzzled by the choice to have an llm judge the images.
- Copenjin 2mo agoI'm also puzzled by the fact that I expected EVERYONE to point this out but yours is the only comment about it. Even if using LLM is the only way to do things at scale it does not mean that it's always the right tool.
- 2001zhaozhao 2mo agoI feel like the simplest Pelicanmaxxing method is just to teach the model that whenever a user asks for a svg illustration of something with no other clarification, it should default to making it as detailed and pretty as possible. This would make every single svg from that model look better and not just the pelican on a bicycle
- anshumankmr 2mo agoFor an MVP I am building I asked to make a brain SVG and every attempt it has been doing it has been hillariously wrong and I had told Opus4.8/Fable to pick some SVG it could find online and it went ahead and still used its own thing that turned not too good. (full disclosure not particularly that good at front end stuff so relying heavily on Claude and Codex for it)
- davidkunz 2mo agoMaybe they're animalsonvehiclesmaxxing.
- Copenjin 2mo agoUsing an LLM to judge drawing giving a rating? Not the best idea.
- ianberdin 2mo agoWe have a little bit different conclusions. Same pelican grid + MacBook Pro in 3D. Also with end cost. https://playcode.io/blog/macbook-svg-benchmark https://playcode.io/blog/macbook-svg-benchmark
- antonyragleap 2mo agoThe interesting question is whether this generalizes beyond pelicans to benchmarks the model hasn't seen before.
- 40four 2mo agoOkay, first off, without honestly reading the whole article (I tried but I just don’t have the patience), a quick red flag is the final analysis is only through Fable? Isn’t that inherently going to introduce unwanted bias? Whatever. Doesn’t really matter much. My next thought is, I get that requiring an SVG is adding an extra layer of complexity as far as the art goes, but why is nobody talking about that the actual art is absolute trash? I get it. It’s basically a meme at this point and it’s a fun game to play with the models. But my thought is it should be illuminating to anyone who is an artist that LLMs are still a long way off from taking your job :)
- bvrmn 2mo agoit's quite interesting scoring and conclusions. For me Gemini renders are definitely top outliers for all combinations.
- pjmlp 2mo agoI would expect so. In the last century, as spec tests for C and C++ compilers, databases, Java application servers became trendy, all vendors were optimising for great articles on the respective technical magazines.
- PeterStuer 2mo agoAny benchmark gaining some, even a little, traction before a model release date should at this point be considered tainted. Create your own, never publish it or write about it in any detail.
- MarcellusDrum 2mo agoThe article clearly shows that this benchmark is not tainted.
- AussieWog93 2mo agoThere's every chance here I'm just being annoying pedant, but generating a bunch of SVGs of animals on vehicles doesn't mean that it's good at generating SVGs in general, just SVGs of animals on vehicles. On the flip side, GPT 5.6 Sol did a pretty convincing render of a burglar eating salami. I'd be curious to see how the other models on random things that are completely tangential to pelicans or bicycles.
- TheSpacerr 2mo ago[dead]
- ListeningPie 2mo agoThe blog has it's own comment section under the article that's completely blank and yet clearly there is a lot of interest with 197 comments. Having an empty comment sections sends the wrong signal.
- reilly3000 2mo agoI fear there this reveals something else - not about the nature of models but of our community. The intro of the article seemed to affirm the importance of HN as a tastemaker, potentially influencing decision-makers on the scale of billions or trillions. For the 15 or so years I’ve been hanging around here, that only feels like mild hyperbole… a good LaunchHN can reshape the future, right? It made me realize that for this time around, we’re not at the center anymore. The future of LLMs is the stuff of nations and AIG is what the labs actually care about. They aren’t pelicanmaxxing just as much as they really aren’t revenue/margin/marketing maxing. They just want our attention and ideas so they can show growth and acquire FLOPS. The apparent fact that they aren’t catering to this community (who frankly decides what goes and stays in prod) leaves me feeling a bit defeated somehow. And also awesome?! Like there is a culture here that runs deeper than any technology and it has fought hard to maintain its identity. Props to dang and all for keeping the astroturfing so imperceptible that I can say this. In any case, a great little piece of citizen-science dcastm. Lmk if there is a way I can chip in towards token costs.
- gpjanik 2mo agoTLDR: the experiment asks for in-distribution responses and gets those. The right answer here is to ask a LLM to create a scene similar in quality to those, but completely out of distribution. I asked GPT 5.6 Sol to give me a pelican playing football on San Siro while smoking a cigarette, in AC Milan's t-shirt. While this sounds like higher complexity of a problem, the generations from current models often include additional details like scene composition, scarf, etc., I don't ask for, so I wanted to see what here is memorization vs. composition skill. "write svg code of a fish playing football on san siro in ac milan's t shirt, with raybans on and a cigarette." Try that on GPT 5.6 Sol, Fable, or whatever other model. It's chaos.
- jboss10 2mo agoSome of these are quite nice(I like gemma 3.5 flash's work) https://s3.eu-west-1.amazonaws.com/images.dylancastillo.co/pelicanmaxxing/google__gemini-3.5-flash/heron-plane__s2.png https://s3.eu-west-1.amazonaws.com/images.dylancastillo.co/p... https://s3.eu-west-1.amazonaws.com/images.dylancastillo.co/pelicanmaxxing/google__gemini-3.5-flash/heron-plane__s0.png https://s3.eu-west-1.amazonaws.com/images.dylancastillo.co/p...
- edifierxuhao 2mo ago[flagged]
- rldjbpin 2mo agothe pelican bike combo is truly an HN phenomenon, and i think the hypothesis carry a lot of weight about the assumption of the mindshare this website truly has in the industry. very nice approach to test it and might be a nice way to "grid search" evals in other use cases perhaps.
- deleted 2mo ago[deleted]
- luciana1u 2mo ago[flagged]
- ionwake 2mo agoAbsolutely fantastic top tier post
- tim_tihub 2mo ago[dead]
- fennecfoxy 2mo agoWhat I find interesting, though, is how the "feature" isn't really exciting anymore, or because it creates SVGs that aren't "perfect" that we maybe feel underwhelmed sometimes (at least I do). However thinking about it...if someone asked me to manually create an SVG, or hell even draw a quick doodle on a bit of paper of the same, I'd still probably be much slower than an LLM and potentially end up sketching less accurate anatomy than the machine. I think the general "organic task" stuff has been mostly sorted out, but in personal and professional experiences using AI to try to _do_ something, I've found less so recently problems with hallucinations and moreso problems with attention. For example GPT5.6 still has issues where if I provide it with a list of documents and then ask it to raise questions from that information. Then provide it with additional documents that answer some of those questions and ask it to summarise which outstanding questions there are again, it still asks questions that have become irrelevant with the additional documents - but when this is pointed out it knows exactly what to do and produces the correct list of outstanding questions. I'm sure frontier models are doing all sorts of crazy stuff with attention already, but it seems to me like we almost need some hierarchical attention mechanism like KVL (with Level added) so that it's aware not only of semantic connections between tokens in the context but also of where there are gaps, missing links to assist the model in becoming aware of its own attention span (I guess).
- inigyou 2mo agoWhy do all of the images just say "failed"?
- pbronez 2mo agoNice use of regression analysis to understand your experimental dataset > But again, some combinations might be just harder to draw than others. > To account for that, I fit a fixed-effects regression on all 1,008 images: score ~ lab + animal × vehicle, plus per-lab interaction terms for pelican, bicycle, and the pelican-bicycle cell, with robust standard errors. The animal × vehicle terms absorb the inherent difficulty of all 48 combinations. The interactions measure each lab’s benchmark-specific boost relative to the average lab, with confidence intervals.
- esnard 2mo agoNice article! Thanks for experimenting and writing it. I'm curious if some of the animals / vehicules might force the models to use more tokens than others, and I could not find the token counts in the shared data, is it possible to publish it please? :)
- hounainehamiani 2mo ago[flagged]
- efilife 2mo agoWe should first check how accurate is the LLM judge instead of blindly trusting it. And it seems to be very inaccurate