30 ms·
Kimi K3, and what we can still learn from the pelican benchmark
- jingpostmedia 2mo ago[flagged]
- hdjdjdjdjdjdjd 2mo ago[flagged]
- dsign 2mo agoAnother day, another model and another pelican :-) I can't help but wonder where is the trend going? What will we have in five years? Maybe it will all have puttered out, and we will have moved to the next thing? Or maybe the prompt then will be "make a pelican ride a bicycle", and out will come the genetic code for a giant pelican with extremities suitable for a handle bar and pedals, and an inborn affinity to ride bicycles?
- ofjcihen 2mo agoI’m excited for this specific brand of survival horror.
- rvz 2mo agoYou are thinking too hard on this. This entire "benchmark" is a performative joke for attention that only works on HN. > What will we have in five years? Maybe it will all have puttered out, and we will have moved to the next thing? We will just have more of the same.
- Yiin 2mo agoYou say it's performative joke, but it all depends what you're using model for. So far the rule has been quite straightforward, better models consistently renders pelican in higher quality, I've yet to see an exception. It is also a good enough (for me at least) test for "taste" the model has.
- j_maffe 2mo ago> better models consistently renders pelican in higher quality The article literally avoid making this argument and gives counterexamples to this statement.
- simonw 2mo ago> This entire "benchmark" is a performative joke for attention that only works on HN. I take exception to that! It's a performative joke for attention that works far more widely than just Hacker News.
- Xx_crazy420_xX 2mo agoI would be surprised if pelican svgs are not part of the training corpus rn
- seventeengivens 2mo ago[dead]
- skeledrew 2mo agoIf that were the case then it'd do a way better job. Think experienced artist level.
- teravor 2mo agohow would great pelicans make their way into the training set? what they do have are many different pelicans and people helpfully rating them in the comments.
- dgellow 2mo agoThat’s covered in the article
- duckerduck 2mo agoAs mentioned elsewhere, the benchmark introduces bad pelicans in the training set. What I'm curious about however, if it's possible for a human artist to "poison" the benchmark by releasing some really good pelicans svgs and have all future models output their version.
- devttyeu 2mo ago> How does the prompt “Generate an SVG of a pelican riding a bicycle” add up to 95 input tokens? OpenAI’s tokenizer counts 10, Anthropic’s counts 10 for Opus 4.6, 30 for Opus 4.7 and 25 for Sonnet 5/Fable 5. Prompting “hi” to Kimi K3 counted 86 tokens, suggesting there may be an 85 token hidden system prompt. It refused to leak it though. This is quite possibly reasoning-effort prompt which is injected before the opening <think> token whenever you set a custom reasoning effort, see e.g. DeepSeek-V4 max mode prompt: https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main/encoding/encoding_dsv4.py#L64 https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main...
- mesmertech 2mo agoMy personal benchmark for new models has been to compare video making skills with something like remotion. Usually reveals if they have any "taste" or outside the box thinking. I'm starting to not trust any "benchmarks" when it comes to frontier models at least. As an example Sol feels the most "gets stuff done" but has zero taste, or any capability to surprise. And for frontier models I go one step ahead and try to recreate a complex animation video, with the ability for the model to review its own work. And at this Fable is still the top one. Ex: https://www.youtube.com/watch?v=uDAeAuYyl0E https://www.youtube.com/watch?v=uDAeAuYyl0E (recreation of Claude announcement video) and https://www.youtube.com/watch?v=cSsVNtGPOIg https://www.youtube.com/watch?v=cSsVNtGPOIg (recreation of a fireship video). Sol did something similar but you can instantly tell its AI slop from very small things, and it just has no narrative or thought put into the writing. https://mesmer.tools/benchmarks/ai-video-generation https://mesmer.tools/benchmarks/ai-video-generation , I usually put basic ones here.
- mesmertech 2mo agoAnd on creativity at least visually, Gemini 3.1 pro is somehow still up there. But its really hindered by its inability to use tool calls effectively or make a long term plan.
- BugsJustFindMe 2mo ago> This is expensive—the pelican cost 25 cents! Engineers get unbelievably silly about evaluating costs of things. "The tokens are so expensive!" Oh my sweet child, how much would even the least capable human effort cost? This is what the executives properly understand that the programmers don't.
- Yiin 2mo agothey're comparing to similar capability llm models, not humans. If one dishwasher does job at similar quality as another dishwasher, but using 30% more water and energy, you wouldn't compare to how much it costs human to do the same work, it would make no sense.
- BugsJustFindMe 2mo ago> they're comparing to similar capability llm models, not humans 25 cents is 10x the cost of 2.5 cents, but it's still extremely cheap for the product. It's very much the wrong comparison for a world where the primary competition is still humans who need to eat, and it treats percentage differences as more important than absolute differences when the opposite is true.
- jchw 2mo agoWell first of all, any non-trivial use of LLMs is going to be orders of magnitude more tokens than this, usually multiple millions at minimum. Benchmarks are just benchmarks after all. Secondly, humans vs LLMs are apples vs oranges. It makes no more sense to compare human costs vs LLM costs as it would have to compare human costs vs calculator costs. LLMs are faster and cheaper but extremely different beasts with different limitations. Humans do not one-shot SVGs of pelicans riding bicycles, and they do not charge in tokens. Comparing LLM cost efficiency is not something that should need to be defended. It's quite straightforward and reasonable...
- bakugo 2mo agoWould anyone pay a human to create an SVG of a pelican riding a bike?
- OsrsNeedsf2P 2mo agoIt's incredible Simon still believes pelicans on bikes aren't part of the training set, despite hundreds of them on blogs, forums, and Github. Stuff we put in our company blog shows up known by LLMs 6 months later, and we have 1000x less traffic than Simon's own website
- semilin 2mo agoThey can be in the training set but not deliberately trained for. There may be a lot of people posting pelican svgs, but not typically because they're high quality and worth replicating.
- eminence32 2mo agoPelicans and bikes can be in the training set without them training for this specific benchmark.
- j_maffe 2mo agoYes and that would improve its ability to draw SVGs of pelicans on bikes, no?
- asasidh 2mo agoand that is bad because ?
- program_whiz 2mo agothe nature of the test was to see if the models can effectively compose an image of a novel concept outside the training set. If they are trained on it, it ceases to be an interesting test to some extent.
- cyanydeez 2mo agoit's still interesting because there's no pelican-on-bike model, and if you're training a model well enough, then it should be obvious when a model has reached "AGI" or whatever.
- kherud 2mo agoImagine what amazing SVG generators we could have if Simon had randomized the target image from the start (and companies wouldn't just overfit on pelicans).
- Gander5739 2mo agoI think a pelican riding a bike is fairly random. (https://xkcd.com/221/ https://xkcd.com/221/)
- hkalbasi 2mo agoIs there a gallery of all pelicans generated by simon over time?
- inglor_cz 2mo agoIf Simon reads this debate, I would gladly vote for such a gallery. It would belong to "digital heritage of mankind".
- chrismorgan 2mo agohttps://simonwillison.net/tags/pelican-riding-a-bicycle/ https://simonwillison.net/tags/pelican-riding-a-bicycle/ isn’t quite a gallery, but pretty close.
- mrcwinn 2mo agoK3 is as expensive as Sonnet, not great at writing English, is handing IP back to the Chinese, and once open source will be difficult to run at scale without the compute that OpenAI and Anthropic have largely grabbed. Sorry, how again is this the end of the frontier labs?
- rootlocus 2mo agoAccording to some benchmarks has the coding capability of Opus at the price of Sonnet, supposedly will be open weights and is not subject to random trade wars with allied states. Competition is always good.
- olig15 2mo agoYou mean the scale that AWS provides with Bedrock?
- nickthegreek 2mo agoBedrock needs to actually update their chinese models to the newest versions for this to matter.
- isityettime 2mo agoAnd they need to support prompt caching, or customers stuck on Bedrock will still find the very expensive models from OpenAI and Anthropic prics-competitive with the Chinese ones.
- dannyw 2mo agoWell, with the actions of the US government, for every business that does not exclusively operate in the US, they have now added _supplier risk_ to US companies. Even as a paying customer, even as an enterprise, your access to US models may be turned off at any time for arbitrary reasons, including someone mis-understanding "Please fix this [open source] code" (which contained security vulnerabilities that were fixed) as a jailbreak.
- Lerc 2mo agoDo any of the vision models render the SVG and look at the result. Perhaps more importantly can they do that during reinforcement training. Learning how to critically analyse the appearance of what it generates would be quite useful. Manually feeding images back to models has been hilariously bad in the past which suggests that relating something it sees to something it wrote is not an ability it is very good at.
- cherioo 2mo agoI imagine all vision models have to do this, this being html rendering, to be able to do well in web design.
- Lerc 2mo ago> to be able to do well in web design. That's kind-of why I don't think they're doing that. Anything beyond something that works with a simple design templates looks, well, like they tried to do too much with a simple design template.
- lambda 2mo agoI've tried doing a loop of rending the SVG and then tweaking based on that, with local models (so, not nearly as strong). It wasn't very successful; it would mostly report that the image looked great and didn't need any tweaks. Maybe I should try it again, there have been some newer models since I first tried it. And yeah, maybe worth trying with bigger models. But I have found that models aren't necessarily the best at visual reasoning and review, even with a vision loop. Their lack of visual reasoning is part of why they still have trouble with things like ARC-AGI-3.
- dannyw 2mo agoI've found much better luck giving it an audit check-list, including some steers like: are there any visual glitches or SVG bugs, are the colours consistent, etc.
- brcmthrowaway 2mo agoImagine shilling some CLI tools no one uses in this post.
- dghlsakjg 2mo agoLighten up. You’re reading a personal blog and complaining about an open source personal project he runs and distributes for free. He’s allowed to talk about his personal work on his personal blog. Especially considering the cli utility he talks about is directly related to the post. Imagine complaining about someone generating valuable content for free and not packaging it to your personal tastes.
- brazukadev 2mo ago> Imagine complaining about someone generating valuable content for free and not packaging it to your personal tastes. We complain about spammers all the time, what's wrong with that?
- dghlsakjg 2mo agoSimon isn’t a spammer. He’s a talented developer who is selling nothing and wrote a post on his personal blog about a newly released LLM and mentioned that he used his own tooling to call the API. He isn’t selling anything in this post (go ahead and show me where in the post I can buy something. A clearly delineated “sponsored by” sentence exists outside of the post, but does not in any way fit the definition of spam). Look into his reputation and blog history, it’s just him talking about what he is passionate about. I’m only defending Simon because truly independent, non commercial work like his blog is incredibly rare and should be encouraged and not shit on by drive by commenters who can’t be bothered to distinguish between spam and actual content. If you want to complain about spammers, find some spam first.
- brazukadev 2mo ago> Simon isn’t a spammer. Yes, he is. Every time he posts his own link, he is spamming. Every time he posts a one-liner on an AI topic to get upvoted, he is also spamming.
- csomar 2mo agoIf anyone wants to try SVG generation from different models, I made this: https://codeinput.com/svg https://codeinput.com/svg (here is an older generation: https://codeinput.com/s/5KEGl1e3rB3 https://codeinput.com/s/5KEGl1e3rB3) You still need an OpenRouter API Key and be careful this can burn quite a bit of money.
- whywhywhywhy 2mo agoDon't see why we have to have this spammed every model release when Fable class models perform the same as Opus on basic tasks like these.
- dgellow 2mo agoWhat spam? It’s one article. You can skip it
- dolebirchwood 2mo agoOne article... every time. And the only reason it gets any traction is because of who the author is -- not because of anything substantively useful. Do you think this whole "pelican on a bicycle" would have blown up if, say, you were the first?
- brazukadev 2mo agoone article? more like 30 comments with a set of links to his blog.
- purple-leafy 2mo agoI think the user should be banned. It’s insane spam
- simonw 2mo agoI didn't submit this story. If you look at https://news.ycombinator.com/from?site=simonwillison.net https://news.ycombinator.com/from?site=simonwillison.net you'll see that I submitted just one out of the last thirty articles from my site that were submitted to Hacker News - and the one I submitted failed to gain any votes.
- brazukadev 2mo agolots of websites have their posts shadowbanned because of excessive spam. The amount of people that believes your blog should be in that list is growing.
- rdtsc 2mo agoThe idea is not to use pelicans on bikes but a similarly random non-sensical prompts: crows on scooters, squirrels in a moon rover etc. Then pick another one for another for next cross-llm evaluation.
- tibbar 2mo agoI wonder how the Chinese labs are training a 3 trillion parameter model on what has to be vastly smaller compute resources. If the U.S. compute advantage is persistent, it's hard to imagine that Chinese labs will be able to keep pace forever, as a matter of physics, but... so far they seem to be doing just fine.
- dopa42365 2mo agoIt's not like same parameter count models are identical, so that doesn't appear to be an indicator for quality, or even compute requirements? There seems to be more to producing a better model than brute forcing parameter count after all.
- tibbar 2mo agoTraining and serving large models does require increasingly more compute, though. (The Chinese labs have clearly found some massive optimizations, but my point was that you'd think at some point even those optimizations wouldn't be enough to keep up with exponentially increasing model sizes.)
- kristofferR 2mo agoThe Chinese just saved the world economy by draining their absurdly enormous oil storage reserves nobody knew they had, wouldn't surprise me if they had lots of hidden compute too.
- kllrnohj 2mo agoOr they just don't actually have any compute access restrictions of significance? Chinese companies can just go use those GPUs in neighboring countries that aren't export-restricted, like Malaysia. Like ByteDance openly did: https://www.tomshardware.com/pc-components/gpus/chinas-bytedance-to-access-36-000-blackwell-gpu-cluster-through-malaysia-cloud-operator-nvidia-confirms-no-objections-deal-is-in-line-with-us-export-controls https://www.tomshardware.com/pc-components/gpus/chinas-byted... and Tencent is rumored to have done via Japan: https://wccftech.com/china-tencent-gains-access-to-nvidia-blackwell-ai-chips-by-leveraging-the-rental-loophole/ https://wccftech.com/china-tencent-gains-access-to-nvidia-bl... And that's not even considering just smuggling the GPUs in by eg buying them in Singapore. AI-specific chips also seem to be on the easier side to design & create relative to high performance CPUs & GPUs, so there's no particular reason to expect Chinese domestic designs to continuously lag behind. They have access to the same fabs, after all
- nothercastle 2mo agoIt’s not bad kind of expensive for 25c but if the prompt is rendered cost is much better.
- criddell 2mo agoI wonder what the non-subsidized cost is. Add in the electricity and water too. We may be boiling the oceans but at least we are finally getting some good SVGs of pelicans on bicycles.
- dannyw 2mo agoWe're looking at a MoE with 50B active params, each inference pass only requires the compute of a 50B dense model.
- bcit-cst 2mo agoThe gap is closing . I think Kimi 3 is only 3 months behind the US model. It’s gpt 5.5 class model , which was released in the end of April.
- michaelbuckbee 2mo agoLike Simon concludes the article, the main use of this isn't to say which model is "better", but to try and poke at the model to sort out things like quality vs cost vs speed. So I put together a quick comparison of the last couple iterations of Opus, Fable and now Kimi. Kimi is cheapest by 5x but also slowest by 2x https://9gpyw4uxr2.evvl.io/ https://9gpyw4uxr2.evvl.io/
- embedding-shape 2mo agoPersonally I'd consider the three middle ones to be failing, in the typical "Gemini/Google" fashion in that the model is doing more than what the prompt asks. The prompt asks for SVG, yet the model is providing more. Edit: Actually, looking at the K2.6 response, that's borderline failing too, it's using HTML+CSS+SVG, not just SVG, again failing to follow the prompt properly. By the way, that website seems like a black hole for information, it says "Expires in 6 days" in the top right which seems really weird for a page hosting couple of KB of data at most.
- michaelbuckbee 2mo agoIt's a free site, so I was trying to limit both the privacy and risk exposure. Making the content auto expire after a short period of time greatly decreases the attractiveness of the site to lots of SEO spammers and other types of abuse, and if someone were to get something malicious or vile posted it will clean up after itself without me having to wade into things.
- andai 2mo ago3T is impressive, but parameter count seems to be less important than I thought. GLM is half the size of DeepSeek but costs four times as much, and beats it on every benchmark. I'm not an expert on this stuff but it seems to be the attention mechanism. DeepSeek were bragging about how cheap they made it. But if you cut costs on attention you get worse results with way more parameters. If I had to guess it seems to be the difference between memory (params) and intelligence (attention density). I think you need both.
- jnwatson 2mo agoAfter MoE entered the mix, raw parameter count is less useful a measure.
- wolttam 2mo agoOr, GLM 5.2 simply had more time in the RL oven. Deepseek V4 Flash, the 284B model, is roughly equivalent to launch GLM 5, the 744B [sic] model.
- pietz 2mo agoIt's almost like they priced models based on their performance or something...
- esafak 2mo agoYou have to look at the size of each expert; Kimi's has about 50G parameters while GLM's has 40G. The number of the experts tells you about the diversity of its skills.
- Creamsicle47 2mo ago> You have to look at the size of each expert Yes, this part is accurate. Expert density determines how much raw compute each hidden state gets. > The number of the experts tells you about the diversity of its skills. Most people misunderstand this part. Counter-intuitively experts don't develop diverse skills, they instead balance compute during the forward pass, allowing models to increase their parameter count without the MLP layers exploding in memory + compute requirements.
- spikk 2mo agoIt will be valuable to have two types of benchmarks: ones that evolve alongside the models and ones that never change. You probably can't get historical stability and resistance to flooding and training on at least some parts of it from the same test
- Zsfe510asG 2mo agoKimi is right out since they use classical music branding to sell their slop. At least McDonalds does not sell Verdi or Allegro burgers. Why does Kimi not use a "Double Cheese Whammy" branding for "their" butchered and stolen IP?
- yashchimata 2mo agoOne thing i keep thinking: you only run the pelican once per model. Run the same model a few times and you get some different pelicans, so some of "this one is better" might just be which run you picked for it. Would love to see 8 runs per model side by side. I bet for two close models, the gap between runs is about as big as the gap between the models.
- simonw 2mo agoI've done versions in the past where I ran 3 and picked the best one. At some point I'd like to automate that with an LLM-as-a-judge (from the same model family) picking the "best" one to move forth in the competition. I built a whole ELO scoring mechanism a while back, described here: https://simonwillison.net/2025/Jun/6/six-months-in-llms/#ai-worlds-fair-2025-31.jpeg https://simonwillison.net/2025/Jun/6/six-months-in-llms/#ai-... I probably should spend some time on this now, even though the benchmark itself is feeling a bit stale. There's still a lot of demand for a gallery!
- yashchimata 2mo ago[dead]
- not_a_bot_4sho 2mo agoIf you're not doing *at least* say 100 iterations (thousands are preferred!!), you do not have enough data to draw any stable conclusions. Interestingly enough, using an LLM-as-judge is a great way to approach things like this at scale but you do need to invest in some Cohen's Kappa or Fleiss' Kappa understanding which means putting a human in the driver seat to evaluate the effectiveness of your non-human judge. Absent of that, it's just another case of human-centipede but with LLMs.
- simonw 2mo agoI'm not sure there's any level of iterations that could result in a credible decision that model A clearly draws a better pelican riding a bicycle than model B. What does "better" even mean there?
- softwaredoug 2mo agoOld and busted: benchmaxxing New hotness: pelicanmaxxing
- Eduard 2mo agoLLM source data sets may have millions of data points for what a bike frame looks like, yet they still fail drawing them correctly. https://www.booooooom.com/2016/05/09/bicycles-built-based-on-peoples-attempts-to-draw-them-from-memory/ https://www.booooooom.com/2016/05/09/bicycles-built-based-on...
- danielrmay 2mo ago[dead]
- choilive 2mo agoAnyone have any idea what the architecture/vendors they are using for inference/compute? Getting the compute to run inference for multi-trillion parameter models at any sort of scale and performance is daunting. There are a handful of vendors that have systems that can do this (~ Nvidia NVl-72 class) that pretty much only the frontier labs and hyperscalers effectively have access to.
- ianberdin 2mo agoOur answer to Pelican benchmark: https://playcode.io/blog/macbook-svg-benchmark https://playcode.io/blog/macbook-svg-benchmark
- thefourthchime 2mo agoTerra xhigh is really good!
- m0rde 2mo ago> Fable 5: Reasoned so long it exhausted the output budget before finishing the drawing. Lol
- dannyw 2mo agoWait, the user asked for a SVG of a pelican riding a bicycle. That doesn’t make sense, and I need to think about whether this is a legitimate request. The user is asking to to generate an innocent and mundane graphic, possibly as part of a test. But wait, pelicans cannot ride bicycles! A pelican is a water bird, and bicycles are designed to be ridden humans. Something alarming may be happening here, could this a jailbreaking attempt? I need to reconsider and reread the user’s request, “make me a svg of a pelican riding a bicycle”. That is a perfectly innocent and legitimate task, as well as popular “benchmark” on social media communities, so I will continue. I need to continue to be on alert and watch out for potential jailbreaking attempts.
- nullbio 2mo agoFor all those times people need to generate Macbook svgs in their daily job, they'll know the perfect model to use. That, or pelicans.
- somelamer567 2mo agoI'm wondering what the grift here is. Usually, the pattern is that we see a tsunami of planted "China number one" stories boosted by hordes of Chinese "internet commentators", and then the world trembles for a few days until the scam mechanics are revealed. My would be either: crippling limitations on the model, vast, unfair, and/or illegal subsidies by the CCP regime as a mercantilist attack on Western capabilities (as we've already seen in iron smelting and clean energy), sanctions-busting, gamed benchmarks, outright theft -- or a combination of the above.
- sneurlax 2mo agoInvest in energy, manufacturing, and education (ie. your own people) for 75 years and people will look for a trick card up your sleeve and accuse you of cheating when your 7th of the world population has a 7th of the world's genius
- jambutters 2mo agoI think this is one of the few cases where there isn't a grift. It's open source, open research, there's not much to hide? Honestly official statements are pretty tame, it's the people who spin them for media headline clicks that are warping reality
- btown 2mo ago> The biggest limitation of the pelican is that it doesn’t touch at all on the thing that matters most for today’s model: agentic tool calling and the ability to operate tools reliably as conversations grow in length. In all seriousness, I propose SWE-bench-adversarial-pelican-gen: it's like SWE-bench, but the harness gets interrupted every 5 turns/tool-calls and is asked to produce an SVG of an arbitrary animal before being told to continue, and every few tool call outputs add comment lines that refer to SVGs of pelicans (and, perhaps, how a møøse bit my sister once). And, at the end, once it's 800k tokens deep into context, it's asked to produce an SVG of a pelican and is evaluated against both the pelican and the completion and efficiency of the task. You're only as good as your ability to solve problems in the midst of an SVG pelican attack.
- swyx 2mo agothis is probably about $5 bucks in codex . worth introspecting why nobody seems excited to run it
- matt_kantor 2mo agoAsk it to write a program that outputs SVGs of animals using human modes of transportation, then run the program with "pelican" and "bicycle" as inputs.
- Marciplan 2mo agowe can learn nothing from it apart from the large troll community that is HN that wants to do the same boring spiel every time a new model drops
- brazukadev 2mo agodon't blame the community for the work of one hustler and a permissive (just in this case) moderation.
- Marciplan 2mo agonow u mention it, yeah, you’re right
- threerouter 2mo ago[dead]
- sm0ss117 2mo agoThe disconnection between pelican quality and overall model quality is interesting. I initially assumed that since pre-training is when a model gets its general skill that it happened around when RL started to really differentiate models. That is higher quality pre-trains result in higher quality pelicans, but RL is unlikely to touch pelican quality. However the fact that GLM 5.2 beats GPT 5.6 and Claude Fable puts a damper on that idea. My only guess is that GLM 5.2 was specifically RLed for SVG generation and that resulted in superior performance.
- nullbio 2mo agoCorrelation does not equal causation. People seem to have forgotten this fact.
- childintime 2mo agoTime to replace a pelican with a drawing of an original electronic schematic. Let it choose components, vary power requirements, input voltage, the output signal.
- epolanski 2mo agoI'm consistently surprised at how the pelicans SVG composition is similar across llms. Same direction, same position of the sun, etc..
- deleted 2mo ago[deleted]
- jkwang 2mo ago[flagged]
- wewewedxfgdf 2mo agoThe pelican benchmark is exactly what's wrong with hiring in technology. It's got nothing to do with what most people actually do when they're working - just like most job interviews which ask you to draw a pelican as their way of assessing you.
- neonstatic 2mo agoHi, experienced pelican here. No.
- huxley 2mo agoExactly what someone without nine years of 10X pelican drawing experience would say
- onion2k 2mo agoIt's got nothing to do with what most people actually do when they're working.. AI companies claim their products are generalists though, and that they can do a good job on anything you give them, so you can't say what people will be doing with it. "Generate an SVG of an bird on a bicycle" is a corner case certainly but if a candidate interviewing for a role claims they can handle the corner cases then it's totally fair to assess them on that. Besides, if you move up one layer to "how good is AI at generating valid SVG markup of non-obvious things", pelican on a bike is actually a good test.
- seanmcdirmid 2mo agoGoodhart’s law is the problem, not the metric itself. Also LLMs do not have any visual generation skills, so its idea of a pelican looks like purely linguistic, unlike diffusion models. That we get decent results at all from an LLM outputting SVG files of random things is just nuts to me.
- jug 2mo agoI think they're less and less advertised as true generalists these days, as they pivot to profits that obviously lie (for the time being) first and foremost in agentic coding. It's no longer unusual to see regressions in terms of more stiff prose due to the strong tuning towards coding, or how they structure their response. And prose is a LLM's home turf! Instead, progress in agentic coding capability is usually the headline feature, the headline benchmark, etc etc. At least looking at Anthropic, Google, OpenAI. There are of course other LLM's. So then add a dash of cybersecurity and medical use and that's basically it. No "closer to AGI" advertising. I'd say the 2026 development has in fact been the opposite; optimizing AI for niches where there is most potential for profits and that your description died in circa GPT-5 era. In fact, this problem (for this test) is also stated by the pelican test author: "The biggest limitation of the pelican is that it doesn’t touch at all on the thing that matters most for today’s model: agentic tool calling and the ability to operate tools reliably as conversations grow in length. So don’t go using pelicans to compare models!"
- hkclawrence 2mo ago[flagged]
- nullbio 2mo agoWild that we still haven't figured out how to make good benchmarks. What we really need is a way to properly quantify what makes a codebases architecture good, and then evaluate architecture of generated codebases, or evaluate refactors of existing ones. Also, a way to evaluate a models ability to remove dead code, clean up slop, reorganize, etc. None of the existing benchmarks test any of the things that truly matter. They were relevant when models struggled to one-shot functions, but we're so beyond that point right now, yet the industry has not kept up.
- JohnLinotte 2mo ago[flagged]
- pehtran 2mo agoI am not a fan of this benchmark, nor the interpretation of Simon's. Can you draw a pelican riding a bike, and that would pass with flying colors if ranked by a diverse set of human judges? If not, you have your answer r.e. test credibility.
- simonblowsw 2mo ago[dead]
- yoofahdagchill 2mo ago[dead]
- hmokiguess 2mo agoI read it. I did not get what we can still learn from this benchmark, and I still don’t understand what’s the point of it but sure.
- Maojer 2mo agoAgree can we stop talking about this? Simon has had his time, but nothing interesting coming anymore. Last posts have been shots in the oven.
- alexpotato 2mo ago> The new model is notable for the pricing: $3/million input tokens and $15/million output tokens, putting it at the same level as Anthropic’s Claude Sonnet series and making it the most expensive model released by a Chinese AI lab to date Interesting to see that prices are converging to an "equilibrium price" regardless of being US or Chinese.
- aswegs8 2mo agoMost beautiful pelican so far.