17 ms·
Show HN: LOL Bench – a benchmark for whether LLMs get jokes
- yakshithk_ 7d agohey, builder here. The short version: LLMs take three tests: explain why jokes work (or don't), write jokes under shared premises or predict which jokes humans prefer The finding so far that surprised me: every model aces explaining real jokes (95%+) but drops hard on explaining why a failed joke fails (81–92%). happy to answer anything about the eval system and open to any sort of feedback!
- sergey_v 7d agoThis is cool. I've always thought LLM humor understanding is an under researched, but important, topic, especially w.r.t "is AGI here yet?". I voted on a few "Which one is funnier?" choices, but honestly none of them were funny. The em dashes in every joke already set the "another LLM slop" mood. I would suggest re-writing the prompts to drop dashes. Suggestion #2: introduce basic quality checks. There was this one that talked about a 2020 event as if it was still indeterminate: "The date for Superbowl 2020 has been announced as Sunday, February 2 ... They haven't yet announced who the Patriots will be playing." Suggestion #3: slider instead of A/B single choice. Sometimes I leaned towards one choice, but it still wasn't that funny to select it; a more nuanced scale would have helped there. Thanks for building this. From the first glance, we humans are safe from machine AGI judging by joke quality today.
- yakshithk_ 7d ago[flagged]