6 ms·
These are ~1000-word stories (prompts at [1]), squarely in the sweet spot where planning logistics are minimal, context windows comfortably fit everything, and
by smallmancontrov 8d ago
These are ~1000-word stories (prompts at [1]), squarely in the sweet spot where planning logistics are minimal, context windows comfortably fit everything, and the appropriate level of abstraction omits any details which could glaringly reveal weak world-model knowledge.
SOTA LLMs from 2 years ago would saturate this study, and while modern LLMs are amazing and I suspect they would do quite well on a similar but more challenging study that forces them out of said comfort zone, this isn't that.
[1] https://www.cambridge.org/core/journals/judgment-and-decision-making/article/bot-or-not-can-people-tell-the-difference-between-stories-written-by-a-human-or-by-an-ai-system/45E6DC0BB90AA648654D5AE243F6C667 https://www.cambridge.org/core/journals/judgment-and-decisio...