7 ms·
I'd make a distinction between youtubers who are incentivized to dial their reactions to the max for everything[0] and people doing silly tests but keeping thei
by senko 10d ago
I'd make a distinction between youtubers who are incentivized to dial their reactions to the max for everything[0] and people doing silly tests but keeping their expectations and reactions real (most famously, Simon's pelican test).
As to why the prompts aren't shared, for these more complex things it's most likely a somewhat messy process (ie. not a single-prompt one-shot creation; some back & forth) that would make the whole thing seem less spectacular.
Likewise for the end result - it's probably cherry-picked what works well. For example, in Claude models' resuts I always get stuck in water (something about height/jump calculations is off), where with Astra I didn't have that problem. These sorts of issues you can only spot if you try to playthrough yourself.
This is just my speculation tho. As for me, I do the tests because I'm interested in the results (easy comparison across time & models) - then I started sharing them because people asked.
I have OpenAI and Anthropic subscriptions so testing these is not an extra cost for me (I'm sloppy around recording the tokens & API-equivalent cost tho - have to improve on this). For the other models, it's total a few bucks per month or so.
So if you're careful about the cost, it's not too much, especially for a serious youtuber who's doing it for commercial reasons.
Finally regarding your comment about cloning the popular enshittified games - I don't think people are going to be doing that for testing, but if you want to have a different spin (or do a close-enough clone for yourself) on a game you loved, the modern AI systems can often deliver!
[0] from the link you posted "i am in disbelief, this is insane, bro what is this, you can't believe it's ai" - yeah...umm, it's not that good :)