5 ms·
But people who are bad at maths are unlikely to be writing about maths. A crude example might be if you search for “2+2=” in the training data, you’re much more
by bfbf 10d ago
But people who are bad at maths are unlikely to be writing about maths. A crude example might be if you search for “2+2=” in the training data, you’re much more likely to find “4” as the next character.
Obviously llms are far more complex than this, but I think this proves the point. The fact you had to add the “recent” qualifier there highlights that llms in general were bad and had to be provided with corrective targeted training data to improve. (And they still can’t count the R’s in strawberry!)
- rmunn 10d ago> And they still can’t count the R’s in strawberry! Really? I do not have the time to survey the modern LLMs to see if your assertion is correct, but if it is then I'm surprised; I would have thought that that one would have shown up so often in their training data that they would be able to answer that question, even if they would then be unable to (for example) count the R's in raspberry, or in some other word where "count the R's in _____" was not widely found in recent online discussion.
- bfbf 10d agoI mean a very quick check in ChatGPT got it wrong today, yeah! I’m not sure what model I was using, but it proves the point!
- classified 9d ago> they still can’t count the R’s in strawberry! My resident Qwythos-9B counts 3 "r"s in "strawberry". So even locally hosted models are catching up.