6 ms·
One way to interpret these results is that the LLMs tested are badly calibrated for this kind of multi-armed bandit problem. Even if the intent is for the model
by aesthesia 8d ago
One way to interpret these results is that the LLMs tested are badly calibrated for this kind of multi-armed bandit problem. Even if the intent is for the model to find and exploit patterns, it's bad at doing it (or rather, at recognizing that there is not in fact any pattern).
- aaron695 8d ago[dead]
- vintermann 8d agoIt may be bad at recognizing it, but if all arms are equally good, that doesn't matter.