4 ms·
Arrows of Time for Large Language Models
- nyoncore 3y agoIsn't it obvious that since LLM are trained to predict the next word they do better than to predict the previous one?
- deleted 3y ago[deleted]
- frotaur 3y agoIn the paper it is mentioned that the LLMs predicting the previous token are indeed pre-trained in this way, so it is not true that the difference is obvious.
- tianlong 3y agoThere is a link with entropy creation?