5 ms·
From the paper: "All models still underfit WebText and held-out perplexity has as of yet improved given more training time."
by vedant 8y ago
From the paper:
"All models still underfit WebText and held-out perplexity has as of yet improved given more training time."