5 ms·
My hunch is that since LLMs are trained on a per word basis (okay, per-token), vacuus verbosity is overrepresented. If you have one normal sentence and one ove
by Zondartul 2y ago
My hunch is that since LLMs are trained on a per word basis (okay, per-token), vacuus verbosity is overrepresented.
If you have one normal sentence and one overly verbose, the latter will have more tokens and therefore more weight.