6 ms·
I think it's most illustrative to see the sample battles (H2H) that LMArena released [1]. The outputs of Meta's model is too verbose and too 'yappy' IMO. And lo
by ekojs 1y ago
I think it's most illustrative to see the sample battles (H2H) that LMArena released [1]. The outputs of Meta's model is too verbose and too 'yappy' IMO. And looking at the verdicts, it's no wonder by people are discounting LMArena rankings.
[1]: https://huggingface.co/spaces/lmarena-ai/Llama-4-Maverick-03-26-Experimental_battles https://huggingface.co/spaces/lmarena-ai/Llama-4-Maverick-03...