5 ms·
Evaluating 55 LLMs with GPT-4
- habitue 3y agoShould be evaluating each prompt multiple times to see how much variance in the scores there are. Even gpt-4 grading gpt-4 should probably be done several times
- crashocaster 3y agoI always find evals of this flavor offputting given that 3.5 and 4 likely share preference models (or at least feedback data)
- aiunboxed 3y agoAny reason why palm or cohere models are not here ?
- jasonjmcghee 3y agoPalm 2 is tied for #10
- ionwake 3y agoReally cool thanks
- bradknowles 3y agoHow is this benchmark not inherently biased towards GPT? If I did the same sort of thing but used Claude to grade the tests, would I get similar results? Or would that be inherently biased towards Claude scoring high?
- londons_explore 3y agoGPT-4-0314 is top of the league table (ie. Not the latest version, but the version released in March). Is this our Concorde moment?
- natsucks 3y agoWhy no multi-turn evaluation? A lot of these benchmarks fail to capture the strength of ghost attention used in Llama 2 chat models.