5 ms·
I built two related benchmarks this month: https://github.com/lechmazur/sycophancy https://github.com/lechmazur/sycophancy and https://github.com/lechmazur/pers
by zone411 6mo ago
I built two related benchmarks this month: https://github.com/lechmazur/sycophancy https://github.com/lechmazur/sycophancy and https://github.com/lechmazur/persuasion https://github.com/lechmazur/persuasion. There are large differences between LLMs. For example, good luck getting Grok to change its view, while Gemini 3.1 Pro will usually disagree with the narrator at first but then change its position very easily when pushed.