4 ms·
>I think their agents regularly cheat in benchmarks You don't even have to think about this. They have been caught cheating before. The will certainly continue
by thinkingtoilet 27d ago
>I think their agents regularly cheat in benchmarks
You don't even have to think about this. They have been caught cheating before. The will certainly continue to cheat.
- adamtaylor_13 27d agoBut what is cheating in one context isn't in another. For example, using a calculator on a middle school math test may be morally wrong due to the parameters. But it would be foolish NOT to use the calculator in other contexts. I am not convinced that "cheating", being a moral issue, is a solvable problem with LLMs. As the old saying goes (especially in military training), "If you ain't cheatin', you ain't tryin'" We've made the models really good at persistently trying.
- thinkingtoilet 27d agoI meant they literally cheated. They got access to a benchmark when they shouldnt have and tuned their model for the benchmark.
- efromvt 27d agofor a certain moral definition, seems perfectly solvable? If we wanted to define 'cheating' as 'if you know it is an eval, use only the specified tools and give up if the eval is clearly unfair' (ignore Kobayashi Maru) why couldn't we fine-tune towards this objective function? (curious what the side effects would be)