7 ms·
Yeah.. feels like we're still so early in terms of effective training and evals. Like the official evals out there that have had so many instances of just plain
by InvidFlower 1mo ago
Yeah.. feels like we're still so early in terms of effective training and evals. Like the official evals out there that have had so many instances of just plain incorrect questions. Or being incentivized to always answer instead of saying you don't know (just like advice to any human multiple choice test taker). Or the "escape hatch" in this case. There's so much money going in, but almost every day, I see "low hanging fruit" type papers where the reaction is like "really?? no one tried that before??".