6 ms·
at this point, it's pretty easy to create evals/benchmarks, and then run the latest model on them. LLMs are so easy to swap out, so having good benchmarks/eval
by peab 2mo ago
at this point, it's pretty easy to create evals/benchmarks, and then run the latest model on them.
LLMs are so easy to swap out, so having good benchmarks/evals are pretty useful.
Even then, a lot of the time the model improvements are so obvious that you don't even need an eval.