6 ms·
I literally say you should take benchmarks with a grain of salt :) > Of course, it's always dependent on statistical noise + host system load, and running suff
by aeneas_ory 6d ago
I literally say you should take benchmarks with a grain of salt :)
> Of course, it's always dependent on statistical noise + host system load, and running sufficiently large benchmarks is simply too expensive, so take em with a grain of salt.
And the savings listed are coming from a benchmark harness that implements different OSS bugs one time with and one without lumen - in those cases the % saved are reproducible (caveat: it was on older models, Opus 4.6 I believe).
Also I explain WHY it saves tokens - because the model doesn’t have to brute force different terms until it finds the match it needs, but uses semantic „distance“ so the embedding does it for the model.
- lucaprata 6d ago[flagged]