7 ms·
I need someone to run actual benchmarks between the two.
by jadbox 28d ago
I need someone to run actual benchmarks between the two.
- swatcoder 28d agoBenchmarks are the BMI of model evaluation. They may have utility in trying to look at the whole landscape of models, but are very misleading when it comes to making 1:1 comparisons or in developing confidence at to how a given model will deliver on your workflow.
- dofm 27d ago> Benchmarks are the BMI of model evaluation. That is such an elegant way to put it.
- NitpickLawyer 28d agoOnly relevant benchmarks are those you make yourself, targeted specifically for your workflows. Anything else is just number go up on a pretty graph, and every model out there is probably benchmaxxed to hell on the public ones anyway. Keep yours private.
- gertlabs 28d agoThese models have gotten a fair amount of attention -- we're hoping it's enough to get them added to some reliable inference providers and OpenRouter, at which point we'll run them on our full benchmark suite.
- RomulusHill 16d agoHi Gertlabs, my inference company has actually started offering Ornith1.5 today for 9B and 35B A3B. If you are still interested, send me a message on X and I'll help you get started! https://x.com/romulushill https://x.com/romulushill https://scalattice.com/models/ornith-1.5-9b/ https://scalattice.com/models/ornith-1.5-9b/ https://scalattice.com/models/ornith-1.5-35b-a3b/ https://scalattice.com/models/ornith-1.5-35b-a3b/ https://scalattice.com/developers/ https://scalattice.com/developers/