6 ms·
You have hundreds of hours with a model that was barely even released hundreds of hours ago? The perception of capability varies greatly between task. For my n
by aroman 2mo ago
You have hundreds of hours with a model that was barely even released hundreds of hours ago?
The perception of capability varies greatly between task. For my needs for example sol xhigh consistently outperforms fable xhigh.
- jayd16 2mo agoYou could run hundreds of agents in parallel and arguably that counts.
- aroman 2mo agoNo, it wouldn’t. The hours in question are human experience, not that of the agent.
- jayd16 2mo agoWhy not? You have far more results to review. If you run a model on slower hardware are you getting more experience? Surely its a factor of model output reviewed and not human time.
- aroman 2mo agoRight, it's about how much time the human spent - the time spent by the machine itself is irrelevant. As you rightly point out: that is why we measure programmer experience by wall clock time, not CPU time :)