7 ms·
The tasks are the thing to really look at here: https://github.com/harbor-framework/terminal-bench-science/tree/main/tasks https://github.com/harbor-framework/
by j_maffe 20d ago
The tasks are the thing to really look at here:
https://github.com/harbor-framework/terminal-bench-science/tree/main/tasks https://github.com/harbor-framework/terminal-bench-science/t...
- walrus01 20d agoI am hoping someone with more free time than myself can contribute some things in the RF engineering domain in the 'engineering-sciences' section. There's some problems out there that will definitely stump even a smart LLM.
- daveguy 19d agoLooks like most things definitely stump even a "smart LLM"... Best score on this is 30%. Which is what you should assume for tasks you give an LLM if they aren't exactly the same as an existing benchmarked task. They're just not that good for the purposes people seem to think they are. Very limited application space.