7 ms·
Is the benchmark measuring one-shot retrieval accuracy, or Coding agent response accuracy?
by esafranchik 4mo ago
Is the benchmark measuring one-shot retrieval accuracy, or Coding agent response accuracy?
- stephantul 4mo agoHey! Co-author here. The benchmark currently only measures retrieval accuracy. We’re interested in measuring it end to end and also optimizing, e.g. the prompt and tools, for this, but we just haven’t gotten around to it.
- esafranchik 4mo agoTwo follow-ups: 1) How do you compare accuracy? by checking if the answer is in any of the returned grep/bm25/semble snippets? 2) How do you measure token use without the agent, prompt, and tools?
- stephantul 4mo ago1) yes! It’s not accuracy, but ndcg 2) we assume that if the agent gets the correct answer in the returned snippets it does not need to read further
- esafranchik 4mo agoWouldn't NDCG/token results vary wildly depending on the agent's query and the number of returned items? e.g. agents often run `grep -m 5 "QUERY"` with different queries, instead of one big grep for all items.
- stephantul 4mo agoThe same holds for semble: the agent can fire off many different semble queries with different k/parameters. I guess the point we’re trying to make is that you need fewer semble queries to achieve the same outcome, compared to grep+readfile calls.