7 ms·
Show HN: Frugon – Find which LLM calls a cheaper model could handle (local, MIT)
I started leaning in on AI heavily this year, as I wanted to get more done autonomously, but then my token usage climbed dramatically to the point where my weekly quota would run out before the end of the week, sometimes a couple of days into the week.
I realised I had to do something about it else I'd have to double my spend. So I decided to start tracking my cost per task type. This revealed that a lot of my spend went to searches/scans or simple things like scouting tasks.
I then decided to turn this into a simple CLI tool that can be used to read your OpenAI-style logs locally, and analyze the cost and compare this spend to other models, then show you how much you could potentially save by switching those calls to a cheaper model.
When you run analyze you get an offline estimate priced against LiteLLM and gated by LMArena tiers. The general savings bands come from the research published by RouteLLM; but you can confirm this yourself using 2 commands --measure (shows the prompt-response output side by side) and --judge (a model chosen to do the comparisons). These send a sample of the prompts from the logs to the candidate models - either the default choice or set by you. This call goes directly to the model provider (never through me) as any normal LLM call would, and the response is shown and judged to either be better or worse or a tie.
It's deliberately small, because I tend to over complicate/think things sometimes: analyze + capture + a few commands, doing three jobs. Cost, quality visibility, routing recommendation.
Nothing is hosted. capture is an optional local proxy on your own machine, and there's no endpoint in the path of your data. You can confirm this by checking the source.
I included a demo so you can check out the output. It has a synthetic 56k call log (a month's worth) showing how costs can drop from $549.46 to $343.91 a month. A 37.4% saving.
Try it:
uvx frugon analyze --demo
or
uv tool install frugon
Then point it at your own logs.
All feedback is welcome, especially any on the routing/quality logic, or anything else, good or bad.
- deleted 2mo ago[deleted]
- arspesk 2mo ago[flagged]
- cyanydeez 2mo agothis would be more interesting as a local LLM anlysis; throw out all the costs, and figure out primary-subagent model architecture, and maximize token generation and prefill. I don't see how anyone can operationalize this information.
- jarodrh 2mo agoThat's an interesting thought, and one I'll take note of, but that would be a different tool. However, if you look hard enough I'd say we're tackling the same issue. I'm just choosing to look at the problem from a cost perspective as opposed to a raw token generation/prefill perspective. The mindset can be applied to both sides, but Frugon prices the cloud side. The end goal is essentially the same and your mind went to that point, "primary-subagent model architecture". This is what the tool helps you figure out. It's not there to hold your hand and explain what your architecture should be, as I wanted it to be a small simple tool that would give the user insight into triggering your exact thought process. The thought process of breaking down their tasks by type, mapping that to individual models regardless of app or dev tool, regardless of cloud or local (the thought process transfers). It shows the user that they could route a portion of their calls to a cheaper model. It's then up to the user to understand the task type and point those calls to different models. To directly answer the operationalization statement, it depends on what the call log is from (app/harness). If harness, then the direction would be to pin models per role per task type; the routing recommendation maps directly to that, and measure/judge allows you to verify this before switching. If app, then this is a similar shape as above, where you would then categorise and pin those calls identified in the recommendation to the model recommended, and as above, measure/judge allows you to verify before switching. You're proposing an optimisation for throughput whereas Frugon is for spend. Same issue, just different lens.
- cyanydeez 2mo agoyes, but $ is operationally useless since $ is model depenedent and as other articles currently on HN show, the model+harness are symbiotic or antagonistic; if you strip out the $ part of it, you can focus on the interdependence of Agent+subagent and you can evaluate that in a stricter sense because I would want to select the model+subagent that improves my outcomes instead of whicheere is cheapest. If you tell me this model has a 99% chance of succeeding at $10 and this one has a 50% chan
- XUEYANZ 2mo agoHaha. somehow i just love the naming. it just makes sense :D
- isadubois 2mo agoThis looks super clean. I'm curious about the --judge command. How does it evaluate if the cheaper model's response is a "tie" or acceptable? Is it using a specific LLM-as-a-judge prompt template?
- jarodrh 2mo agoThanks. Yes, that's exactly right, and aptly named so in the code. A few things to note: the prompt is actually deliberately tie-biased; the tie only then breaks on clear material differences: factual error, missing information the prompt requested, or just a complete failure to follow the instruction. The judge is explicitly told that response/output length, wording, style and formatting are not quality differences. To ensure there's no favouritism, the judge only sees the outputs (A and B), it never sees whether the output is from the current model or from the candidate; and to ensure no funny business occurs, the outputs for A and B are randomized. The default model is chosen automatically from the highest-tier model in the log, but this could invite self-bias (this is cautioned in the report output), so you can override this with --judge-model You can check out the prompt for yourself in src/frugon/measure.py - search: JUDGE_PROMPT_TEMPLATE
- Capitanai 2mo ago[flagged]
- tokoi 2mo ago[flagged]
- ipanditshashi 2mo agoUseful but how it compares with other providers model
- jarodrh 2mo agoThanks. I'm not sure if you're asking "how does frugon compare provider models" or "can it compare across provider models" or even "how does frugon compare to other tools"... On the first: It does this utilizing pricing from LiteLLM registry, and quality tiers from LMArena. On the second: A log of calls for OpenAI can get recommendations for Gemini or Anthropic or DeepSeek etc. and vice versa. On the third: It's free, local and offline. Hosted tools normally require you to plug them into your live traffic, whereas frugon reads logs you already have. And you have the benefit of exploring the source code.
- cnxiaom 2mo ago[flagged]
- pauldavis 2mo agoI think this is great, and the next frontier is to analyze how well calls can be handled by a local model. Realistically, to do that usefully requires response time as a new dimension of judging: Can a local model provide an acceptably accurate response in an acceptable amount of time?
- jarodrh 2mo agoThanks. Yes, local models are gaining a lot of traction. The measure/judge step uses LiteLLM, so it does sample local models. I just tested "--candidates ollama/llama3.2:1b", and that works - ignoring the lack of rich UX for local/unpriced models, as I was focussing on cloud cost, but you've inspired me to give this area some polish. Noted: "response time as a new dimension of judging" - Added to the roadmap. Try it and let me know if you hit a wall.
- santiago-pl 2mo agoIf you experience any issues with LiteLLM, you may try GoModel - the AI Gateway I'm working on. It consumes ~60x less resources and is more reliable :)
- indigodaddy 2mo agoWhat we need is an AI gateway/router that will actually first analyze the input tokens and then decide what model to use. If it's so simple that a dirt cheap qwen 3.5 flash or whatever will be fine, then it chooses that. If it deems we need GPT 5.6, then it uses that, etc. does anything like this already exist?
- jarodrh 2mo agoYes they do, it's quite the popular topic atm. RouteLLM (OSS - frugon's savings bands are based on their research), OpenRouter's auto mode; I actually commented on a /show post not long ago: Wayfinder Router (neat project) - https://news.ycombinator.com/item?id=48704373 https://news.ycombinator.com/item?id=48704373 and more. My stance on the router point is a config-led (deterministic) one, even if the eventual industry consensus is a hybrid of dynamic (AI-led) + deterministic. Both sides still require evals/evidence for policy creation, which is what frugon provides.
- westurner 2mo agoEvals and OpenInference (OpenTelemetry) might be useful. Costed opcodes (like the shelved eWASM opcodes cost chart) would be useful for this model routing problem as well. Is this the cost to converge problem, the minimize cost to converge upon sufficiently low error problem, or the minimize cost and error problem? EA methods: mutation, crossover, selection Gradient descent as a mutation, crossover, and selection pattern; back up when the error/cost stops decreasing for too long and try a different branch. A simple experiment: vary only a nonce in the prompt and compare output value. The nonce is a parameter. The model is a hyperparameter.
- jarodrh 2mo agoVery interesting thoughts. This is a whole discussion on its own. The ones that I mulled over that I feel have legs in my line of thinking are: > minimize cost to converge upon sufficiently low error problem - This is the one. Frugon does an easy/hard split to keep quality within a certain tolerance, and find the lowest costed model within that constraint. The judge handles the "sufficiently low error" side of that via sampled prompt logs. > Evals and OpenInference (OpenTelemetry) might be useful. - Evals already a given. This is frugon's wheelhouse, however it currently only supports OpenAI-style jsonl. I'll definitely consider adding support for OpenInference. > A simple experiment: vary only a nonce in the prompt and compare output value. The nonce is a parameter. The model is a hyperparameter. - This is an intriguing idea. This would actually enhance/strengthen frugon's stance on the "minimize cost to converge upon sufficiently low error problem". Noted.
- westurner 2mo ago> minimize cost to converge upon sufficiently low error problem Is this a convex optimization or non-convex optimization problem? > Evals Re: pytest-evals, mcpbr, agentevals, foundry-toolkit; and devtools-mcp; https://news.ycombinator.com/item?id=48532642 https://news.ycombinator.com/item?id=48532642 OpenInference and OpenTelemetry have a schema for agent sessions. I finally discovered agentsview and ctxrs/ctx for indexing and searching agent session logs. Perhaps such an index is also useful for agent optimization > Nonce Accuracy is sensitive to nonce and temperature without variance in the model hyperparameters
- kliukovkin 2mo ago[dead]
- mujia 2mo agoA cheaper model can look fine per request, but one weak answer may create another call or a human review step. That seems like the hardest cost to capture.
- jarodrh 2mo agoTrue, but human review isn't visible in the logs, hmm, however I suppose retry patterns can, but frugon doesn't currently track this, so we can't quantify this yet. Frugon analyzes your existing logs offline on cost and quality tier. The "looks fine per request" is exactly the reason why --judge exists. Earlier, cyanydeez inspired the consideration of a metric "effective cost per judged success", which could also answer this point. Try it on your logs, and tell me if anything falls through the cracks.
- modgate 2mo ago[flagged]