Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
Aperswal
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
Aperswal
2mo ago
Currently building a LLM router and got a few of my startup friends to try it out. But, I feel like their advice is going to be slightly biased because they don't wanna hurt my feelings. Not sure though if I should send out ads or emai
2.
▲
Show HN: FlexInference LLM Router
(flexinference.com)
3 points
by
Aperswal
2mo ago
|
0 comments
3.
▲
by
Aperswal
2mo ago
If you give an SLA of lets say 10 seconds to find cheaper inference. I look around different providers that offer flex/batch and try to fulfill in that time. If I can't then I escalate to a default request. This means you definite
4.
▲
by
Aperswal
2mo ago
The basic logic for default routing or flex routing?
5.
▲
by
Aperswal
2mo ago
This isn't open sourced, as of now. The routing is largely pass through or standard conversions between API surfaces. The SDK uses the OpenAPI specs for it's given API surface. The eval is in the hero of the website itself, everyd
6.
▲
Show HN: Made a Free LLM Router
(flexinference.com)
2 points
by
Aperswal
2mo ago
|
6 comments
7.
▲
Show HN: A router that drops costs by roughly 45%
(flexinference.com)
2 points
by
Aperswal
2mo ago
|
1 comments
8.
▲
by
Aperswal
2mo ago
Hey everyone! My name is Adi, I am the founder of FlexInference. I built this to reduce AI spend for all my side projects. I then turned it into a proxy for Anthropic, OpenAI, and Gemini. I also then built a python and typescript SDK with s
9.
▲
by
Aperswal
1y ago
Btw the link above is broken because of mismatched quotes But u can check out the repo at: https://github.com/TrySita/AutoDocs
10.
▲
Show HN: AutoDocs – Reduce AI costs and never manage context again
(github.com)
3 points
by
Aperswal
1y ago
|
2 comments