Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
abdik
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
abdik
29d ago
not really counter-positioned, they are one of the providers we route to. elevenlabs sells models - voices, stt, and their agent platform on top of them. we don't sell any models: we measure all of them (elevenlabs included, their scri
2.
▲
by
abdik
1mo ago
i don't have a thesis on voice as a form factor for search. the demand we serve already exists: businesses answer phones. receptionists, outbound campaigns, clinic front desks and these calls happen at scale today, and the teams runnin
3.
▲
by
abdik
1mo ago
thanks! yes, some of them do, cuz speed is a per-provider capability, not universal. And, you can see it in the gateway code (minimax, hume, xai tts adapters all handle a speed param). that unevenness is actually a routing constraint by it
4.
▲
by
abdik
1mo ago
People who built or building voice agents immediately get this. openrouter is the openrouter for audio models. the conflation is "audio models" vs "voice ai", and the mental model that untangles it: think batch requests.
5.
▲
by
abdik
1mo ago
yes, it does support, when you are creating an api key, you can point out narration or transcription use case, then you will be able to see. let me know how it goes or what use cases are thinking of for non-realtime models?
6.
▲
by
abdik
1mo ago
looks cool, checking it out!
7.
▲
by
abdik
1mo ago
exactly, we see the same thing, around 95% cases are still cascaded, even tho STS has been improving a lot
8.
▲
by
abdik
1mo ago
Synthetic-data fine-tuning is the other credible answer to domain vocabulary. Curious whether you re-benchmark the fine-tune when new base models ship?
9.
▲
by
abdik
1mo ago
What purrcat259 said, and I think that the page should explain it, we will add a tooltip. for some languages CER is more relevant than WER. Thai and Mandarin have no word boundaries, so we score them by character, and Japanese gets a readin
10.
▲
by
abdik
1mo ago
agree with omneity here. Whisper's initial-prompt trick is exactly that, and several hosted vendors have equivalents (custom vocabulary / keyword prompting). Domain vocabulary is where STT models separate the most in our runs. for
11.
▲
by
abdik
1mo ago
Actually, we measured exactly this recently. The strongest open-weights speech-to-speech model we have run is NVIDIA's NemotronLabs VoiceChat 11B - no provider serves it, so we hosted it ourselves and ran the same scripted call every m
12.
▲
by
abdik
1mo ago
cool app, and agreed that on-device keeps eating the single-user cases, we are seeing dictation and translation are exactly where local models shine. We benchmark the open models on the same boards as the hosted ones, but there are still a
13.
▲
by
abdik
1mo ago
Turn-taking specifically does not need a listening panel, but it is measured mechanically. 200+ real human clips, and we score end-vs-wait decisions: did the model decide the caller finished speaking, or just paused mid-thought. The best de
14.
▲
by
abdik
1mo ago
Yes, on the hosted side (agents platform): full sessions come with VAD and turn-taking handled - we set them up and tune them for your use case, so that is the closest thing to conversation in a box. If you run your own orchestration, the g
15.
▲
by
abdik
1mo ago
there is a good progress on on-devise models, but not ready for production yet to fit in devices. But as soon as there is are some good results, we are going to benchmark them and put in https://benchmarks.speko.ai/
16.
▲
by
abdik
1mo ago
thanks! actually, we have the filipino already, can you check out and share your feedback?
17.
▲
by
abdik
1mo ago
Fair pushback. On end to end: we measure those too, same methodology: https://benchmarks.speko.ai/s2s . If the single models win, we route to them the same way, so we do not care which architecture (s2s or cascaded) wins. Fo
18.
▲
by
abdik
1mo ago
The main difference from gateway is we help with picking the right voice stack, which seems to be a big problem for users: we benchmark the models continuously and route based on those measurements for your language and constraints, and the
19.
▲
by
abdik
1mo ago
Thank you!
20.
▲
by
abdik
1mo ago
Yes, i added it. that's the right link.
21.
▲
Launch HN: Speko (YC S26) – OpenRouter for Voice AI
(speko.ai)
118 points
by
abdik
1mo ago
|
69 comments
22.
▲
Show HN: YC Interview Simulator (Voice of Garry Tan)
(yc.speko.ai)
3 points
by
abdik
4mo ago
|
1 comments
23.
▲
by
abdik
5y ago
The logo is similar to ours https://www.lovo.ai/
24.
▲
by
abdik
5y ago
Check out https://nomadlist.com/
25.
▲
by
abdik
8y ago
YY9mtbXH2xpcyTvPeFVqh6Guo7ISe47HGtfcrqt11nsDy9hJDAQ50er6KpCCmILl1ztJ6xdC/7vPSTyiTEQfUYP05ZMSsv7e5IAa3xO0U4VZr/9rTEEub/a0epxZujTJSlazNsdYlFRMrUDekVsqIxq6bjlf3v5lQdIZxQMGOscp4cbfWgMZuf0yFiZb0t7S2W4I5UsHyxmR/dZg7eHx2An+CTei
26.
▲
Ask HN: Can you share indoor mobility related resereach ideas?
1 points
by
abdik
10y ago
|
0 comments