Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
sumanyusharma
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
sumanyusharma
1y ago
Congratulations on the launch. Few qs: How do your agents decide a suspected issue is a validated vulnerability, and what measured false-positive/false-negative rates can you share? How is customer code and data isolated and encrypted
2.
▲
by
sumanyusharma
2y ago
How is this different from Integuru? They posted a few weeks back here: https://news.ycombinator.com/item?id=41983409
3.
▲
by
sumanyusharma
2y ago
Appreciate the info; I'll double-check Fidelity again!
4.
▲
by
sumanyusharma
2y ago
I'm actually pretty interested in what you're building. Sure, Vanguard and Fidelity are well-established giants, but they've barely moved beyond standard ETFs for decades. Having the option to tweak weightings at a more granu
5.
▲
Hamming AI (YC S24) Is Hiring a Product Engineer
(ycombinator.com)
1 points
by
sumanyusharma
2y ago
6.
▲
Hamming AI (YC S24) Is Hiring a Founding Engineer in SF
(ycombinator.com)
1 points
by
sumanyusharma
2y ago
7.
▲
by
sumanyusharma
2y ago
No plans for acquisition :) Building product, talking to customers and making something people want!
8.
▲
by
sumanyusharma
2y ago
We're focused on end-to-end evals focused on function-call accuracy, style, tone & latency of the conversations between our sims and your voice agent. Less focused on pure TTS evals at the moment!
9.
▲
by
sumanyusharma
2y ago
Pipecat looks awesome! I'll run the examples over the weekend and try to see what the integration hooks need to look like: https://github.com/pipecat-ai/pipecat/tree/main/examples It should be prett
10.
▲
by
sumanyusharma
2y ago
Likely outsourced call centers since call complexity is low to medium. We also expect rapid adoption in industries like customer service, healthcare, and retail, where 24/7 availability could be high-impact for businesses and convenien
11.
▲
by
sumanyusharma
2y ago
Should be fixed now; could you try again please?
12.
▲
by
sumanyusharma
2y ago
We forgot to enable non-US numbers in our config for the demo. (oops) We're working on a fix right now!
13.
▲
by
sumanyusharma
2y ago
I am curious - how was the team solving this at Kea?
14.
▲
by
sumanyusharma
2y ago
I use Superwhisper (no affiliation, just a happy user), which runs a local Whisper model, to create most of my email drafts and post-meeting notes. I find Whisper more accurate than Mac’s built-in speech-to-text, plus I’m faster at speaking
15.
▲
by
sumanyusharma
2y ago
I'm curious to learn more about what's blocking the widespread adoption of the LLM capabilities. Lack of knowledge, reliability, or something else?
16.
▲
by
sumanyusharma
2y ago
This tracks. Text evals to test core logic and voice evals for overall end-to-end performance!
17.
▲
by
sumanyusharma
2y ago
It's a bit of a catch-22. Making current voice agents reliable is incredibly time-consuming and complex. This challenge has kept many teams from pushing their agents into production. Those who do launch often release a very limited, ba
18.
▲
by
sumanyusharma
2y ago
Our customers, who build voice agents, are often asked by their customers to make their voice agents more human-like and flexible. Their clients — businesses like pest control and automotive repairs — value providing a personalized experien
19.
▲
by
sumanyusharma
2y ago
Bolna looks awesome! We've considered going open-source, but we're not sure how to effectively manage a community. I'll reach out async!
20.
▲
by
sumanyusharma
2y ago
Absolutely agree that creating effective evals requires domain expertise. Right now, we're co-building evals with customers, but we're identifying which aspects can be productized. Regarding text-based evals — part of testing voic
21.
▲
by
sumanyusharma
2y ago
Yes! We're aiming to build a tool that both engineers and non-engineers love. We've discovered that it's often faster for non-technical domain experts to iterate on prompts in a structured, eval-driven way, rather than relyin
22.
▲
by
sumanyusharma
2y ago
I wonder if a more optimistic version of this could be used to train humans and improve their skills. I'm thinking along the lines of LeetCode / Project Euler, but more dynamic and personalized! Few examples: 1) Customer service:
23.
▲
by
sumanyusharma
2y ago
Yes! Drive-through customers can be very impatient. We tried to make the demo persona maximally annoying. Testing for edge cases is especially important because getting an order wrong can cause health hazards, long line-ups, and churn!
24.
▲
by
sumanyusharma
2y ago
Nice! What's the use case your agent solves for? I'm happy to spin up some scenarios that are more relevant for you instead of our stock demo personas :) Feel free to email me at sumanyu@hamming.ai
25.
▲
by
sumanyusharma
2y ago
Yup, we named it after Richard Hamming. His essay 'you and your research' was deeply influential during my undergrad; I re-read it every quarter. Our current product draws inspiration from Hamming distance because we're compa
26.
▲
by
sumanyusharma
2y ago
I appreciate your candid feedback. Our aim isn't to push for replacing humans but to ensure that when companies do use LLMs, they work as intended and don't create more problems. We'd rather see well-functioning systems than
27.
▲
Launch HN: Hamming (YC S24) – Automated Testing for Voice Agents
129 points
by
sumanyusharma
2y ago
|
66 comments
28.
▲
New code-focused LLM needle in the haystack benchmark
(github.com)
6 points
by
sumanyusharma
2y ago
|
1 comments
29.
▲
by
sumanyusharma
2y ago
Hi HN - In collab with UWaterloo, we published a new code-focused needle in the haystack benchmark. TLDR - GPT-3.5-Turbo showed lower accuracy on the BICS benchmark than the BABILONG benchmark at the same context length and target depth, in
30.
▲
Can LLMs find bugs in large Python codebases?
(hamming.ai)
5 points
by
sumanyusharma
2y ago
|
3 comments
More ›