Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
bcherry
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
1.
▲
by
bcherry
6mo ago
this really reminds me of the "worst volume control" from reddit https://uxdesign.cc/the-worst-volume-control-ui-in-the-world...
2.
▲
by
bcherry
7mo ago
you mention voice ai in the announcement but I wonder how this works in practice. most voice AI systems are bound not by full response latency but just by time-to-first-non-reasoning-token (because once it heads to TTS, the output speed is
3.
▲
by
bcherry
8mo ago
This isn't really the author's point, but I think one effect of AI and the forthcoming robotics revolution will be the unrolling of a lot of consolidated supply chains for all sorts of products. It could usher in a renewed era of
4.
▲
by
bcherry
11mo ago
wow thanks for leaving this comment - i now realize two things: 1. the farmer's almanac i thought of when i saw the title and even read the article is not going anywhere 2. i have never before heard of the farmer's almanac referre
5.
▲
by
bcherry
11mo ago
they'd have to be extra careful with cpython, it's got a lot of include
6.
▲
by
bcherry
1y ago
yeah i think they shot themselves in the foot a bit here by creating the o series. the truth is that GPT-5 _is_ a huge step forward, for the "GPT-x" models. The current GPT-x model was basically still 4o, with 4.1 available in som
7.
▲
by
bcherry
1y ago
"The sculpture is already complete within the marble block, before I start my work. It is already there, I just have to chisel away the superfluous material." - Michelangelo
8.
▲
by
bcherry
2y ago
Chat is a great UX _around_ development tools. Imagine having a pair programmer and never being allowed to speak to them. You could only communicate by taking over the keyboard and editing the code. You'd never get anything done. Chat
9.
▲
by
bcherry
2y ago
a little glossed over, but they do point out that most important improvement o1 has over gpt-4o is not it's "correct" score improving from 38% to 42% but actually it's "not attempted" going from 1% to 9%. The i
10.
▲
by
bcherry
2y ago
disagree - good products meet their users where they are and bury complexity under the hood. i can't imagine trying to use a calendar app (or any app really) that refuses to operate in any mode other than UTC.
11.
▲
by
bcherry
2y ago
It's kind of interesting because I think most people implementing RAG aren't even thinking about tokenization at all. They're thinking about embeddings: 1. chunk the corpus of data (various strategies but they're all som
12.
▲
by
bcherry
2y ago
hey sorry about that - ran into a snag with the API but we got it back online an hour ago! hope you get another chance to take a look! reply
13.
▲
by
bcherry
2y ago
hey sorry about that - ran into a snag with the API but we got it back online an hour ago! hope you get another chance to take a look!
14.
▲
by
bcherry
2y ago
this one was actually so much fun I built it into the defaults https://playground.livekit.io/?preset=doom
15.
▲
by
bcherry
2y ago
It sure can https://playground.livekit.io/?preset=0tfwypgx7&instructions...
16.
▲
Show HN: Speech-to-speech playground for OpenAI's new Realtime API
(playground.livekit.io)
10 points
by
bcherry
2y ago
|
11 comments
17.
▲
by
bcherry
2y ago
yes and you can use it in text-text mode if you want. a key benefit is for turn-based usages (where you have running back and forth between user and assistant) you only need to send the incremental new input message for each generation. thi
18.
▲
by
bcherry
2y ago
correct - you should also be able to save a lot by skipping their built-in VAD and doing turn detection (if you need it) locally to avoid paying for silent inputs.
19.
▲
by
bcherry
2y ago
keep in mind that this is just v1 of the realtime api. they'll add realtime vision/video down the road which can also have wide applications beyond synchronous communication.
20.
▲
by
bcherry
2y ago
yes it transcribes inputs automatically, but not in realtime. outputs are sent in text + audio but you'll get the text very quickly and audio a bit slower, and of course the audio takes time to play back. the text also doesn't cur
21.
▲
by
bcherry
2y ago
No, it's the same thing as ChatGPT advanced voice. Full speech-to-speech model.
22.
▲
by
bcherry
2y ago
LiveKit, Cartesia, Deepgram, and Vercel
23.
▲
by
bcherry
2y ago
there's a difference between "click bait" (a misleading title specifically crafted to drive instinctual interest but which is not an accurate summation of the content) and titles accurately describing something truly interest
24.
▲
by
bcherry
3y ago
I wonder if they're allowed at SFO?
25.
▲
by
bcherry
3y ago
"net income" may be a bit simplistic for the moment, but agreed that it doesn't need to be twice as good. it simply needs to be the best. just because the marginal value of adding people goes down, doesn't mean you shoul
26.
▲
by
bcherry
3y ago
it's inconsistent. when i set it up on friday and used it over the weekend, i couldn't read the text on my iphone or the keys on my keyboard. i'm using it at work right now, and have no problem - it's extremely sharp. no
27.
▲
by
bcherry
3y ago
analyst estimates are that they sold about 160,000 units through preorders. At $3500 per unit, that's $560M.
28.
▲
by
bcherry
3y ago
I agree that I/O has always been confusing and non-intuitive. O stands for "On"? For "Off"? I means eyes-shut? O means eyes open? The new symbols aren't any better but the UX of "push this button to turn
29.
▲
by
bcherry
3y ago
I suspect the truth (if this claim is true) is a lot more nuanced than "it did grade-school math", and there's more context surrounding this claim which insiders have and makes it much more interesting.
30.
▲
by
bcherry
3y ago
You don't need to accept their definition, neither does anyone else. But they do need to have a definition that they accept themselves, because it's used throughout their charter. https://openai.com/charter
More ›