10 ms·
(Atty from OpenAI here) GPT-Live-1 is the first version of a new generation of models, and we believe the full-duplex architecture + delegation enables entirel
by athyuttamre 2mo ago
(Atty from OpenAI here)
GPT-Live-1 is the first version of a new generation of models, and we believe the full-duplex architecture + delegation enables entirely new ways of human-AI interaction.
Would love to hear your feedback!
- 6thbit 2mo agoCan it delegate to just one agent at a time or can it spawn multiple subagents for different tasks?
- athyuttamre 2mo agoThere are many delegation models possible: 1. The voice model delegates to one agent. 2. The voice model delegates to multiple agents, and keeps track of tasks. 3. The voice model delegates to an orchestrator agent, which then delegates to sub-agents and keeps track of tasks. YMMV depending on the exact product experience you care about, because there is a tradeoff between latency and layers of delegation. Our current implementation is backed by one model, but you can imagine this getting much better with time.
- worldsavior 2mo agoWhat made you to try again?
- famouswaffles 2mo agoDoes video/image input still work with these duplex models?
- athyuttamre 2mo agoImage input is supported, but video is not today. We're working hard to bring it to you soon.
- dandaka 2mo agoCan I connect it to my skills/tools? Example case, I have a knowledge base and event log in my company. I need a brainstorm companion, which will have full access to this knowledge, can converse about it and can invoke skills/tools available in the repo.
- athyuttamre 2mo agoIn ChatGPT, Voice doesn't yet support connectors, but we're hoping to add support soon! Once GPT-Live launches in the API, you can also build custom integrations yourself.
- BoorishBears 2mo agoDo integrations supporting streaming input? One big gap I've run into for UX is most realtime voice harnesses wait for a full response from tools, and at most support the model filling the dead air until then It'd be a game-changer to be able to have the model start replying with partial information streamed from the tool call, then seamlessly continue with additional information.
- alooPotato 2mo agohonestly its not that big of a deal, being able to call tools is the high order bit
- BoorishBears 2mo agoIt rubs me the wrong way that someone would respond to something I need with 0 context just to tell me it's not a big deal. Some people have standards in what they build.
- alooPotato 2mo agook sorry - can you explain more why it'd be a game changer to stream the tool call. Whats an example tool call that this would help with. Most of the time the tools I want called are at the end of the conversation (i.e. "summarize what we talked about and email it to me"). Or if its in the middle of the conversation, we can just carry on the conversation while the tool call is happenning in parallel..
- vessenes 2mo agoCan it sing? Is it an end to end multimodal model?
- ls-a 2mo agoAre you still using LiveKit for the back-and-forth architecture
- Sean-Der 2mo agoWrite up about the architecture is here https://openai.com/index/delivering-low-latency-voice-ai-at-scale/ https://openai.com/index/delivering-low-latency-voice-ai-at-...
- k2xl 2mo agoWhen is it rolling out? Currently on ChatGPT Pro but not seeing it yet?
- Fraterkes 2mo agoHey! Bit of an unusual question maybe: if this stuff further exarcerbates the loneliness epidemic and atomization of society, will you be able to live with yourself you think? If you hear about teenagers only spending time with your chatbot in 5 years, will you feel some amount of personal responsibility or not? Always curious to hear you guys' perspective on that kind of stuff!
- z7 2mo ago"Hey Leibniz, how do you live with yourself knowing that your binary system helped eventually replace human conversations?"
- deleted 2mo ago[deleted]
- vinaigrette 2mo agoFair question, although I think he really has a hard time living with it ...
- dag100 2mo agoOh no, now I'd better blame Hitler and Stalin's ancestors for their misdeeds. Of course, the SS soldier bears no responsibility, however, as he was just doing his job.
- overfeed 2mo ago> if this stuff further exarcerbates the loneliness epidemic and atomization of society, will you be able to live with yourself you think? Looking at the 30,000-foot view of how society is set up: laws, economic system, employee incentives, etc, do you suppose it matters what the individual contributors think? I say this not to absolve anyone of responsibility, but to point out the obvious outcomes of our incentives across the strata (polity -> shareholders -> boards -> C-suite -> employees) I will bet you dollars to donuts, somewhere inside OpenAI is a frequently-used revenue dashboard, but not for loneliness - if anything, OpenAI will make horny models and tout itself as a solution to loneliness, a la character.ai - if that earns them more money.
- quotemstr 2mo ago> GPT-Live-1 is the first version of a new generation of models, and we believe the full-duplex architecture + delegation enables entirely new ways of human-AI interaction. Awesome. Are you guys able to share anything about the model architecture? I've been interested lately in split-transformer RVQ-based conversational agents, e.g. via stuff like https://arxiv.org/abs/2412.10208 https://arxiv.org/abs/2412.10208 (ResGen) and https://arxiv.org/abs/2603.18090 https://arxiv.org/abs/2603.18090 (MOSS-ITT) and of course Moshi (https://arxiv.org/pdf/2410.00037 https://arxiv.org/pdf/2410.00037). Intuitively, decoupling semantic and audio-timeslice-space generations with coupled but distinct histories is right model architecture, not just for these sorts of assistants, but for domains like robotics too.
- mycocola 2mo ago- As models get better, have you considered some kind of filter or particular cadence to serve as a reminder that the user is not talking to a human? - The videos felt scripted and dishonest
- ollin 2mo ago[dead]
- iandanforth 2mo agoCan we have less terrible voices please? Nothing that sounds like a bubbly millennial. Literally anything that has gravitas.
- athyuttamre 2mo agoHave you tried some of our deeper voices like Spruce? Would love to hear what your ideal voice is.
- iandanforth 2mo agoI have. They are not to my taste. Here's a voice I really enjoy listening to. https://elevenlabs.io/app/voice-library?search=zNsotODqUhvbJ5wMG7Ei https://elevenlabs.io/app/voice-library?search=zNsotODqUhvbJ... If I could have Christopher Lee, or Stephen Fry, I would.
- znpy 2mo ago> Would love to hear your feedback! I'm currently on the 20 $/mo subscription and using codex meaningfully, and i'm loving this. I am considering bumping my subscription to the 100 $/month and this might be the reason i switch, BUT: i really envision me using this also through other means as well (eg: agents like openclaw/hermes) in agentic ways. Will this be supported? I can make OpenAI stuff the center of my agentic AI life, but I need it to be interoperable.
- athyuttamre 2mo agoWe're adding support to the API soon, which will let you integrate with any agent in the background. Would love to see the community go wild with it. You can sign up to be notified here: https://openai.com/form/gpt-live-1-in-the-api/ https://openai.com/form/gpt-live-1-in-the-api/
- hersko 2mo agoHow does it compare to the realtime-2 model?
- djtriptych 2mo agoI'm interested in how you can present simultaneous rich visual information about what is happening the side delegation work. i.e. how will full duplex & delegation enable/enhance desktop flows w/o corresponding leaps in UI.
- bariswheel 2mo agoIf it's not able to connect to my other apps like gcal, it should suggest that I use the standard gpt client, instead of telling me 'I don't have the capability to do that' which is misleading as some might presume the normal client doesn't have that capability either.
- interloxia 2mo agoAny feedback from different locations or cultural groups? One group's expectation of interruption for pleasant conversational flow can be just as off-putting as another's expectation of patient silence.
- sandspar 2mo agoI like it! I was watching a YouTube video of someone driving in a big city. I saw an interesting skyscraper and asked Chatgpt what it was. After Chatgpt answered, I asked for a photo of the building to confirm, and Chatgpt helpfully showed it in-chat. It was the correct building! And because the voice is so frictionless to talk to, I asked about what company owns the building, then that company's industry, then how that industry works in this particular country etc. I probably wouldn't have bothered going down a rabbit hole like this if I'd had to type. Voice is much easier than typing. Anyhow it's fun! Thanks for making it!
- justanotherunit 2mo agoHey! Nice work, been looking into having a conversational modal integrated with my dev workflow. Are we seeing any conversational layer integrated with codex soon?
- NotSuspicious 2mo agoSuper nice to be able to finally have voice mode search stuff in the background. The more conservational vibe is also a nice touch. The real game-changer we're all waiting on, though, is being able to vibecode/deeply interact with an OS through voice mode. Until then voice mode will continue to be a niche, interesting product. Once that's implemented, though, it will be one of the most important technologies available. There's a lot of people I know that don't use a lot of the frontier capabilities of AI right now because of the required typing interface. Once that's changed and you can interact with these capabilities through voice mode (and ideally on your phone) I definitely see them and lots of people like them becoming big users of AI.
- pietmichal 2mo agoathyuttamre is a coward if he doesn't answer this