7 ms·
(I work at OpenAI.) It's really how it works.
by gdb 2y ago
(I work at OpenAI.)
It's really how it works.
- xanderlewis 2y agoI like the humility in your first statement.
- Induane 2y agoI like their username.
- belter 2y agoYou might be talking to GPT-5...
- moab 2y agoPretty sure the snark is unnecessary.
- egillie 2y agonot snark. if only hn comments could show the right feelings and tonal language
- ayhanfuat 2y agoWas it snark? To me it sounds like "we all know you Greg"?
- xanderlewis 2y agoThis was my intention.
- moab 2y agoI misunderstood; my apologies.
- colecut 2y agoI don't think it was snark. The guy is co-founder and cto of OpenAi, and he didn't mention any of that..
- renewiltord 2y agoI downvoted independently. No problem with groupies. They just contaminate the thread. Greg Brockman is famous for good reasons but constant "oh wow it's Greg Brockman" are noisy.
- deleted 2y ago[deleted]
- theboat 2y agoI love how this comment proves the need for audio2audio. I initially read it as sarcastic, but now I can't tell if it's actually sincere.
- xanderlewis 2y agoIt’s completely sincere. I’m surprised by the downvotes. Greg Brockman needs no introduction.
- bombcar 2y agoAnd here I thought it was just a GNU debugger fan or something.
- latexr 2y ago> Greg Brockman needs no introduction. Even if that were true¹, it doesn’t mean everyone would know their HN user name. ¹ Greg may be well known within a select group of people but that’s way smaller than even users of ChatGPT.
- xanderlewis 2y agoI clicked through to see his bio; I didn’t know his username.
- jamestimmins 2y agoWith this capability, how close are y'all to it being able to listen to my pronunciation of a new language (e.g. Italian) and given specific feedback about how to pronounce it like a local? Seems like these would be similar.
- taytus 2y agoThe italian output in the demo was really bad.
- thegabriele 2y agoWhy would you say "really bad"?
- bzudo 2y agoIt doesn't have hands.
- nijuashi 2y agoThis was the best joke I’ve heard this year.
- mark38848 2y agoSo good!
- DonHopkins 2y ago"I Have No Hands But I Must Scream" -Italian Ellison
- rezonant 2y agoJoke of the day right there :-)
- GaggiX 2y agoI'm a native Italian speaker, it wasn't too bad.
- baq 2y ago> (I work at OpenAI.) Winner of the 'understatement of the week' award (and it's only Monday). Also top contender in the 'technically correct' category.
- swyx 2y agoand was briefly untrue for like 2 days
- behnamoh 2y ago> Winner of the 'understatement of the week' award (and it's only Monday). Yes! As soon as I saw gdb I was like "that can't be Greg", but sure enough, that's him.
- Uptrenda 2y ago[flagged]
- jasondigitized 2y agoBro what?
- Uptrenda 2y ago[flagged]
- mttpgn 2y agoLicensing the emotion-intoned TTS as a standalone API is something I would look forward to seeing. Not sure how feasible that would be if, as a sibling comment suggested, it bypasses the text-rendering step altogether.
- skottenborg 2y ago"(I work at OpenAI.)" Ah yes, also known as being co-founder :)
- leozq 2y agoyes, also known as a programmer loves coding a lot:)
- terhechte 2y agoRandom OpenAI question: While the GPT models have become ever cheaper, the price for the tts models have stayed in the $15/1Mio char range. I was hoping this would also become cheaper at some point. There're so many apps (e.g. language learning) that quickly become too expensive given these prices. With the GPT-4o voice (which sounds much better than the current TTS or TTS HD endpoint) I thought maybe the prices for TTS would go down. Sadly that hasn't happened. Is that something on the OpenAI agenda?
- deleted 2y ago[deleted]
- passion__desire 2y ago[flagged]
- 999900000999 2y agoHow far are we away from something like a helmet with chat GPT and a video camera installed, I imagine this will be awesome for low vision people. Imagine having a guide tell you how to walk to the grocery store, and help you grocery shop without an assistant. Of course you have tons of liability issues here, but this is very impressive
- rfoo 2y agoCan't wait for the moment when I can puta single line "Help me put this in the cart" on my product and magically sells better.
- smokel 2y agoThis Dutch book [1] by Gummbah has the text "Kooptip" imprinted on the cover, which would roughly translate to "Buying recommendation". It worked for me! [1] https://www.amazon.com/Het-geheim-verdwenen-mysterie-Dutch/dp/9080348163 https://www.amazon.com/Het-geheim-verdwenen-mysterie-Dutch/d...
- DonHopkins 2y agohttps://en.wikipedia.org/wiki/Steal_This_Book https://en.wikipedia.org/wiki/Steal_This_Book
- seanmcdirmid 2y agoOr tell the AI to optimize paper clip production as much as possible.
- macintux 2y agoJust the ability to distinguish bills would be hugely helpful, although I suppose that's much less of a problem these days with credit cards and digital payment options.
- krainboltgreene 2y ago> Imagine having a guide tell you how to walk to the grocery store I don't need to imagine that, I've had it for about 8 years. It's OK. > help you grocery shop without an assistant Isn't this something you learn as a child? Is that a thing we need automated?
- bjtitus 2y agoIs it possible to use this as a TTS model? I noticed on the announcement post that this is a single model as opposed to a text model being piped to a separate TTS model.
- cchance 2y agoThis is damn near one of the most impressive things, can only imagine especially with live translation and voice synthesis (eleven labs style) you'd be capable of to integrate with something like teams (select each persons language and do realtime translation to each persons native language, with their own voice and intonations would NUTS)
- purplerabbit 2y agoThere’s so much pent up collaborative human energy trapped behind language barriers. Beautiful articulation. This is an enormous win for humanity.
- deleted 2y ago[deleted]
- sensanaty 2y agoBy humanity you mean Microsoft's shareholders right? Cause for regular people all this crap means is they have to deal with even more spam and scams everywhere they turn. You now have to be paranoid about even answering the phone with your real voice, lest the psychopaths on the other end record it and use it to fool a family member. Yeah, real win for humanity, and not the psycho AI sycophants
- throwaway11460 2y agoLet AI answer the phone. I love it, can't wait. I hate answering phone but some people just won't email, they always call.
- rane 2y agoWill the new voice mode allow mixing languages in sentences? As a language learner, this would be tremendously useful.
- selfmodruntime 2y agoI've always been wondering what GPT models lack that makes them "query->response" only. I've always tried to get chatbots to lose the initially needed query, with no avail. What would It take to get a GPT model to freely generate tokens in a thought like pattern? I think when I'm alone without query from another human. Why can't they?
- throwthrowuknow 2y agoTrain it on stream of consciousness but good luck getting enough training data.
- kolinko 2y agoJust provide empty queey and that’s it - it will generate tokens no prob. You can use any open source model wirthout any promot whatsoever
- djur 2y agoYou might not have a prompt from another human, but you're always receiving new input.
- Filligree 2y agoAnd humans malfunction pretty badly without input. Even solitary confinement quickly drives them insane.
- selfmodruntime 2y agoYes, but that's the fundamental difference. Even if I closed my eyes, plugged my ears and nose and laid in a saltwater floating chamber, my brain will always generate new input / noise. (GPT) Models toggle between a state of existence when queried and ceasing to exist when not.
- kristiandupont 2y agoYou could just let the GPT run in a loop and it too would continue to generate tokens.
- ALittleLight 2y agoIn my ChatGPT app or on the website I can select GPT-4o as a model, but my model doesn't seem to work like the demo. The voice mode is the same as before and the images come from DALLE and ChatGPT doesn't seem to understand or modify them any better than previously.
- sumedh 2y agoGPT-4o text version is available not the multi modal one.
- jacobsimon 2y agoI couldn’t quite tell from the announcement, but is there still a separate TTS step, where GPT is generating tones/pitches that are to be used, or is it completely end to end where GPT is generating the output sounds directly?
- derac 2y agoIt's one model with text/audio/image input and output.
- jacobsimon 2y agoVery exciting, would love to read more about how the architecture of the image generation works. Is it still a diffusion model that has been integrated with a transformer somehow, or an entirely new architecture that is not diffusion based?
- rrr_oh_man 2y agohttps://community.openai.com/t/when-i-log-in-to-chatgpt-i-am-prompted-for-a-login-failure-and-a-message-something-went-wrong-please-make-sure-your-devices-date-and-time-are-set-properly/508758 https://community.openai.com/t/when-i-log-in-to-chatgpt-i-am... Sorry to hijack, but how the hell can I solve this? I have the EXACT SAME error on two iOS devices (native app only — web is fine), but not on Android, Mac, or Windows.
- dmarinoc 2y agoAre you blocking some of your traffic? I had the same issue until I (temporary) disabled NextDNS just for signing in. Sadly, the error returned is not related to the cause.
- rrr_oh_man 2y agoNo VPN. Mobile Internet and different Wifis. Turned off everything on the devices, from Safari content blockers to IP masking. Nothing seems too help.
- andybak 2y agoMay I just say this launch was a bit of a mess? The web page implies you can try it immediately. Initially it wasn't available. A few hours later it was in both the web UI and the mobile app - I got a popu[ telling me that GPT-4o was available. However nothing seems to be any different. I'm not given any option to use video as an input, the app can't seem to pick up any new info from my voice. I'm left a bit confused as to what I can do that I couldn't do before. I certainly can't seem to recreate much of the stuff from the announcement demos.
- sumedh 2y agoThe website clearly says that the text version is available now but the multimodal version will be released over the coming weeks.
- dpflan 2y agoWho's idea was the singing AIs? What specifically did you want to highlight with that part of the demo? I imagine that there is a lot of usage at the HQ, human + AI karaoke?
- hpeter 2y agoI can't wait to try it out, it sounds too good to be real. It will be fully available in Eu with the GDPR compliance?
- deleted 2y ago[deleted]