7 ms·
I'm stunned that people are taking this accusation as a fact. OpenAI is no stranger to rivalry with Anthropic but 1. it's not like user data is sitting around
by postalcoder 9d ago
I'm stunned that people are taking this accusation as a fact.
OpenAI is no stranger to rivalry with Anthropic but 1. it's not like user data is sitting around on some kitchen table somewhere and 2. I consider OpenAI to be as economically motivated as any other actor in this space and playing around with user data like that would destroy their business.
There are things that Buckmaster alleged and things that he speculated. The entire training data thing is speculation. If this is pissing you off, then you ought to evaluate how you ingest information.
- defmacr0 8d agoI am stunned anyone is giving OpenAI the benefit of the doubt
- black_rabbit_ 8d agoI'd be surprised if all of it is organic discussion, shall we say. I reckon The Bot Factory just possibly might dogfood the astroturf machine.
- sensanaty 8d agoI guarantee you most of the comments regarding this aren't real humans. The homepage is full of crap meant to distract from what OAI did here, the comments are full of OAI employees. Dead internet theory pushed to the max
- thereitgoes456 9d agoHe asked whether they used their chats as training data and received no response. Any speculation here seems quite appropriate?
- deleted 9d ago[deleted]
- dash2 9d agoHe didn't even make that accusation! > I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer. The shocking/interesting thing would be if it was trained on the sessions. I think it's very implausible that they gave the model access to someone else's sessions as input. That would be a huge privacy violation and would probably blow up a large proportion of their enterprise business. Does openAI train on user conversations in general? I assume so. But so fast as that? That seems unlikely in general. I expect OpenAI will come out denying this.
- deleted 7d ago[deleted]
- actionfromafar 9d agoCouldn't the Enterprise have a different fine print?
- johnnienaked 9d agoIt wouldn't be shocking at all. They stole human data to train the first models and they've been stealing it ever since to train new models. Stealing mathematicians private chats and private research and taking credit for it would absolutely be par for the course. Enterprises are well aware of it and are fully on board. You didn't think every corporation in America has an OpenAI subscription because the models were good, did you? The whole reason they have subs is to train them on YOUR WORKFLOWS lol
- johnnienaked 9d agoThey stole it.
- FiberBundle 9d agoWhenever I see comments defending AI companies, I look at the account's creation date, and interestingly almost all of them were created post 2024.
- az226 9d agoI don’t think you understand how brazen big tech companies are in practice.
- paxys 8d agoThis is how internet discourse works on Reddit/Twitter/HN and the rest. Someone said something which confirms your biases so it’ll now be treated as a fact and repeated endlessly in the echo chamber.
- hgoel 8d agoEspecially after the blatant cover up of their uncontrolled bot swarm infesting the internet, and the feckless "hopefully we do better" response upon being caught, I don't think OpenAI deserves much grace until they properly explain themselves. We had all assumed that surely the supposed smartest engineers in the world, with access to the most computing and a direct view of model capabilities, would take sandboxing and cybersecurity much more seriously than they have turned out to do. It follows that while we might assume they take user data privacy seriously and have tight controls on who can access it, it's possible they do not actually do that. At this point any initial trust is dead and has to be re-earned.
- haxiomic 8d ago> playing around with user data like that would destroy their business. Their entire business is based on stealing data. They can make a calculation that the cost stealing data is less than the cost of the positive publicity they can shape for solving Millennium NS
- sherburt3 8d agoGiven the history of OpenAI and current litigations, I would say they've developed a bit of a reputation for not respecting intellectual property. I'm dubious they have some unbreakable moral code that would prevent them from viewing and using user data.
- ozgung 8d ago> The entire training data thing is speculation. I think it's safe to assume AI labs DO train on your data and it's very hard to prevent that. I've just checked my inaptly named "Help improve our AI models" toggles. The toggle on the Claude settings had magically turned on. I asked about how this can happen. Claude says they show re-consent modals when terms change, and it is a "real and fairly common pattern" to re-opt in without noticing. All my work and conversations since I don't know are now part of their training corpus. No way to take it back. Google's Gemini/Antigravity didn't have opt-out toggles at all last time I checked. Codex also has a separate "include environments" setting which is hard to find (found it in Codex Cloud) and I don't know what it does. Lots of Dark UI Patterns here even if we assume they keep their promise. For this incident, Occam's Razor says their internal models somehow saw a version of the mathematicians' logs, during or after training. Maybe indirectly. These systems are literally designed to collect data. Privacy and safety is not trivial to achieve on the users' side. Simply because it's against the labs' best interest.
- cma 7d agoYou can opt-out of training in Gemini on personal plans, but it disables your chat history, just to be vindictive; there is no technical reason and the other companies don't do this.
- cyclopeanutopia 8d agoBut the whoreshippers of The Holy Dollar will tell you it's all good and justified.
- iaw 8d agoAbsence of evidence is not evidence of absence. With the behaviors we know OpenAI engages in the accusations are wholly believable.
- ajkjk 8d agoEverybody knows it's not a sure thing, it's a question of trustworthiness. OpenAI is not trustworthy at all; this random researcher is and seems honest so far. iThe fact that people are corroborating Bubeck being a piece of shit in other settings add to credence. But nobody is over here saying it's an indisputable certainty. And your (2) is probably false, their history of deception suggests they would do just about anything as long as they didn't think it would backfire on them publicly.
- Tanjreeve 8d ago>The entire training data thing is speculation Quite literally in the terms of use.
- timdiggerm 7d ago> I consider OpenAI to be as economically motivated as any other actor in this space and playing around with user data like that would destroy their business. It's a gamble that this would be overlooked compared to the reputation they build for solving the thing
- sp527 7d ago> I consider OpenAI to be as economically motivated as any other actor in this space and playing around with user data like that would destroy their business. And that's exactly why you can only see evidence of this when the stakes are high enough, such as solving a Millenium problem. The legalese they outputted in response to the incident left them an escape hatch that permits the possibility of theft. People familiar with corporate damage control should recognize the verbal maneuvering, often used to paper over actual guilt. There was also already a high prior of shadiness. The company is run by someone who is close enough in reputation to "known sociopath". It's actually far more reasonable to assume that OpenAI stole user data to generate a breakthrough. They have an unbelievably high economic motive to do so. And even a low chance that they might be doing this implies catastrophic risk to anyone with valuable knowledge.