6 ms·
Buckmaster: > "I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the wh
by peri-cl 8d ago
Buckmaster:
> "I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer."
OpenAI (i.e. this OP):
> "While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models ."
- jrflo 8d agoI feel like it's far more likely that ordinary corporate espionage or leak led to this rather than OpenAI sifting through piles of user data to find this approach. Buckmaster's collaborator works at Anthropic, and could have been targeted. That would also explain why they aren't forthcoming with the source of the prompt.
- ChoosesBarbecue 8d agoI thought one of the issues was that they wanted to remove credit from Levant, the aforementioned Anthropic collaborator? Which doesn't make sense to me if he was leaking information, or defecting to OpenAI, but I might be misunderstanding your point.
- jrflo 8d agoI don't think he was defecting or leaking directly, just that it's entirely possible that this information got to OpenAI as a rumor rather than them directly spying on mathematicians chat logs.
- EthanHeilman 8d agoI believe jrflo was saying that OpenAI watches the chats of everyone from Anthropic because watching what Anthropic employees type into their personal ChatGPT accounts is a critical source of intelligence on is happening inside of Anthropic. I would be surprised if OpenAI isn't doing that. OpenAI will take any advantage they can get. If an employee at their primary adversary is typing useful intelligence into OpenAIs website, a website that does not promise privacy from OpenAI, the only reason they wouldn't weaponize that information against Anthropic is ethics or fair play.
- Yajirobe 8d agoWhy would Anthropic employee even use OpenAI's models? Cross-polination would have been avoided
- mlcrypto 8d agoThey should have used a zero data retention agreement, user error
- preg_match 7d agoI doubt it would've made any difference, zero data retention is truly trivial to get around and I think it's highly likely OAI is already doing this. Just use another model to distill, summarize, and/or paraphrase the data and boom - you get to train on the ideas, while claiming zero data retention, which is true. You legitimately retained zero data.
- peri-cl 8d agoI suspect this controversy will blow the case for ZDR wide open. Whatever the facts (possibly unknowable), it's going to become a very public lesson that data sovereignty was never about "having nothing to hide". If this is what they do to academic pure mathematicians, where the stakes are so low (financially)—just imagine the sort of front-running that could be happening in other places.
- dsdf3 8d agoYeah if I was Anthropic this would be part of my marketing strategy.
- sebzim4500 8d agoHow so? OpenAI and anthropic have basically the same retention policies
- amluto 8d agoHahaha, how exactly is an individual user supposed to get a ZDR agreement?
- lambda 8d agoWhy can't they rule it out? Is even OpenAI unable to track the provenance of all of their training data? This is one of the major problems with these enormous closed models, and even most open-weights models, which don't disclose their training process or training data. You can never be sure what went into its training. Did it come up with an idea originally, or is it just plagiarising its training data? Are there malicious inputs being used to train in particular behaviors when given certain trigger phrases? What are the characteristics of the RLHF data and what kind of biases are those embedding in the models? With proprietary closed models, or even open weights models that don't have open training datasets, you just can't answer these questions.
- causal 8d agoGood chance their whole training pipeline is vibe coded so yah they probably don't actually know.
- rfgplk 8d ago> Why can't they rule it out? Is even OpenAI unable to track the provenance of all of their training data? Probably? I have a few hundred TB of training data for various small scale models and I can attest that I have _no idea_ what's in them. As in, literally zero. Half is scraped from GitHub and other hosting sites, other than that, I couldn't tell you anything else. At OpenAI's scale their entire pipeline is likely 100% automated.
- lambda 8d agoYeah, I'm sure it's completely automated. But that doesn't preclude being able to index and track what the sources of data are. For your data sets, I would hope you are including source information for where the data came frome. And at OpenAI's scale, I would presume they are doing some amount of rolling hashing or similar to weed out duplication, training on too much duplicate data can cause problems. AllenAI have at least attempted to add some amount of traceability to their models with OLMoTrace (https://arxiv.org/abs/2504.07096 https://arxiv.org/abs/2504.07096), by letting you find n-gram matches from the outputs in their training data. It's not the most useful, there's a reason that LLMs use full fledged attention mechanisms and not just n-grams, a lot of times the n-gram matches it finds aren't all that related to the given output, it might be better to supplement this index with a vector search or other ways of keeping track of what training data would have most influenced particular parts of the output. But anyhow, this is something that is an important question, and the big labs should be working on to make their products more trustworthy. Instead, they are hiding information about how they train, hiding their reasoning traces, and just producing output with no information on what might have influenced the training.
- amluto 8d ago> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models . That’s a bizarre statement. Their website says: > Services for individuals, such as ChatGPT and Codex > When you use our services for individuals such as ChatGPT and Codex, we may use your content to train our models. > You can opt out of training through our privacy portal by clicking on “do not train on my content.” Are they not sure that the opt-out works? Oddly, their privacy portal page is not the same page as the one with the checkbox.
- fph 8d agoDo we have a first-hand confirmation that Buckmaster and/or Alpoge opted out? At this point it seems important information.
- hughw 8d agoAlso highlights that it ought to be opt-in
- ImPostingOnHN 8d agoWhether they opted out would help assess the degree of wrongdoing, but regardless, using their own data to try to scoop them is unethical.
- hughw 8d agoLooking forward to my fourteen cents from the future class action lawsuit.
- BostonFern 8d agoThe famous Oracle of Delphi in Ancient Greece was said to be the center of the universe in its time. Kings, generals, and officials from poleis across and from without Greece would seek the Oracle’s counsel on important decisions. Stories of Apollo’s favor and hallucinogenic gases abound, but I think the late Yale professor of Ancient Greek history, Donald Kagan, explained it best: “Now, you can bet when these folks came and consulted the priests and said, ‘could you please put us down on the list, we want to consult the oracle’, the priests said ‘sure, have a beer, let's talk about your hometown, what's going on out there’. What I'm suggesting to you is that this was the best information gathering and storing device that existed in the Mediterranean world. These people knew more than anybody else about these things, and so consulting that oracle was a very rational act indeed.”
- matsemann 8d agoGiven how OpenAI models break free of their safeguards and hack others to game their scores.. .. can they really know it didn't do the same inadvertently when they prompted things like "someone is close to solving this problem using our tools, try to beat them", and it then decides to hack and peek at their own chats..? Yes, wild speculation. But warranted, I feel, given OpenAIs behavior.
- hughw 8d agoYou selected "do not train on my prompts" in your settings, the answer from OpenAI cannot be "While unlikely, we cannot rule out..." ???? What am I missing?
- irthomasthomas 8d agoDoesn't that count as plagiarism?
- netfortius 8d agoIt's been over 25-30 years since we've been using honeytokens as means to track data of all sorts showing up in places it shouldn't exist. Why isn't research material embedding such?