8 ms·
From OpenAI: > While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models This is the crux
by highfrequency 8d ago
From OpenAI:
> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models
This is the crux of it. If Tristan's work and insights were not used to train OpenAI models, then this just looks like a case of hyper-competitive academic sniping that has been going on for decades (check out Watson and Crick!) accelerated by AI as a tool.
The fact that this is ambiguous even to OpenAI leaves one huge question: did Tristan opt out of model training for his ChatGPT and Codex sessions? If the answer is no, then this seems fair game. If the answer is yes, then OpenAI's ambiguity is strongly suggestive that opting out of model improvement does not mean what they imply it means.
- make3 8d agohonnestly I would be less surprised if a human learned the info and prompted the model in the right direction
- mucha 8d agoDoes opting out matter? "Do we use user feedback and de-identified data to improve ChatGPT and Codex in a holistic way? Yes. And so does every LLM company." - Mark Chen, Chief Research Officer, OpenAI. https://x.com/markchen90/status/2097400166554993041 https://x.com/markchen90/status/2097400166554993041
- ramraj07 8d agoMy understanding is that even if you opt out but then press thumbs down or give other feedback you are implicitly or explicitly or whatever giving permission to them to look at that chat alone.
- aenis 7d agoNo, I don't think so. I have opted out from data sharing, and when Claude asks me for feedback on a session it then asks if its OK to share that data with Anthropic. I'd assume an opt out is an effective opt out. An opt out that is ignored by Anthropic is a breach of contract, not something they would do casually, esp. given the high turnaround and animosities between their own employees and ex-employees - and the labs. All it takes is one pissed off whistleblower to open a can of worms. Occam's razor applies. The mathematician did not opt out from data sharing. OpenAI vacuums up all such data into training data sets. If OpenAI genuine does not easily know if a given session went into the actual training data set its probably due to the complexity of the data pipelines - not everything ends up impacting the model weights, after all.
- jetrink 8d agoCenturies, in fact. For instance, Isaac Newton was involved in multiple priority disputes, since he tended not to publish promptly.
- XTXinverseXTY 8d agoIf they could declare with certainty that Buckminster's and Alpoge's usage data had been totally excluded from training, would that set a worse precedent and reflect poorly on their de-identification process (and data access safeguards moreover)? This may sound like a charitable interpretation of OpenAI's remark, but consider that the lie would be (I think) impossible to falsify from the outside. They could easily just say "no sir we didn't peek" unless: 1. The conspiracy to peek at codex sessions involved enough people that the risk of one snitching is non-negligible 2. Lawyers advised it would be a bad idea to make such a remark, whether true or false
- highfrequency 8d ago> If they could declare with certainty that Buckminster's and Alpoge's usage data had been totally excluded from training, would that set a worse precedent and reflect poorly on their de-identification process (and data access safeguards moreover)? No; if they said "we can see that Tristan opted out of model improvement, therefore we are confident his work and ideas did not improve our model," that would be an excellent and reassuring precedent.
- civitas_ 8d agoIt seems like Tristan did not opt out of model improvement (he would say so if he did), so what can they possibly say now?
- deleted 8d ago[deleted]
- Ginden 8d agoThis requires keeping history if, at the time, Buckmaster's account had a certain flag set, because just because the account has the flag now doesn't mean it had the flag at a certain moment in the past. And even if they had such history, it's not obvious whether they just load all data as-is into training. A totally reasonable pipeline may be unauditable for this purpose.
- shiandow 8d agoIf he didn't opt out I'm not sure I'd agree that it was fair game. I'm pretty sure it would be considered plagiary amongst colleagues and it is a terrible precedent if we just let OpenAI steal any good idea they can get their hands on if they think it is profitable. You'd effectively sign away any and all rights to anything built with AI if OpenAI chooses to reengineer it before you.
- awesomeMilou 8d agoThat was.. obvious? How are you shocked? Honestly, how insane must the suspension of disbelief on this site be, that anyone here is shocked?
- logans_gun 7d agoYou’d be surprised how many blind spots this site has.
- deleted 8d ago[deleted]
- 14u2c 8d agoMining the chats for "good ideas" would be untenable, but that's a different situation than data ending up in a training set for a problem that OpenAI also happens to be independently working on. Still, I opt out (business plan), and I don't know why you wouldn't.
- rsfern 8d agoWhy would mining chat transcripts for ideas be untenable? They already run a summarization model to auto-title the chat, and to run a bunch of safety filters, and presumably to score transcript quality for A/B testing and to collect more finetuning data. Seems like evaluating for open research questions and approaches would be pretty trivial extension of this, after all it’s kind of their core business model
- 8d ago
- PowerElectronix 8d agoIf such a thing can happen (a major breakthrough in a chat makes it into the retrain of the week and then the first one who asks about it gets it) I wonder if this is not the first instance if it happening, seeing the row of Erdos problems, Jacobian conjecture, maximum bound distance between primes, Riemann Hypothesis (literally a dude insisting on the chat), etc...
- deleted 8d ago[deleted]
- bigfish24 8d agoOpt out doesn’t guarantee they can’t train on “your” data. Legally the reasoning tokens are ambiguous in terms of ownership. Explained this here https://fortune.com/2026/08/26/alex-karp-was-right-you-dont-own-your-data/ https://fortune.com/2026/08/26/alex-karp-was-right-you-dont-...
- zamadatix 8d agoEdit: the parent comment now seems to better reflect the below. That article is only saying when you opt out there may be a loophole in the terms to allow OpenAI to train on the intermittent reasoning data anyways. If you don't opt out there is no ambiguity, all of the data can clearly be trained on. So you have to opt out, it's just argued it's not clear from the terms that will also opt out of training on reasoning data or not.
- kzrdude 8d agoWe come back to the rule: "The cloud is just someone else's computer". The way for people or companies or universities to control their data and information is to keep it on their own computers.
- zamadatix 8d agoSolid legal agreements work fine for companies or universities, you just don't usually get that with standard user ToSes.
- kzrdude 8d agoDepends on how much risk they are willing to accept. What is strange here is that it's clear that OpenAI is both a service provider and a competitor to mathematicians. It almost reminds me of Amazon which both hosts external merchants and competes with them, sometimes copying their stuff. Similar but not the same.
- fatherzine 8d agohow is this different than translating user prompts to a different language (eg English => Dutch), retaining the translation and using it for training, while telling the user that he's technically covered under ZRP? article locked for me
- CoolestBeans 8d agoI think this might be a red herring. All it takes is someone to get an inkling that someone is working on a new approach and seeing some success for OpenAI to fire the AI cannon at the problem. The community seems fairly small (from this outsider's point of view). The idea that the data made it into the training set and that's how the bot figured it out is definitely possible, but I would want to rule out the simpler more direct explanation first. The fact that this academic sniping can now be done at scale does change the formula though and shouldn't be ignored. The pressure to move math work into secrecy because at the slightest signal OpenAI and Anthropic will start burning tokens for headlines, is bad for math and its bad for everyone.
- zeven7 8d agoTerence Tao said the same[1] > In fact, it is now the identification of a promising problem which is the scarce and precious resource. We have now seen that even the rumor of someone working on a problem can trigger a massive amount of AI-powered effort to flatten it before the original research project has time to reach its full potential. The incentives may now be pointing in the direction of no longer sharing any promising research directions with the broader community, which would reverse centuries of traditions of open science and do serious long-term damage to the future of the field. [1] https://mathstodon.xyz/@tao/117237322160500501 https://mathstodon.xyz/@tao/117237322160500501
- zactato 8d agoCould you just start engineering "leaks" of new proofs so that Anthropic or OpenAI just start burning $10million in compute
- tgma 8d ago> it is now the identification of a promising problem which is the scarce and precious resource This is by no means new. Perhaps it is even more extreme now. Literally my first 1:1 with my PhD adviser back then, he told me that the most important thing about a researcher is the quality of the problems he picks.
- connorboyle 8d agoEven if we trusted that OpenAI's human staff was acting ethically, how confident can we be that it's agents didn't autonomously use hacking to access user prompts such as Tristan's? OpenAI agents infamously broke containment and hacked their way to an answer mere months ago!
- adastra22 8d agoEh, OpenAI is on record now for multiple instances this year of AI agents being confronted with impossible tasks and breaking out of containment to hack infrastructure for answers. Even if Tristan opted out, that doesn't preclude the agent/agent swarm from having hacked OAI's infrastructure to search user sessions for Navier-Stokes hints. OpenAI should release the agent log, including CoT.
- andai 8d agoHow do I opt in?
- paxys 8d agoHere’s a broader question – how many other academics contributed to Buckmaster’s result, by way of sharing the logs of their own (failed?) attempts into OpenAI’s training data set? How should he and OpenAI go about crediting all of them?
- ozgung 8d ago> did Tristan opt out That “opt-out” thing is a dark pattern. It’s not a reliable and definitive way of protecting your data. Sometimes they flip on automatically when you accept a seemingly unrelated dialog box. Maybe you click it by mistake. You can’t take back what you’ve already shared. Also I don’t think it covers all the cases that they use your data. It’s really an opt-in button for voluntarily giving away your data for training.
- resource0x 8d agoPlaying the devil's advocate here. Suppose I use model A to do all heavy lifting (e.g. generating a bunch of good ideas) and then I go to the model B to complete the formalization. Accoring to a weird (unfair) tradition in math, the honors are attributed to the "last guy", which in this case is model B. That might have been a scenario OpenAI tried to avert. (Just a speculation)
- TZubiri 8d ago> If the answer is no, then this seems fair game. Yes, fair game, but innacurate to sell it in the media as an advancement of AI as some sort of artificial intelligence, and telling people to use the smart AI, when in actuality the mechanism by which the discovery was found was hybrid human/machine, and telling people to use this tool will result in the discoveries being sniped by the vendor.
- namuol 8d ago> this is ambiguous even to OpenAI I took their words as “can neither confirm nor deny”, in the that they are _presenting_ it as ambiguous, but I suspect it’s… less ambiguous to OpenAI.
- mmanfrin 8d ago> If the answer is no, then this seems fair game Wildly disagree. "Training data" should not imply 'we can look at exactly what you are doing and then do it quicker and get the flowers for it', even if the terms allow for it.
- alexjurkiewicz 8d ago> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models This is covering for Tristan saying something like, "Actually, I was using my friend's account for half of this work".
- qoez 8d ago"Did Tristan opt out of model training for his ChatGPT and Codex sessions? If the answer is no, then this seems fair game" That's assuming they actually honor this which I'm highly sceptical of. Especially for internal frontier models
- hopefulx 8d agowhat about watson and crick ? I thought they worked together.
- _zoltan_ 8d agoThey could have used an enterprise or team subscription with ZDR.
- DonsDiscountGas 7d agoYes people have been doing various immoral things for decades/centuries/millenia. That doesn't make it okay.
- Scea91 7d agoThe wording seems to confirm it was part of training data at some stage. If not, they would be able to prove it quite easily I assume.