6 ms·
First, OpenAI is not claiming that the model wasn't trained on those sessions. What they've said is “We (the researchers and the agents) did not see any of thei
by jjwiseman 8d ago
First, OpenAI is not claiming that the model wasn't trained on those sessions. What they've said is “We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem.” and “We did not use their prompts or proofs to prompt our models or direct our agents.” and “While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models.”
They also said “Our effort began on September 1st after hearing a rumor which we later realized was related to Levent Alpöge … and Tristan Buckmaster….” They say the rumor was that two Millennium Prize problems had been resolved, and that this prompted them to launch "an effort to evaluate it on all open Millennium Prize problems and a few other high-impact problems."
It's not obvious to me that's an unethical thing to do, if it happened as they described.
- efxhoy 8d ago> we cannot rule out that de-identified data derived from their usage of our products helped improve our models.” implied the humans sessions could have been (and probably were, why wouldn’t they be?) in the training set? If I was trying to make a model smarter and I had transcripts from the smartest mathematicians in the world I’d make sure the model trained on them.
- tedsanders 8d agoWe were also curious and we looked further into this. We've determined it was impossible for Dr. Buckmaster’s Codex prompts over the last two months to have influenced the system in any way, including training. This goes beyond what we said earlier, when we were less sure. If prompts were submitted earlier than that and training was not opted out, there may be a chance they made their way into our training pipeline in some form. But this would be a droplet in an ocean and unlikely to have made any difference, in my opinion. (I work at OpenAI.) Source for the updated claim: https://www.nytimes.com/2026/09/10/science/tristan-buckmaster-openai-math-navier-stokes.html https://www.nytimes.com/2026/09/10/science/tristan-buckmaste...
- mtgentry 8d agoThis may be true but nobody trusts your employer. The shadiest drips downward too, with the mob-like way they treated Dr. Buckmaster.
- nsagent 8d agoCan you speak to why in both cases, the problems OpenAI's models solved used the same techniques the mathematicians were exploring, which also happened to be niche approaches to the problem. As an NLP researcher myself, I find that coincidence highly suspect unless the models focused most of their attempts on the predominant approaches (they are trained for MLE after all).
- tedsanders 8d agoI'm not a mathematician and I don't want to speculate about anything I can't back up. All I know about Navier-Stokes is from my graduate fluid dynamics class at Stanford a decade ago (where I received a poor grade). However, I don't want to leave you hanging, so what I will say is: - I've heard some people say the model's solution is quite different from theirs (but I have no clue how to personally assess the spiritual truth of this, so please give it zero weight) - Thousands of agents costing millions of dollars searched for ideas, and they were encouraged to explore a diversity of approaches, so it wouldn't be too surprising to me if the approaches they tried overlapped with other mathematicians', especially considering the models have knowledge of so much published math research - This model has been beastly at solving all sorts of math problems (if it was Euler in particular, I'd agree that would look suspicious/lucky) - The Euler regularity disproof itself took ~100 agents working for ~50 hours (if it was very quick, and then the subsequent NS work took a long time, I'd agree that would look suspicious/lucky) I understand the skepticism, but from what I know internally at OpenAI, we have zero reason to believe our models did anything fishy. It's hard for us to prove a negative, especially when you have to take us at our word, so I understand why people still feel suspicious. Edit: Reminds me a bit of the Scarlet Johansson voice cloning accusations and FrontierMath cheating accusations, where the rumors of misbehavior seemed to travel faster than the truth. In both of those cases, we hadn't done what was accused, but suspicions persisted nonetheless.
- falserum 8d agoAs with all press releases I assume it was written/re viewed/redacted by their lawyers, so: > no specific user data was accessed in order to solve this problem Data was accessed in order to <other purpose> (and then accidentally used in training) Also, is llm’s answer to the prompt actually “user data”? > We did not use their prompts or proofs … So they used llm’s answers to those prompts. > … to prompt our models or directew our agents. So they trained the model on it. (Training is not prompting and plain model is not an agent)
- deleted 8d ago[deleted]
- magicalist 8d ago> It's not obvious to me that's an unethical thing to do In terms of work in mathematics, something I personally would not do based on ethical grounds would be to hear a rumor that some researchers are taking a certain approach and may be nearing a solution, use a model that was possibly contaminated with intimate knowledge about that approach (though later they investigated and think it wasn't), and then commit millions to tens of millions of dollars and untold amounts of hardware to try to beat them to it. If I had done this, I also wouldn't have pestered the researchers on a Sunday night to meet immediately so we could negotiate a nice way of presenting the actions I had decided to take. Even if you don't think it was unethical, it was never going to be received well in the community that was especially going to care about this work, and who are very much peers to many of the people working on this solution, so it was at the least an enormous (and well-deserved) own-goal that their unveiling of their solution to NS went like this.
- dwaltrip 8d agoI heard they also tried to strong-arm them into removing the name of their collaborator who happened to work at a different company (Anthropic)... I haven't looked into it myself, but if true, that seems incredibly scummy.
- luma 8d agoThe flip side is that the Anthropic researcher is clearly pushing the case against Open AI and one might have reason to suspect their motivation and version of events for the same reasons. What a mess.
- alexgoodhart 8d agoI think people are too reserved in their unwillingness to operationalize ambiguity. Ambiguity is constantly being thrown in our face, with internal audits and other laughable attestations of virtue that amount to a pantomime of transparency / good faith. Why should I care if a company claims they find no evidence of wrongdoing? Is that the threshold for privacy/trust? “We don’t care if it appears that we’ve been dishonest unless there’s hard proof.” They can simply design proof keeping to terminate at the places their dishonesty is implemented. For me, when there is a clear motive to be dishonest, a corporation should be assumed to be dishonest unless there are robust transparency measures and a regulatory environment shown to be providing a cost to dishonesty. Without it, all you do is burden yourself while the powerful entity moves ahead with its selective dishonesty and the rewards there reaped.
- pred_ 8d agoThose statements were about NS, though; I don't think they've made similar statements for the non-sofic groups?
- robotpepi 8d ago> It's not obvious to me that's an unethical thing to do, if it happened as they described. What!? Even if everything OpenAI said is accurate (big hypothesis there!), it's highly unethical to rush a solution because others have jsut had success. And that's the only beginning.
- alper 8d ago> not claiming that the model wasn't trained on those sessions The math group inside OpenAI may be training or fine tuning their own models which given some reward functions would definitely bias their usage of the training data towards things that look like math. You can launder all of it without a human "directly" doing anything.
- syllogism 7d agoYou can't take any statement like this remotely seriously. We live in a world where NSA officials can testify before congress that they don't "collect" data, because that's true under some baroque definition of "collect" that they invented and didn't tell anyone else about. Similarly you have no idea what definition OpenAI intend for terms such as "specific user data", "accessed" etc. And we have no idea what non-excluded possibilities actually did happen that they simply omit from their statement. In practice OpenAI and many others have created a situation where they're actually unable to make any credible denial of anything really.