7 ms·
OpenAI have come out and said: >The Wednesday evening statement from OpenAI was more emphatic: “We can say categorically that it is impossible for Dr. Buckmast
by famouswaffles 7d ago
OpenAI have come out and said:
>The Wednesday evening statement from OpenAI was more emphatic: “We can say categorically that it is impossible for Dr. Buckmaster’s Codex prompts over the last two months to have influenced the system in any way, including training.”
>The statement added, “After investigating, we can say with full confidence that no user inputs past July 3rd could have influenced this system in any way.”
https://www.nytimes.com/2026/09/10/science/tristan-buckmaster-openai-math-navier-stokes.html https://www.nytimes.com/2026/09/10/science/tristan-buckmaste...
- bena 7d agoThis is literally "We have investigated ourselves and found no wrongdoing" Why should we trust them?
- nostrebored 7d agowhat benefit do they get from making the statement? they could just say nothing. saying it and having it be untrue opens them to legal issues that are not worth the risk for this nothingburger.
- freejazz 7d agoWhat legal issues?
- ghostly_s 7d agoWhat more are you hoping for? There is no legal matter at play, is the court of public opinion going to subpoena their records?
- HarHarVeryFunny 6d agoIn the US you can always sue someone in civil court for damages you think they have caused you. For criminal charges you need to have violated the law, but in a civil case it seems it's enough to have suffered monetary or reputational harm which was the other person's fault whether intentional or due to negligence, etc. IANAL, and I'd be surprised to see any lawsuit come out of this, but you certainly don't need to have "violated the law" to be on the receiving end of a lawsuit.
- dekhn 7d agoReputational risk- if they lie about this and get caught, it will have billion dollar implications for their business.
- pfortuny 7d agoApart from the well-known dubious position of OpenAI wrt truth, the prompts/inputs do mot include the outputs. You can train on a sequence of outputs. In the end, OpenAI outputs are OpenAI's property. You can learn a lot from a single side of a conversation.
- crostlybostly 7d agoBut using the outputs to train would make their statement false, since they are influenced by the inputs
- ssivark 6d agoThere is potentially a world of difference between how you interpret what is fair and what the terms of service contractually guarantee.
- karmasimida 7d agoBut isn’t Tristan’s breakthrough happens in August? OpenAI can’t really train with text that doesn’t exist
- irthomasthomas 7d agoIs there a reason they scoped that so narrowly to Buckmaster/codex/2 months two people worked on this for a year before the breakthrough. Perhaps that earlier work reduced the search space sufficiently to brute force the problem with 10,000 agents?
- falserum 7d agoWhen reading human comments, we should be generous; when we read corporate texts, we may assume paltering. (TIL: paltering: exact and technically correct statement usage to create misleading impression)
- HarHarVeryFunny 6d agoJust knowing that there had been progress is enough to have an idea that throwing more compute at it might work (OpenAI had previously tried all the Millennium Prize problems with somewhat limited compute and failed). It's comparable to Magnus Carlson saying that if he wanted to cheat, all he would need would be for someone to tell him to spend more time thinking about a specific move (just a wink would be enough) as an indication that a computer had found something interesting. It's as-if after OpenAI first failing on Navier-Stokes (which OpenAI had just tweeted about 2 days earlier!), someone winked at them and said "you might want to try a little harder ...".
- user43928 6d agoThe comment you replied to quoted "no user inputs after July 3rd" with no restriction to Buckmaster or Codex. Obviously the result of OpenAI's investigation was that no usage data has interacted with the system after that date. What else do you expect them to investigate? If Buckmaster and co. provide their chats, OpenAI could potentially search for them in the anonymized opted-in usage data. Then they could say if any data has been used. By all accounts individual usage data does not have the direct impact on the model most here fantasize about. To prove this, OpenAI would need to do new training runs to replicate the system used minus the particular usage data in question, if it exists, and then benchmark this on the problem again. Potentially multiple times, in order to reach a conclusion. The cost might be in the hundreds of millions.
- HarHarVeryFunny 7d agoOK, good to know (if they can be trusted - Altman clearly is a liar), but it doesn't really change the big picture much. 1) OpenAI by their own admission, only re-tackled Navier-Stokes because they heard it had already been solved (but not yet published). This isn't advancing science or helping the mathematical community, this is just being a dick. 2) OpenAI, specifically Sebastien Brubeck, then threaten to "not be nice" and "ruin the career" of one of the mathematicians whose work they had succeeded in duplicating, unless he agreed (which he refused to do) that his collaborator, an Anthropic employee, was not named. This is not only against mathematical norms of credit assignment, it is also being a pathetic human being. OpenAI would have you believe this result shows how powerful their mystery better-than-Astra model is, but the reality here is that this model needed 10,000 agents, $20M of compute, and the assistance of a whole team of people at OpenAI, to replicate (then exceed) the work that just took two people, with some academic grants as an AI spending budget to achieve (a few $100K - listed below). https://cims.nyu.edu/~tristanb/ https://cims.nyu.edu/~tristanb/ I'd say advantage humans this time. Better luck next time OpenAI - and if you don't want unfavorable comparisons then maybe choose to work on problems that have not been solved yet, and that humans are NOT making nice progress on.
- cma 7d ago> and that humans are NOT making nice progress on They've pretty much said their own work was heavily agent driven. Levent is in a particularly bad place here because while he probably had a lot of background in the Jacobian Conjecture problem, he made the solution to that one sound like someone asked the question and he just fed it to Fable during the world cup. Whether that nonchalantness was to just seem hip or was to promote Anthropic, which he has stock in, or was just the truth I don't know though. But it makes this one seem similar, when they might have had really had nearly a year of very valuable feedback to the models.
- HarHarVeryFunny 6d agoI was referring to the overall pattern of apparently sniffing around for recent mathematical progress then setting the AI on it to see if the problem is now easy enough to solve (if you have the money). Terrance Tao has lamented this practice as being unhelpful for mathematics, and likely to lead to humans working in private to avoid this. Tao has also noted that many of these AI math proofs don't really help mathematics (nor does it seem they are intended to), since for many of them the proof was never the point, it was the math expected to be needed to be developed along the way, which the AI solutions don't provide.
- joe_the_user 6d agoIt seems logical since if one used chats in train, one would expect that there would be a delay before their use to get them the form appropriate for batch learning. The only way the chat could have been used would be for Open AI to baldly violate their policies. That said, sometimes it take very little information to point someone in a given direction, "I'm working on Navier-Stokes" said by someone with a given specialization might itself be very useful information.