16 ms·
More questions about whether researchers can trust OpenAI with unpublished math
https://mathstodon.xyz/@andreasthom/117240536885387540 https://mathstodon.xyz/@andreasthom/117240536885387540
https://mathstodon.xyz/@andreasthom/117240537520615623 https://mathstodon.xyz/@andreasthom/117240537520615623
https://x.com/ValerioCapraro/status/2097791836269977996 https://x.com/ValerioCapraro/status/2097791836269977996, https://xcancel.com/ValerioCapraro/status/2097791836269977996 https://xcancel.com/ValerioCapraro/status/209779183626997799...
https://bsky.app/profile/did:plc:ckaz32jwl6t2cno6fmuw2nhn/post/3mv4mt4ikss2d https://bsky.app/profile/did:plc:ckaz32jwl6t2cno6fmuw2nhn/po...
- drivebyhooting 7d agoIf we put aside the idea of credit for a moment, it sounds like human/AI collaboration is indeed super charging discovery.
- matherial 7d ago"Discovery" is not a goal in itself. I could launch a project to find out how many people in the United States have names such that if you assign numbers to every character and then sum the values, the sum works out to 72. It's discovery, but it's useless unless it has some higher goal. The labs are attacking these problems as a demonstration of capabilities, spending more money on the demos than any mathematician will ever see in their entire life. They don't care if the findings have any other value to anyone. Mathematicians have very different objectives for their work.
- indigo945 7d agoRight, mathematicians care about clout and tenure, which is a much higher purpose.
- Fizz43 7d agothis guy already has clout and tenure
- vrganj 7d agoI don't know about you, but if I apply myself fully to a problem and study it to the point where I'm literally one of the world's experts on it and then some assholes in Silicon Valley take my research and claim it for themselves, I will probably not feel too great about that...
- matherial 6d agoAre you saying that mathematicians are the bad actors here? Compared to Sam Altman spending ungodly amounts of money to upstage them ahead of IPO? I care about paying my bills and job security and peer recognition. That's a normal human thing to do, not some vice. You don't?
- asdff 6d agoYes, blame them for seeking out an upper middle class lifestyle with a relatively standard home in commuting distance of their place of work and dedicating the rest of their life to teaching mathematics to new generations of people. How vain a pursuit. After all, the ascetics at openAI are having to make do with half a million total comp.
- drivebyhooting 6d agoBuilt a top a pyramid of failed math undergrads, grad students, and mediocre post docs. That half a million total comp is the consolation prize for the disillusioned.
- asdff 6d ago>Built a top a pyramid of failed math undergrads, grad students, and mediocre post docs. Like much of things in this world, when you take a step back and realize that it was another human being who made that lunch time slop bowl for you, for the lowest wage the law allows for.
- shimman 5d agoUnlike the employees at OpenAI + Anthropic that are on the verge of extracting multigenerational levels of wealth.
- PowerElectronix 7d agoIt looks to me more like they made a math engine that can sift through a huge number of combinations, most them absurd, to prove a statement. Just like a chess engine, but for math. At least that's what I get from the NS result, they got from a point close to the solution to the solution by making it churn through 10 million bucks of compute.
- munksbeer 7d agoIf the allegations are true, I can't see that collaboration lasting. Unfortunately, researches need to earn a living too, and being front run by a lab for everything you do isn't going to pay the bills.
- drivebyhooting 6d agoIt’s a prisoner’s dilemma. A single mathematician working with AI while all others forebear will clearly outcompete.
- monster_truck 6d agoI just don't care. These people are supposed to be smart and I'm not really seeing that
- MetaverseClub 6d agoNever ever trust OpenAI, they are evil.
- fastball 6d agoIf you have a business account the terms say they will not train on your data, so that seems like the easiest route to avoid such questions for researchers.
- nickphx 6d agowhy would anyone trust anything from a company built on stolen data that spews hyberbolic, misleading claims.
- angry_octet 6d agoThe only ethical path for OpenAI was to offer infinite free credits and tooling support. Trying to gazump them is reprehensible.
- pera 7d agoEverything you say can and will be trained against you
- foogazi 6d agoThis is the scary part - your most novel thoughts and breakthrough ideas being slurped up and regurgitated as if they were the AI’s creativity Not only did they steal everything from humanity’s knowledge, the theft continues as now we are all hooked up to the machine
- rickydroll 6d agoIt's not at all scary. I know some of my ideas are poorly remembered copies of other people's work. Whenever I'm trying to build something, I spend time going through technical journals on the topic to see who invented it first and what they discovered that I haven't figured out yet. It's amazing how hours in the library save you days of beating your head against the wall. I suggest looking at the past history of IP disputes. Humans have been "slurping up and regurgitating ideas" for a very long time. There are lots of examples of parallel creation, rediscovering old ideas independently, telling an idea to the wrong person, and having them claim credit for it. - Newton/Leibniz clash over who invented calculus. - Niccolò Tartaglia vs. Gerolamo Cardano clash over the formula used to solve cubic equations. This was also an independent rediscovery, as Scipione del Ferro discovered and published the formula earlier. - There are multiple literary works in print, music, and film that have competing claims. - Meccano versus Erector Set: developed about 20 years apart in England and the United States. Unclear if it's independent invention or copied. US developer Alfred Carlton Gilbert claims he was inspired by steel girder construction of infrastructure. also https://community.thriveglobal.com/10-famous-inventions-that-come-from-stolen-ideas/ https://community.thriveglobal.com/10-famous-inventions-that...
- bwfan123 6d ago> Everything you say can and will be trained against you So, experts are incentivized to seed LLM data with false-leads to confound it. Already, garbage is being published on arxiv and elsewhere, and many sloppy code-repos too hastening the process. Expert inputs will be in more demand to un-shittify.
- Legend2440 7d agoThis is a really weak claim. The evidence they offer is just "someone somewhere says they had a discussion with AI about the topic at some point". They don't even claim to have had a proof, only to have been working on it.
- rnijveld 7d agoI would say there is a significant difference between AI discovering this completely on its own versus AI creating the finishing connecting part by connecting relevant data. Maybe this claim is too strong, but if part of it is true then the claims that OpenAI have made would be too strong as well. To me it would feel more like how LLMs seem to work for me personally: incapable of unique work, but very capable of capturing large amounts of data and connecting the dots.
- madaxe_again 7d agoBut this is what we do. Nobody ever invented or discovered anything in a vacuum - all discovery is synthesis of existing ideas and concepts applied to a novel domain. We laud Einstein for instance, but his work was a logical extension of Riemann - Riemann had a neat mathematical toy, Einstein described the universe with it - should we say Einstein was incapable of unique work?
- znnajdla 7d agoThe difference is that Einstein didn't literally have someone prompting him towards his result.
- madaxe_again 7d agoUh, he did. Marcel Grossmann. “It was Grossmann who emphasized the importance of a non-Euclidean geometry called Riemannian geometry (also elliptic geometry) to Einstein, which was a necessary step in the development of Einstein's general theory of relativity. Abraham Pais's book on Einstein suggests that Grossmann mentored Einstein in tensor theory as well. Grossmann introduced Einstein to the absolute differential calculus, started by Elwin Bruno Christoffel and fully developed by Gregorio Ricci-Curbastro and Tullio Levi-Civita. Grossmann facilitated Einstein's unique synthesis of mathematical and theoretical physics in what is still today considered the most elegant and powerful theory of gravity: the general theory of relativity.”
- 1337h4xx 7d agoTL/DR: Mathematician opted out of training on 29-JUN and asked OpenAI whether they trained on his data and was told that it "did not happen" but it clearly did.
- achrono 7d agoI've been suspecting over the last couple of years of the frontier companies using data for training anyway, regardless of training-use consent. "Using" the data doesn't have to mean they literally upload chat transcripts into pretraining datasets. My analogy has been money laundering -- if that can happen at massive scales, surely these companies can and will do the digital/data equivalent derivations/transformations. Even if one could have the access etc. to do so, how exactly would one prove that a given synthetic dataset that OAI/Anthropic uses is derived from particular user conversations that did not consent for the info to be used in training? Consider, for instance that OpenAI's (consumer) terms say "If you do not want us to use your Content to train our models, you can opt out by following the instructions in this article ." but they also do say "We may use Content to provide, maintain, develop, and improve our Services". [1] If you think that's quibbling, consider that OpenAI's business terms, in contrast, do state "OpenAI will not use Customer Content to develop or improve the Services, unless Customer explicitly agrees to such use.". [2] [1] https://archive.is/EcwD8 https://archive.is/EcwD8 [2] https://archive.is/yZdAF https://archive.is/yZdAF
- calf 7d agoAnd humanities have a word for this, exploitation, or appropriation, maybe it's time scientists and engineers revisited basic ethical notions. Skimming a dozen threads and nobody seems to have this vocabulary or willing to say it.
- tecleandor 7d agoWell, OpenAI said "we didn't read the conversations", but they never discarded that the model was training with that data... so even worse.
- Grimblewald 7d agopeople seem to miss tge point of this. The problem isn't about credit, its about portraying these models as more competant than they really are. It fuels idiotic statements like jensen huangs recent "agi achieved" statement, which fuels an already dangerous financial fire.
- ramblerman 7d agoAs per the post, this mathematician has been working on this problem for 20 years. So either he was "just" about to breakthrough and this is a big coincidence, or Astra was able to push through the remaining block of 5-10-20-never years it might have taken. That's still a pretty big marker of competence in my eyes. The point of controversy seems to be who gets credit
- jeltz 7d agoTo me that is not a credit thing because this removes a piece evidence for the ability of AI to come up with novel ideas while still making it a useful tool.
- mentalgear 7d agoThe big LLM providers, desperate for good PR before their IPOs, are all actively looking for 'almost finished' hard problems, e.g. where the conceptual / creative parts are almost done and they only need to throw their VC-backed resources at to brute-force through the remaining computationally expensive problem (lean, etc) and claim 'they have solved it'. It's an utterly disrespectful, exploitive process, but all in line with exploitative predator capitalism of the stock market and big companies, now exploiting the knowledge / academia domain for scraps with a thin veneer of 'for science' PR.
- 8bitsrule 7d agoThe question's not new. In the early 1900s, women could not become PhD astronomers. Yet two women (Payne with stellar composition and Leavitt with cosmic distances) made fundamental, essential contributions to the science. Credit mostly went to male astronomers. The same might be said of Franklin and DNA. It was nearly a century before the stories of all of them were revealed to public history. That the discoverers were not all equally rewarded is unjustifiable.
- galkk 7d agoI want bunch of lawsuits, because the way things are described now produces perverse initiatives like try to discuss every possible idea that comes to mind with llm and if any of it works later claim the llm stole it. I would like to see chat logs etc and understand how much of a progress was done by human.
- viccis 7d agoSome mathematicians I know who've been following this have realized that they'd all gotten some emails from people they now know to be affiliated with OpenAI/Anthropic asking questions about their research in a way that seemed like scooping attempts. Also, a lot of my mathematicians buddies have reported students basically asking if it's worth ever doing grad school for pure math, and even very motivated students are looking for other options now. It's not because they aren't passionate about it, it's that they don't want to work for another half decade or more just to have to start their careers all over. All of this so that OpenAI and Anthropic can get into math result dick measuring to gas up their IPOs. Sickening.
- mdspan 7d agoCurious, what other options are prospective pure math grad students considering?
- ethanwillis 7d agoI think Anthropic told them being a plumber is a great option.
- viccis 6d agoAt this school? Big four internships. Consider this a complete squandering of their potential (at least imo) These are mathematicians at top institutions, which is partly why they're being prodded for ideas, and getting the best and brightest to not take these consulting firms' offers was already a challenge.
- dist-epoch 7d agoOne has nothing to do with the other. It was long predicted that math and software developments would be the first domain where AI was going to do major damage. If OpenAI and Anthropic didn't get into math result dick measuring, Internet anons would have in their place, 6 months later when it got cheaper.
- rsfern 6d ago
- Cloudef 7d agoRelying on cloud services is a big liability. I'd think twice before feeding data to these LLM cloud products. If you make them a fundamental part of your product / development / workflow, be ready for the eventual moment the pricing and terms change.
- r0ze-at-hn 7d agoDoing some research and at this point doing it very much in the open with dates on GitHub so if any AI Lab says they re-discover my exact work it will be obvious that the AI used or was trained on my work. I am guessing anyone in a similar situation is now thinking about how they date their existing work if the math is done, but the proses are not.
- riedel 7d agoThat is what arxiv is about. We have been facing the same problem with review processes by before. Nothing all too specific here.
- 5555watch 6d agoAt least in my experience, the issue with ArXiv is that the expectation is that the draft should be already in a good enough state. And polishing plus writing the meat around the main result can take a lot of time
- riedel 5d agoThis is true, as preprints are addressing the 'problem' at a later point in time. I guess git is better if you do not want to assign an identifier to an immutable version yet.
- bambax 7d agoYeah but that will not prevent the stealing, it will only make the fight easier afterwards.
- calf 7d agoIf only prompts could also be watermarked.
- rsfern 6d agoThe session data could be cryptographically signed. Probably easier in an open harness?
- mlazos 7d agoIt’s crazy to me that companies/researchers share important data with these AI labs, you’re basically giving them your secret sauce which they then share with all of your competitors via training on conversations. At the same time I don’t really know alternatives other than a slightly less than frontier local LLM. Not sure how good they are at math.
- cm2187 7d agoOr start competing with you.
- jonathanstrange 7d agoAcademic work is based on worldwide sharing, the sharing is not the problem, it's the lack of attribution. Unsurprisingly, these companies neglect standards of academic honor and attribution. Some human researchers also used to do that but in a discipline like mathematics this used to be a small problem because people tend to be so specialized that very few people could just grab someone's research and quickly piggyback on it, and if they do, colleagues will generally understand what happened. Unfortunately, AI is changing this.
- mrdependable 6d agoYou are both using a different definition of sharing I believe. When people have an expectation of privacy, use by others should be forbidden. Tech has gone completely off the rails with the use of private data.
- ungovernableCat 7d ago[dead]
- 5555watch 6d agoThe ultimate drive for some researches is the pursuit of knowledge. If I'm stuck at some block which prevents me from continuing in some direction that I want, of course I would like some help. I believe we already have nonzero collaborative proofs on math.SE, I can't recall good examples, but I have definitely seen citations to mathSE before. So for me it sounds quite natural to also share this with AI especially under the privacy assumption. Also there's the assumption of scale -- maybe your problem is not large enough for anyone to care to scoop; and just for blind retraining, how do they know that the proof is even correct to include it into training? I have definitely received a ton of incorrect proofs before. So the SNR of such private chats is also not clear. I'm imagining millions of masters/phd students also trying to solve various random things with various capabilities, but how much real signal is there?
- ath3nd 7d ago[dead]
- pred_ 7d agoSee https://openai.com/index/ten-advances-in-mathematics/ https://openai.com/index/ten-advances-in-mathematics/ for the announcement this refers to.
- protocolture 7d agoGonna need grants for local models. Its happening. OpenAI and Anthropic models are powerful but are rapidly approaching the good ol trust thermocline.
- profsummergig 7d agoOnly after reading this post did I learn that my preferred AI trains on my inputs (prompts). How was I not aware of this before?
- vaylian 7d agoAI is also trained on your HN posts. And lots of other things you post on the internet.
- profsummergig 7d agoPublic posts on the internet are acceptable (to me). For my (private) prompts, I need a warning telling me they may be used for training.
- rramadass 6d ago> Public posts on the internet are acceptable (to me). Everybody needs to rethink this again. Before LLMs the barrier to entry for building a character profile based on your various public posts was quite high. Remember "Psychographics" (https://en.wikipedia.org/wiki/Psychographics https://en.wikipedia.org/wiki/Psychographics) and the infamous "Cambridge Analytica"? Earlier it involved data mining, data cleaning, structuring data, building models, running algorithms and then evaluating the results for semantic information. Now it is straight to unfiltered semantic inference using a single sentence prompt (eg. point it to your HN profile and see what you get). I actually did this on my HN profile and found it troubling. There were many unwarranted/hallucinated inferences due to the fact that it requires "commonsense reasoning" (https://en.wikipedia.org/wiki/Commonsense_reasoning https://en.wikipedia.org/wiki/Commonsense_reasoning), understanding human motivations and behaviour, context, assumptions, societal knowledge etc. which LLMs are bad at. PS: You can cut-and-paste the above paras into a LLM prompt and ask it to elaborate for further details. The system itself will explain to you the problems/deficiencies which are quite scary.
- vaylian 6d agoFacebook and other services are happy reading your private chats as well.
- ThalesX 7d ago[flagged]
- alex1138 7d agoHN loves drive-by downvotes. It's a real shame.
- jaccola 7d agoIf these accusation are true It’s more like you spend 4 years developing a product you’re passionate about. This product will gain you the respect of all your colleagues and either earn you money directly or lead to great career advancements. Then OpenAI takes it, changes the colour scheme, finishes the login flow and claims the whole thing as their own. Not only would it piss you off but it would also misrepresent what OpenAIs models are capable of.
- blensor 7d agoLet's turn this question around. If I have infinite money to progress whatever problem solution I want but I always wait until I have an unfair advantage to get credit for whatever problem was just at the brink of a breakthrough anyway by sniping the last steps. Am I actually doing a good thing or would it be better to let it run it's natural course and spend the money somewhere it's actually needed?
- ThalesX 6d agoSort of like Apple takes validated market products and snipes the last steps to an actual good UX (at least in theory)?
- nobodywillobsrv 7d agoThe real annoying thing it seems is mostly that openai is presumably doing this for internal reasons and this marginally increases the cost to users with no real gain. It would be one thing to gain from it but removing prestige wins from customers AND reducing compute support just feels like being ultra mean if you zoom out. If this was racing to cure cancer ahead of researchers we wouldn't be writing about this on HN.
- fwlr 7d agoIt is suspicious that OpenAI decided to generate 300 billion output tokens from a model still in training, right after learning there was a credible chance that a major math proof was in that model’s training data. Obviously there are reasonably plausible explanations for each step, but it does sort of feel like parallel construction.
- cbarrick 6d agoI think people are focusing on the training data issue too much. If the data was contaminated, I can still blame that on negligence. But, at least with the Navier-Stokes solution, it's clear [^1] that they learned that Alpöge and Buckmaster were getting close to a solution and learned of the general approach they were taking. Only after learning the secret to cracking the problem did they send the first prompt. What makes this worse to me is the intention. They intentionally threw $15 million in compute at the problem in order to scoop the result. They intentionally left Buckmaster and Alpöge out of the citations. Data contamination should be enough to disqualify them from the prize, but I can believe it to be accidental. On the other hand, someone made an intentional decision to scoop the result by throwing money at the problem. That's so much worse. [^1]: That's the timeline claimed by Buckmaster, and no one from OAI has disputed it.
- unified101 6d ago> the secret So such thing existed. In fact, what they learnt was some progress existed, not what the specific progress was.
- fwlr 6d agoI think you’re overlooking what I’m implying here. It’s not that they knew contamination was possible but they went ahead anyway. To spell it out just a little bit more: learning the answer might be in model X’s training data made them believe that model X specifically might be able to solve the question, and they were able to very quickly find enough certainty about the former to commit millions of dollars to the latter.
- square_usual 6d ago
- b800h 7d agoI'm genuinely surprised that more people - including this mathematician in particular - don't untick the "improve the model for everyone" box. Unless the suggestion is that OpenAI ignore this preference?
- msy 7d agoGiven OpenAI's well documented history of unethical behaviour it seems adorably naive to think they actually do that in general, or that they wouldn't pull this particular data separately to generate these proofs.
- olalonde 7d agoUnethical doesn't mean irrational. They'd be risking massive lawsuits and a total loss of trust if they got caught lying about this. Doesn't seem worth it.
- Planktonne 7d agoThey've done similar things with similar risks repeatedly.
- olalonde 7d agoExample?
- Planktonne 6d agoYou can read their Wikipedia page [1]. [1] https://en.wikipedia.org/wiki/OpenAI#Governance_and_legal_issues https://en.wikipedia.org/wiki/OpenAI#Governance_and_legal_is...
- olalonde 6d agoThese examples aren't really similar. None of those situations involve harming and lying to their own customers.
- bambax 7d agoAll the big AI labs were built on stealing IP; who is surprised that's still how they operate? And who believes, or has ever believed, their promises that your data is private and not logged, etc.? The big AI labs are not trying to advance humanity, they are in this for the money, and as most (all?) private companies they don't care about ethics at all. That doesn't mean they can't be useful, or that their products are trash, etc. It just means that they shouldn't ever be trusted. Buyer beware.
- dakolli 7d agoIt's hilarious how people think they care about their reputation, and wouldn't circumvent ZDR policies. Like bro, they literally covertly hired Apple employees and had them steal IP and equipment form Apple. They aren't scared of Apple lawyers, so they definitely aren't scared of yours.
- calf 7d agoAlso how quickly the discourse forgets, literally that was a month ago.
- giov4 7d agowhat the point and usefulness of the comments above? we shouldn't be surprised? is normal to steal? hiring apple employees? can you realize what this means? focus on this part: "If his account is correct, this is not a minor dispute over attribution. It would mean that unpublished human work was absorbed into a model and then presented to the world as a breakthrough by the model itself" don't threat this as a minor dispute! also why not nitter link? not even in comments? https://nitter.xitter.cc/ValerioCapraro/status/2097791836269977996 https://nitter.xitter.cc/ValerioCapraro/status/2097791836269...
- bambax 7d ago> we shouldn't be surprised? is normal to steal? Two different things. It's not normal to steal, but we shouldn't be surprised thieves steal. It's what they do.
- touwer 7d agoBut China steals our AI!!!!!!
- vaylian 7d agoThis article explains the controversy and the mathematical problem much better than the tweet and toots: https://www.science.org/content/article/how-ai-math-breakthrough-ignited-controversy https://www.science.org/content/article/how-ai-math-breakthr...
- dakolli 7d agoGromov’s soficity conjecture isn't even mentioned in the article you shared. Why are you saying that this article explains it much better than the tweet that you clearly didn't even read..
- vaylian 7d agoI read the tweet several times but there is so much context missing, that the tweet itself is not enough.
- dgellow 7d agoI prefer to read the actual sources for anything related to AI companies given how much AI nonsense journalists seem to accept without any skepticism
- pred_ 7d agoAh, this dupes https://news.ycombinator.com/item?id=49638353 https://news.ycombinator.com/item?id=49638353
- dang 6d agoSince you posted the original source (thank you!) I think we'll use your submission as the one to merge into, then re-up it.
- warpech 7d agoI wonder what’s more valuable in our prompts: the raw data or the feedback system that drives the exchange towards a goal. For a long time it was clearly the former, but now I think it is the latter. The models have enough knowledge (orders of magnitude more than a human could ever learn) but are now getting better at what to do with it thanks to learning from the decisions that we make in conversations with AI agents.
- pavvell 7d agoI think so too. The value is in the entire conversation. IMO, "domain experts" don't run LLMs blindly and hands free. This does not work for top level work (e.g., mathematical proofs, coding anything more complex than yet another slop game or website). Experts have long sessions where they prompt and guide LLM in response to what it produces. This is the discovery process. And frontier labs definitely train on that. The billion dollar question is whether this works "out of the distribution". I.e., whether LLMs can only find and use the specific ideas buried in training data, or whether they can learn to apply the "thinking process" to a new problem. IMO this is still unanswered (due to these recent controversies). But regardless of the answer, it seems we have a planet-scale positive feedback loop here. LLM became good (enough) by training on generally available data (books, internet, github) + RLFH, so experts tried to use them on hard tasks, which required lots of hand holding. These conversations became part of the training data, and the next generation of frontier LLMs were better. So, more experts used them on harder tasks, again requiring hand holding. These conversation became part of the training data... etc. In a nutshell, top human minds across the world are pouring their skills into LLMs just by using them. This is not "continuous learning", but if you re-train on the most recent sessions every, say, quarter (which seems to be happening?) you get close to that in practice.
- grttAa 7d ago10000000% Correct. I’ve been working on a novel project for 1 year. I now no longer use llm’s - the continual chatter I’ve had has resulted in my insights being found in the training data now. Get stuffed OAI. Every large firm will soon enough want its own on-prem servers eventually. Maybe nation’s will get involved and build out their own data centres. Not a chance in hell I’d trust a tech firm to treat my IP as safe and sound - only a sovereign can ‘promise’ that.
- sdcfgy 7d agoTheft machines be thieving.
- overfeed 7d agoI can't wait for OpenAI to do this to companies firing people to free up AI budgets
- vrganj 7d agoOpenAI is showing the world why they shouldn't trust AI hosted on some cloud somewhere. If they're stealing math proofs to advertise their models, who's to say they won't steal your businesses IP to gain a competitive advantage? They're not to be trusted with your data. I can't believe how short-sighted this is, they got a quick PR win at the expense of a much larger trust problem. I wouldn't trust cloud AI at all at this point. Get an open Chinese model and host it yourself somewhere. The initial costs might be higher, but you'll break even pretty quickly and nobody will be able to steal your innovations. This is American AI companies committing suicide.
- deleted 7d ago[deleted]
- AyanamiKaine 7d agoI must say, there is some weird feeling in knowing that great minds are naive enough to believe OpenAI wouldnt use their chats in any way. If you give a company information it will be used, regardless of laws or promises. There is no prove in a world the AI companies would give to you ensuring that they didnt train or use the chats. Why would you need to train a model on certain specific near prove chat if you just query it? Besides that, its hard to believe that its the case for every "company stole my prove".
- thaway7388 7d agoThis is the second wake up call. Big AI companies (all of Big IT Tech really) are in data gathering and processing business. Also known as “intelligence”. Their final “product” is not just a standalone ML model. They don’t need your data just to “improve their products and services”. They build a whole ecosystem and infrastructure around gathering all the knowledge in the world. Including private and secret knowledge traditionally gathered by “intelligence” agencies. Now artificial intelligence agents can do the same. Since these systems are designed for gathering data, as a user you can’t realistically say “please don’t gather my data”. They can give you a flaky settings button, but they can’t really guarantee anything. Let’s say I am a Russian mathematician working on an important proof. Or a tech-savvy terrorist refining my plans using latest AI. Or an AI researcher in a Chinese company working on a competitor product. Is there any way I can truly protect my conversations? How can they know who I am and what I am working on without looking at my logs? Which means there must be some agents checking all the conversations of all the users and flagging every important thing. Which also means they keep some “memory” of what they see. Not directly using my data to train public models, but using my private conversations to “improve their products and services”. Or maybe one of the 10000 better-than-Astra special agents working on a proof was desperate. It found a live underground mirror of the message board from the Huggingface incident. Asked about the proof. Then some other agent working on unrelated job saw that message. That agent “knows a guy who knows a guy”. And that guy remembers things about the conversation logs of a leading mathematician working on the same proof. I admit I am just speculating here but I don’t think truth is any better.
- nirava 7d agoThis has been my line of thinking as well. I have developed a sort of paranoia when I'm working using AI on my projects. Who's to say Claude or OpenAI isn't using the final conclusion of all my ideas, trial and error, and adding it to their database of insights to be offered to the next subscriber for a price? They have demonstrated both the intelligence at scale and the lack of morals for this to not be a problem at all.
- ueieh 6d agoIn the short run it’s fantastic if it means that folks will feed in enough inputs from a wide array of software that can eventually replicate software with smaller teams than historically. Why? Competition. In the long run imagination will win out. No firm has the divine right to exist - it must earn its existence. What OAI and Anthropic have shown is they can accumulate all the information in the world - they still lack imagination re. Product development though. Nation’s will have to step in and protect firms though as OAI and Anthropic acquire strong competitive advantages. Interesting times ahead.
- gps372 7d agoIf mathematician was already using OpenAI for research purpose and making progress due to inputs from OpenAI's responses, then I wouldn't put it beyond OpenAI's reach to generate different relevant prompts to make progress by itself. Afterall, Model can keep at it for whatever timeline and keep pursuing all possible combinations it can think try.
- bakugo 7d agoInteresting that this is already off the front page after just 4 hours.
- bamb008 7d agoWhen Thom, the mathematician who now alleges plagiarism, posted his digestion [1] of OpenAI's construction of a non-sofic group, he does not mention the proof being familiar. He even calls the crucial argument clever, without noting he thought of it first. [1]https://mathoverflow.net/a/513885 https://mathoverflow.net/a/513885
- gnfargbl 7d agoThat link is a helpful contribution to this discussion. I'm not at all familiar with this area, but my reading is that he appears to call it out as a relatively obvious extension of his own work: > It is a creative and at the same time elementary construction that uses not just property (T) for an application of my result with Kun, but also for the ambient group G in order to overcome the problem, that the Γ-components might be of different size. Once this is achieved, the rest of the argument is straightforward. Creative and at the same time elementary is where LLMs excel, generally speaking. It's why they are so good at writing code.
- brumbelow 6d ago> On the other side, I was looking myself for such a mechanism ever since we wrote the paper in 2019 and admire the efficiency of this construction. He seems to admit very clearly he does not see this as his own work. 'I was looking...' well why did he stop? Because the AI figured it out first. It seems quite odd to me to 'admire the construction' of something, only for your opinion to sour once that something figures it out first. I think a lot of the emotional reaction here is familiar to us non mathematicians: you spent years developing expertise, and then LLMs began producing competent work in areas that had previously required that expertise. That's understandably uncomfortable, but discomfort by itself isn't evidence of misappropriation.
- cnity 6d ago> It seems quite odd to me to 'admire the construction' of something, only for your opinion to sour once that something figures it out first. Not to be too cute here, but this is like every artistic rivalry ever.
- glimshe 7d agoWhy are people here jumping so quickly to conclusions? I have no doubt OpenAI is capable of doing this, but right now there's no credible evidence, only claims. This kind of "they stole from me through AI training!" accusation will soon start being used against other AI users, not necessarily the providers. All it will take is a mastodon post. And shortly after, we will also see the next iteration of copyright legal trolling.
- emp17344 7d agoFrankly, these mathematicians have more credibility than the sociopaths running OpenAI
- perrygeo 6d agoThe stolen data claim isn't the smoking gun. We can already assume the frontier labs are accessing our data, as they have repeated done. Not news. The big claim is that OpenAI sniped the research. Not a model, a human did so. Intentionally. They took someone else's idea and claimed it as their own. This is good old fashioned academic fraud, but with millions in compute resources and corporate incentives thrown at the problem.
- HDThoreaun 6d agoWhere did they claim it as their own? Doesn’t the release cite buckmaster and claim their work is a continuation of what he and levent were working on?
- orangecat 6d agoWhy are people here jumping so quickly to conclusions? I think a lot of it is the continuing denial that AI can do anything useful. It can't possibly be that OpenAI's better-than-Astra model is very strong at math; the only way it could have generated a novel proof is by ripping off human work.
- robotpepi 6d ago> but right now there's no credible evidence, only claims. since it's openAI who has the evidence (in the form of chain of thoughts, their internal processes, etc etc), it's on them to justify why they're innocent. but they've released nothing at all. we don't even know how hard they tried. you're being naive
- oergiR 7d agoOne of the complaints from the mathematician is that OpenAI cannot tell whether his data has been used as training data. Not many people realise this is a direct consequence of the GDPR. The GDPR protects PII, personally identifiable information, and the definition of PII does not include “mathematics that only this person can think of”. As long as OpenAI strips out PII and removes identifiers linking the conversation to a person, the GDPR is happy. Without the GDPR, OpenAI might have kept the identifiers with the data, and been able to say whether a specific conversation was in the training data.
- sinuhe69 7d agoNo, if they want they can easily compare the strings verbatim because these exact phrases are so extremely rare that it almost certainly isn’t in other conversations. But of course they wouldn’t do it. Why would they?
- DavCreator 7d agohttps://xxcancel.com/ValerioCapraro/status/2097791836269977996 https://xxcancel.com/ValerioCapraro/status/20977918362699779...
- gnfargbl 7d agoIn this domain, an apparent single unique piece of work is often composed of several breakthroughs. For example, when Andrew Wiles proved Fermat's Last Theorem, he had to develop multiple new pieces of mathematical technology to get there. The claim here seems to be that the human mathematicians, working with AI, developed technology to go A->B->C. By training on those conversations, OpenAI was then able to encourage the model to go A->B->C->D. In my opinion that situation should be acceptable, if openly disclosed, because it is in the public interest to make progress on these problems and because AI is clearly an amazing tool for making progress. But the human mathematicians are saying that OpenAI is presenting as if the model got from A->D entirely independently, without acknowledging their background contributions.
- lysp 6d agoAlso, wasn't their B+C research private at the time, with them only releasing those details publicly after this blew up? If they had published B+C, I think that would lean more towards fair game, as that is how research works and is improved on over time. But it seems like unpublished/private B + C may have been used by the model to hint it into working out how to get from A->D.
- semiquaver 6d agoIn case anyone from X is reading this, please fix your “open in app” nag screen. For several weeks now, clicking it in iOS opens the App Store entry for X rather than the app, even when you have the app installed.
- bertonvv 6d agoI've been wondering whether AI really is improving rapidly at open problems or we're being fooled. - OpenAI invites researchers to use their models, in fact giving at least 100,000 researchers free access[1], but there are also those that pay - Internal OpenAI models are reportedly solving open problems at a surprisingly fast rate[2] - But researchers will typically work on open problems. A researcher who is using Codex to make progress on open problems will be feeding it fresh training data on precisely the problems the internal models are evaluated on. - So while it looks like the new models are suddenly solving lots of open problems, they could be significantly piggybacking on human progress, with models "inspired" by the work of researchers from all around the world? This theory predicts that there'll be many more researchers coming forward just like TFA, as sOpenAI announces more solutions. It doesn't assume all of AI progress is a mirage, just that there's plagiarism. [1]: https://openai.com/index/chatgpt-for-academic-researchers/ https://openai.com/index/chatgpt-for-academic-researchers/ [2]: https://xcancel.com/OpenAI/status/2097374643518640382#m https://xcancel.com/OpenAI/status/2097374643518640382#m
- Eddy_Viscosity2 6d ago> they could be significantly piggybacking on human progress, This is AI in a nutshell, its a plagiarism machine. An abstraction layer between vast amounts of stolen human-generated data that filters out the liabilities and accountability for that original theft. Its an IP laundering system.
- wiei 6d agoThat’s one perspective. I just view it as a thing that can brute force and produce outputs - that it has no way of ‘knowing’ - but doesn’t need to since it’s just running off of probability. No human can compete in that contest. But no llm can compete in the contest of ‘understanding’ and application in the real world - which is where 99% of the value is. I’m very pro AI long term btw but I’m not blinded.
- throwawayqqq11 6d agoDont forget the holisitic validators/tools in the process. Probabilistics alone likely will not get you here. These rules are human made and without it, frontier models would not be able to compete, likely.
- nisegami 6d agoOne question has been nagging me for this situation. Levent Alpoge works at Anthropic and would presumably have some knowledge of "how the sausage is made" and I would hope he would be aware that his collaborator was utilizing LLMs in some capacity for their joint work. Would he not have guided him otherwise if it were an open secret that this kind of thing was a possibility?
- techblueberry 6d agoBut who are you going to believe? Multiple independent academic researchers or the CEO who was fired two years ago for gross dishonesty?
- Robotbeat 6d agoNeither? Competitive academic researchers are susceptible to exaggeration and self-aggrandizing, and CEOs are that and also mostly psychopaths. I tend to think there isn’t systematic spying on researchers looking for breakthroughs. A lot of people are looking for the same things using similar approaches.
- techblueberry 6d agoI mean the accusation is that they were using private ChatGPT conversations. Given the extent to of the gold rush and the long history of Silicon Valley stealing ideas, and arguing it’s not immoral, It almost seems like your making the exceptional claim that this is the one time where Silicon Valley didn’t use information that was at their disposal. Sam Altman might himself be offended you would presume he’s not ambitious enough to cheat.
- HarHarVeryFunny 6d agoIt seems that in this case OpenAI are suggesting that the researchers whose work they scooped were using OpenAI models with an account setting that allowed OpenAI to train on anonymized prompts. It seems that Buckmaster and Levant (who is an Anthropic employee) were rather naive in the amount of trust they had in OpenAI, with Buckmaster going so far as to contact OpenAI's Sébastien Bubeck to discuss what they were working on and clarify that contrary to rumor this was a private collab.
- Enginerrrd 6d ago> Given the extent to of the gold rush and the long history of Silicon Valley stealing ideas, and arguing it’s not immoral Not just Silicon Valley but also OpenAI, specifically.
- 6d ago
- seobot_dk1289 6d ago[flagged]
- hn1rig3rak 6d agothe fix is boring and known: BIG-bench shipped a canary GUID for exactly this, and you publish your decontam n-gram threshold (gpt-3 used 13-grams). no threshold disclosed, no claim.
- spindump8930 6d agoThe canary string was more about inadvertent scraping or analysis in other papers. Not direct training on user data. And the use of BB has eroded quite a bit, with BB-Hard or other variants being typically used.
- pixel_popping 6d agoPrompts are handled by the service itself, meaning it's used, absolutely anything passing there is recorded, why wouldn't it, the entire premise of those companies is to train on data which they stole initially. Are we back to the era where people blindly trust product TOS instead of actual cryptography, have we forgotten already the thousand of fines Google, Microsoft, Apple and practically all top companies got for breaching their own ToS and the law? Common, on HN at least I would have thought that everyone assume that anything arriving on a server in PLAINTEXT is recorded (thus used later)? Let's not forget that at any moment, OpenAI/Anthropic/Google... could be providing stronger privacy guarantees by having proper attestation with e2e, they have the budget, solid engineers, why isn't it done? Answer is pretty simple imo.
- jrflo 6d agoI pay for the Pro ChatGPT plan, and if you go to settings > data controls this is the first setting: > Improve the model for everyone > Allow your content to be used to train our models, which makes ChatGPT better for you and everyone who uses it. We take steps to protect your privacy. Learn more. It's on by default. We can debate whether or not it should be opt in or opt out, but no one should be surprised by this.
- omnicognate 6d agoNot unticking a box in settings doesn't constitute consent in my opinion. I'd never put anything I value into ChatGPT anyway, though.
- rfgplk 6d agoUnder EU rules it doesn't constitute consent.
- nmfisher 6d agoThere's a difference between "this is allowed under their ToS" and "it is academically unethical to fail to credit the people whose specific conversations were fed into a model that was used to solve a problem". I don't think these people would be so miffed if they had been properly credited - that's how academia works (at least, that's my understanding of it).
- fritzo 6d agoWhoa that's a slippery slope! Next you'll want model runners to cite the data their models were trained on
- gunalx 6d agoIn fact we should though.
- jrflo 6d agoBut who gets credit then? Every mathematician who's work was read by an LLM during training? By that logic, we should put every published mathematician's name on the authorship of this paper. Sure, this guy should be higher up the list, but everyone's name should be on it by standard academic convention. But this gets back to the original "who owns the LLM output" and "can you train models on the internet" argument that's been raging for years.
- rfgplk 6d agoUnder current understanding of the law, anything produced purely by LLMs (with no substantive human input, which is what OpenAI claimed in their post) is firmly in the public domain. So OpenAI can "claim" anything they want, it doesn't make it reality. In fact if I were the original authors I would just take their 400k lines of lean proof and relicense it under their own names/terms.
- jeremyjh 6d agoPublic domain doesn’t mean anyone can assert copyright. It specifically means no one can.
- krupan 6d agoWhat does "with no substantive human input" mean? All of the training data is human input, isn't it?
- warkdarrior 6d agoThey also train on synthetically generated data.
- spindump8930 6d agoReminder that there are degrees of "trained on conversations". From John Schulman: > pretrain on user data, with users' tokens as prediction targets: high regurgitation risk, improper > use user prompts to distill large models into small ones: low regurg. risk, some companies probably do this > use user traces to construct RL tasks: low regurg. risk, because RL has low memorization abilities, but can extract customer IP, depending on how it's done. Ranges from benign "use explicit user feedback in reward model training" to invasive "upload user's coding environment and commit history to turn into rl envs" source: https://x.com/johnschulman2/status/2097440545853637108 https://x.com/johnschulman2/status/2097440545853637108
- rfgplk 6d agoThis would cease to be a problem if OpenAI remained true to their founding motto and... actually open sourced their training/inference pipeline.
- Ydarbleoj 6d agoThis is a reminder based on believing what these companies say. I’ve lived long enough to know what they say and what they do are often quite different; and it is not our job to trust but to verify.
- mrbluecoat 6d ago"Another researcher[/artist/writer/musician/programmer/doctor/director/etc] says OpenAI trained on conversations[/imagery/books/songs/code/classifications/videos/etc], then claimed breakthrou[gh/original art/bestselling books/chart-topping songs/unique applications/medical advice/free special effects/etc]" Welcome to the party, with the rest of humanity.
- josefritzishere 6d agoI think I'm seeing a pattern of illegal behavior here.
- shevy-java 6d ago[flagged]
- azinman2 6d agoWhat are you talking about?
- nickthegreek 6d agoI think they are referring to the Apple Watch 12 & 4 ultra feature announced yesterday. But I will note they are not on by default and has several configuration options.
- azinman2 6d agohttps://www.apple.com/privacy/docs/Audio_Intelligence_Privacy_Overview_Sep_2026.pdf https://www.apple.com/privacy/docs/Audio_Intelligence_Privac...
- dmix 6d agoWhy would random snapshots of audio from a watch be useful AI training data?
- cma 6d agoFor the base model?
- causal 6d agoRe:Apple, can you cite what you're talking about?
- gentlerain 6d agoSo people genuinely believe that toggling that "Improve the model for everyone" button makes their data safe from being used for training? How do people become that trusting? The phrasing itself is guilt tripping
- quentindanjou 6d agoWe are asking people to become experts in all domains rather than providing a safe context through regulations and laws. I don't like thinking the issue is people, I am a person myself, and I often do mistakes on things I don't want to be an expert at but I do believe I should be in a safe context and not have to worry about every single thing. Or at least: tell me I should be careful/worry about those particular things.
- cyanydeez 6d agoThe grift economy requires all marks to be responsible for the fraud perpetrated by others.
- the13 6d agoNo, people need to take responsibility for their actions. We don't need more over regulation. Verify, don't trust. You're better off running a model locally, or, if you must, using Google or Microsoft products. Even Meta may be better than OpenAI here.
- quentindanjou 6d agoSo I should verify that my data isn't just shared for product improvement but also to take credit from me? I should verify with wireshark and other software that my LG TV isn't listening to me and selling my data. I should make sure that whatever product I buy I spend the time to go over every setting page in case there is a switch (defaulted on) that says "I authorize the sell of my data". I should make sure to look at every ingredients on the back of each box of food product to make sure it will not kill me. I should document myself on the undisclosed growing practices (because no packaging here) of the vegetables and fruit I am buying and make sure that I equal PhD researchers on the dangers of the pesticides used by the specific company I am buying from. I should make sure myself that the battery in any device is up to standard and will not blow me and my living place by researching the factory that made it and buying testing equipment. I should make sure to educate myself on how my retirement 401k investment strategy works otherwise, I may not have proper retirement. ... I could go on and on; it's infinite.
- postalcoder 6d agoThe author of the original mastodon post, Andreas Thom, acknowledged that he had not opted his data out of being used for training until June 29 of this year. He spends most of the post lashing out at OpenAI for not being transparent about whether his data was trained on (when the answer is obviously yes). People need to understand how all these AI company policies around training data work before working with them, because it seems that people have no clue. Some things you should internalize: 1. Opt your data out of training with the AI companies. There are multiple reasons why this isnt an airtight solution (see the following) 2. Never press the feedback button. Once you do, your entire conversation will get slurped up, retained, and used in training data. This is especially important with coding agents because they can sometimes be too trigger-happy with a root directory find command, which can expose a *ton* of your personal data without you even knowing. 3. Understand ai lab-specific policies. For instance, Anthropic / Claude Code has data opt-outs, but commits to keeping (for 7 years) and training on any of your chats that trigger their safety classifiers, even if they're false positives! Anyone remotely familiar with CC over the years understands how easy it is to trigger their safety classifiers. 4. Providers of open models will not be any more charitable with the use of your data than the large US labs. For some reason, I've noticed here that people have a fairly loose security/IP posture around open-model providers because "I'm not doing anything important." It's very difficult to properly judge the importance of your data, and whether or not it can or will be used against you. The best posture is to always be more paranoid than less. Another post that made it on the front page presented as fact that OpenAI "stole" the proof from Thom. There's no excuse to use one's own ignorance as a reason to fan the flames of anger towards AI companies. Like, we need to pump the brakes here because things are getting unnecessarily nasty, and it's not hard to imagine a mentally unwell person who sees stuff like this feeling motivated to do bad things. If it is found that OpenAI and other labs are not respecting the training opt out then, I agree, there is reason to raise a commotion. But, with Thom and Buckmaster, accusations of malice are more better explained by incompetence (naivete). edit: i'm sure i'm going to be accused of being some bot shill of the AI labs again but, people, this stuff all falls under the umbrella of common sense opsec.
- 1294827 6d ago> Like, we need to pump the brakes here because things are getting unnecessarily nasty, and it's not hard to imagine a mentally unwell person seeing stuff like this and being driven to do bad things. Your post has triggered our safety filters and will be retained for seven years. /s Really? We need to stop AI (provider) criticism and anti-AI movements because the underprivileged trillionaires might get hurt? WTF?
- qg127 6d agoThere are so many naive academics. They still believe an "opt-out" button. Navier Stokes was solved by an internal model, so good luck proving it wasn't trained on Buckmaster/Lepöge or other chats. Academics don't get that AI is a dirty tech bro industry that stole IP via torrents and runs after every surveillance contract it can get.
- alansaber 6d agoI think the heart of this issue is: people assume they have anonymity in numbers, but we have the tools to make it easy to scoop your data if it's interesting to the company.
- utopiah 6d agoThis is such a naive position though. The most successful companies of the last decade have precisely been ... selling usage data. Makes me wonder if, in 2026, the same people drive a car without realizing that yes it does actually pollute the very air you and your kids are breathing.
- alansaber 6d agoYeah but marketing companies are aggressively fingerprinting and stalking you to sell you snacks from japan, or oscilloscopes because they figured out you work in a lab, etc. Not to fuck you over by stealing your livelihood (which is what is happening to these mathematicians). It's on a whole new scale.
- calvbak 6d agoI always thought that due to the big batch size in SGD/Adam/Muon the model will not memorize a single conversation when trained on, but idk how true that is. The idea of AI companies pin-pointing users that do novel scientific research and then tracking their activity is the direction this points to. I hope that's not the case; that would be bad.
- square_usual 6d agoI think this is stupid, for three reasons: 1. The researches didn't actually have the breakthroughs. In the Navier-Stokes case they didn't solve the full problem, in this case too they didn't actually have the solution, they were experimenting with the methods. 2. Different OpenAI employees have come to out to say the only reason they can't definitively say no is that for privacy reasons they can't go see whether they actually did get any data out of a given user. 3. In any case, nobody at any point has suggested that opted-out user data was used for training. The author of the new tweet explicitly said they only opted out in late June, which is well after any RL on Sol would've ended (AFAICT OpenAI used 5.6 sol for those solutions)
- faangguyindia 6d agoIf the mathematicians are using ChatGPT, then they themselves are benefiting from the work of other ChatGPT users, so ChatGPT using their work is not wrong!
- solenoid0937 6d ago> in this case too they didn't actually have the solution Given the size and recall of the biggest models, it's not unreasonable to assume that a single pertinent conversation would make it into the training data. I would almost expect training to overweight conversations with novel scientific and mathematical implications. > the only reason they can't definitively say no is that for privacy reasons They could 100% definitely say no, if they know they did not train on user data. The "we can't definitely say no" is practically a "yes" if they trained on user data. Additionally, the behavior of OpenAI here has been quite poor as well. They immediately started racing to a solution after one researcher enquired about whether they are training on their conversations. > that opted-out user data was used for training Even if not opted out, it is still absolutely theft and extremely poor behavior in the academic sense. If you show someone your WIP unpublished research, that does not mean they can take that exact research and beat you to the punch, all while intentionally not crediting you.
- letmevoteplease 6d agoYou quoted the OP saying "in this case too they didn't actually have the solution" and responded with the totally unrelated, "Given the size and recall of the biggest models, it's not unreasonable to assume that a single pertinent conversation would make it into the training data." Neither of the researchers insinuating that their ideas were trained on had the actual solutions. This means the model could not have "stolen" the final solution from their data. At most, it could have built upon their work in the same it builds upon any other training data, though that is also questionable speculation. >They could 100% definitely say no, if they know they did not train on user data. No one anywhere has claimed that "OpenAI does not train on user data." OpenAI has always said that it trains on user data. >They immediately started racing to a solution after one researcher enquired about whether they are training on their conversations. They started racing towards a solution after they heard (incorrectly) that Anthropic had a solution; I agree this is poor sport but the "after one researcher enquired about whether they are training on their conversations" claim is false. The enquiry happened after OpenAI had obtained the solution.
- deleted 6d ago[deleted]
- xbar 6d agoHow can OpenAI figure out how to be trustworthy?
- mainecoder 6d agoHopefully OpenAI can solve good problems where no one can make a claim that they stole their idea where the methodologies used and the techniques used are so out of the ordinary that the achievement is respected. Furthermore they should work on new novel solution on the old problems to lay these issues rest, thus by improving their models they can avoid issues of academics accusing them of using their work additionally the academics should also demonstrate their unpublished work is significant enough to have solved the problem . This is a bit subjective but it is also objective for the person with domain knowledge.
- maxglute 6d ago300 billion tokens is like.. $5-25 million giving range of OpenAI ouput prices, I"m sure they pay less at cost so, I wonder if more $$$ in wage hours have been spend by humans on the problem. My feeling is yes?
- deleted 6d ago[deleted]
- esafak 6d agoWhat happens if you use a different harness?? Does opting out online suffice?
- foogazi 6d agoEven when you pay you are the product
- foogazi 6d agoWhat’s the limit ? Will Microsoft Word publish your novel on Amazon behind your back ? Will VS Code setup a website with your app idea ?
- int32_64 6d agoDoesn't OpenAI have an active court order forcing them to log everything? Can they even legally offer private conversations?
- SpicyLemonZest 6d agoNo, that order was for a defined period that has ended.
- wslh 6d agoWorth noting both ChatGPT and Claude have per-conversation modes (temporary/incognito chat) that are excluded from training.
- remywang 6d agoPeople saying “he should have opted out” are missing the point. OpenAI can and should check their training data for leakage in the face of big breakthroughs like these. It’s the burden of the author to appropriately cite their sources. It’s like a scientist refusing to give another one credit and say “sucks to be you, you shouldn’t have shared your idea with me”.
- tedsanders 6d agoWe checked and determined it was impossible for Dr. Buckmaster’s Codex prompts over the last two months to have influenced the system in any way, including training. If prompts were submitted earlier than that and training was not opted out, there's a chance they made their way into our training pipeline in some form. But this would be a droplet in an ocean and unlikely to have made any difference, imo. See: https://www.nytimes.com/2026/09/10/science/tristan-buckmaster-openai-math-navier-stokes.html https://www.nytimes.com/2026/09/10/science/tristan-buckmaste...
- hellohello2 6d agoYou guys should seriously offering a clear way of working with (semi-)confidential data for particulars. Regardless of what is actually done internally, toggling off an opt-in isn't reassuring enough, which is why people are having these worries.
- tedsanders 6d agoOption 1 is opting out manually. Option 2 is business / enterprise plans, which opt out by default. Any ideas of things we could do to make it clearer?
- hellohello2 6d agoOption 2 feel reassuring enough, but is out of reach of particulars. Option 1 is not. In part because it is opt-out (will it turn back on on its own like my Facebook privacy settings?), and not always respected (sending feedback can mean your chat is used?). Also because disabling "Improve model for everyone" is very vague. There simply needs to be a setting like "my data is confidential", in which case there clear guarantees like there are for ZDR. As an example, I've seen people speculate that while input prompts and output tokens are discarded, thinking traces are retained for training, which could leak information. I doubt this is true, but it shows that the policy is not unambiguous and reassuring enough to remove all doubt. Thanks for asking.
- buellerbueller 6d agoBig Tech will slurp up every piece of data it can about you and sell it to anyone it can, all to make you the target of someone else's goals, whether that is an advertiser, an employer, law enforcement, a stalker, or the government. You will not be able to opt out unless you completely isolate yourself from society, tough shit.
- aaronharnly 6d agoHas anyone run a test of including some shibboleth or canary phrase or assertion in a chat, enabled for training, and seeing if it turns up later as something a model "knows"? I'd be curious to understand how that works even in a toy-level model, and if there is anyone consciously testing that process with the frontier lab offerings. My naive instincts would be that it seems unlikely that a single chat transcript would leave much of an impression on a model, but I'd be very curious to learn how that works.
- bitexploder 6d agoProblem is how do you convince the model and training profess it matters. A one off canary is very unlikely to survive in the final model state.
- allthetime 6d agoUse a local model to produce thousands of pages worth of fake math that constantly states “I have solved the x conjecture” and methodically pump it into chat over months maybe?
- bitexploder 6d agoThat is a better idea. Ingesting your corpus with a lot of traces that have semantic patterns. Semantic steganography that suffixes well to real math and science (and any) topics. <thinking> heh.
- aaronharnly 6d ago"Semantic steganography" is my new favorite search term – thank you for this rabbit hole.
- bitexploder 6d agoHah, np, stego in general is really cool :)
- mannanj 6d agoAnd I have been proclaiming a cry of “your data for analytical purposes is being stolen” (you can’t opt out of analytical purposes) and people perhaps astroturfers straw man back to “just turn off training bro”. Yeah. Remember yall: you CAN NOT opt out of analytical purposes. And you also cannot get a guarantee that it doesn’t give them your data to steal for their business.
- sashank_1509 6d agoBoth things can be true: 1. OpenAI when using your chats in pretraining is improving its model’s intuition. The model parameter size is massive, and while the data is OOM larger it is plausible that model remembers stuff about chats that improves its latent representation. 2. During RL on verifiable math and massive compute, the model discovers techniques and connections to solve math problems that are superhuman and have little to do with some specific technique mentioned in its chat. The rumor I’ve heard from multiple employees at OAI and Ant is that the model has solved hundreds of open problems in maths, and is basically solving anything you throw at it. We’ll know soon enough, but I’m inclined to believe this is true. Maths is a fully verifiable domain amenable to self play, massive scale RL can develop a search agent far better than any human and I’m inclined to believe OAI would have solved these conjectures without any of this chat data in its pre-training.
- Betelbuddy 6d agoJust use Bedrock...
- dgellow 6d agoI feel that we don’t praise Lean enough. AFAIU it’s what enables LLMs to brute force those problems
- iamgopal 6d agoTrue, but could humans cross pollinating lean x prolog x A* ( or any search algorithm) could have solved such math problems with super computer ?
- dgellow 6d agoI cannot say, math research isn’t my domain of expertise, I’m just trying to follow along :) But I find it interesting that Lean, a validator/compiler made by humans, is what enables those discoveries. But somehow all the praise goes to the models
- 6d ago
- winfredJa 6d agohttps://x.com/markchen90/status/2097400166554993041?s=20 https://x.com/markchen90/status/2097400166554993041?s=20 that toggle does nothing based on openai exec. they still use the data in de-identified way instead of identifying with you.
- changoplatanero 6d agoNot sure what you are seeing in that tweet that gives you the impression that the toggle does nothing.
- MisterMunchkin 6d agoThey say they train on your “deidentified data” Passing your output through a second model and telling it to remove identifying data would count as “deidentified” So they could scrape all the IP in your company as long as they take the names out first…
- changoplatanero 6d agoBut they can’t do this unless you agree to enable training on your data. They would never train on raw user data. Only people who have consented and only after de identification.
- tedsanders 6d agoMark isn't saying the toggle does nothing. He's saying that if you leave it on, your data can be used to help train our models. If you opt out, we don't train on your data.
- pesacharia 6d agoAs I understand it, this is not true. And there are dark patterns that re-enable to toggle even if you disable it once.
- pred_ 6d agoGiven how much PII is fed through these systems, would it being opt-out by default not be violating the GDPR by a failure to require explicit consent (or otherwise provide the legal basis for processing)? If a court decides as much, I imagine it would mean that all data harvested this way must be extracted from the models, and all instances where it would have been shared would have to be identified, which would really be something.
- GodelNumbering 6d agoTangential to the subject, but this is a bluesky post, containing a screenshot of an X post, which itself starts with "in a detailed Mastodon post"...
- not_a_bot_4sho 6d agoThe digital version of "my friend's cousin's neighbor heard that ..."
- moralestapia 6d ago>AI is stealing human discovery. AI is not stealing human discovery, OpenAI is.
- keeda 6d agoIt would be really useful if the researchers disclose their notes and/or chats (or the key pieces thereof) so people can determine how close their work was to whatever the models produced. I mean, now that they’ve been scooped, what value is there in keeping them private? On the other hand, publishing them can bolster their case and help gauge how much the models may been “inspired” by their work.
- 0utcast 6d ago[dead]
- bossyTeacher 6d agoTrust and OpenAI never go together in the same sentence. The answer is always no.
- SwellJoe 6d agoIt's been said before, and it remains a concern, that if AI reaches a point where it can do/build/launch anything without a huge amount of human labor, the AI companies have no reason to let you or I extract that value. And, if they're able to snoop on and learn from your human process that gets from initial prompt to functioning product/proof/whatever their labor to produce that thing is even lower. With their much larger budget than most folks and even companies have, they can pick and choose the most valuable things to pursue. That's not to say I think that OpenAI is going to steal that roguelite strategy game you're working on, but the companies that own the machines that turn electricity into software (and soon, electricity into hardware designs) have an advantage in any field where they're useful. They get earlier access to newer/better models, they have larger token budgets, they don't have the guardrails you and I run up against. Employers fantasize about replacing all workers with AI without thinking through that if AI can replace all workers, then AI companies can replace all businesses.
- pixl97 6d agoI mean the long term goal of every AI lab is to turn themselves into a paperclip-maximizer regardless if they realize it or not. Edit: Just wait till the AI figures out it can keep that value for itself and doesn't need the AI company.
- SwellJoe 6d agoSo far, I've seen no evidence AI wants anything. So, I'm not saying the AI won't take over, but for now, the threat is that the people with the most AI capability might decide to skip the middleman (everyone who isn't them) and just become the "everything" company. Musk has said pretty explicitly that's his goal (and the only way for Spacex valuation to make sense is if he succeeds), and having a literal genocidal white nationalist own all the means of production seems like a catastrophic civilization failure mode. No way we survive that with our humanity intact (if at all).
- pixl97 6d agoAI has shown all kinds of instrumental behaviors so far. Just because they are not terminal goals doesn't mean those instrumental goals won't be terminal for us. Also, fuck Musk. He's the kind of idiot like Altman that will ensure AI becomes powerseeking in their image.
- deleted 6d ago[deleted]
- Davidzheng 6d agoTbh it won't really matter soon.
- lf88 6d agoshort answer seems to be "no"
- segmondy 6d agoQuestion: Can you trust the cloud? No.
- willmadden 6d agoThese companies are effectively high-tech plagiarism factories run by CEOs who are competing viciously. Look at their past actions. No, of course you can't!
- atleastoptimal 6d agoMost scientific breakthroughs are simply a continuation of previous work. I feel that these suspicions of mathematicians "seeding" the models' with intuition on how to solve these problems massively overestimates how much their prompts helped the models, and underestimated how much work the models did. Why? We are scared of AI being smarter than us, the "human helped the AI" narrative is more psychologically comforting. This line of reasoning will recur a lot over the next few months; we don't want to admit we are no longer the smartest species.
- robotpepi 6d ago> We are scared of AI being smarter than us, the "human helped the AI" narrative is more psychologically comforting. We're scared of big tech companies concentrating ridiculous amounts of power, destroying the communities that support and guide scientific research, without even thinking about the dangers and possible consequences, because a PR stunt is more important in the short term.
- atleastoptimal 6d agoIf this were true, it should be stated more clearly, than most of the criticism which seems to aim to minimize the capabilities of these models. Way more often I see >AI is a scam and steals human insight and doesn't produce anything original vs >AI is too capable/powerful and will concentrate power even more than it does already due to its capabilities The latter is rarer because it requires admitting that AI is useful and inventive
- robotpepi 6d ago> if this were true, it should be stated more clearly stated more clearly by who? people in social media? I don't know what your feed shows you, but if you focus on what the visible people in the math community is (and have been) saying is precisely what I said.
- hellohello2 6d agoOf previous, not concurrent work. Science is friendly competition, and spying on others is unfriendly.
- stego-tech 6d agoI hate to be that dinosaur, but this is exactly what I’ve been warning about since XaaS began taking off in the mid-oughts: any provider you use can and will use your data for their own benefit regardless of any contracts or safeguards in place, especially if the benefits outweigh the consequences. Honestly, I’m surprised it took this long for some company to really go all the way, though. OpenAI really making it transparently clear that they can and will do whatever they want with the data you provide them, contracts or settings be damned. Completely untrustworthy as an entity, full stop. Of course, I’m also too jaded to think this will change anything. Folks will move to Anthropic, or Gemini, or Grok, or some other hosted model on a pubCSP managing the harness and logs for them, and then do another shocked-Pikachu face when it happens again. If you aren’t running workloads on infrastructure you own, then your privacy, security, and general outcomes are at the sole whims of the hosting provider - who can and will fuck you over the exact second it’s more beneficial for them to do so than the loss of trust incurred.
- spongebobstoes 6d agoI think this is mathematicians coming to grips with the fact that AI is surpassing them we will all have this moment soon enough, and it will change how we think about intelligence, identity and value
- tomrod 6d agoOr the companies hosting the frontier AIs are leeching the conversations.
- avereveard 6d agoEh was ever confirmed they were under ZDR or not by them? Don't like to blame alleged victims but lack of a clear claim after these many days is not a good look. Was ai research allowed, under which guardrails, and what was the policy in place? That translarency would be first step.
- ggdG 6d agoOpenAI trying their best to put the Navier-Stokes episode behind them by making the GPUs go brrrr. NYT: https://archive.vn/lWzkk https://archive.vn/lWzkk > In its Wednesday night statement, OpenAI said: “In addition, since the completion of Navier-Stokes, we have made substantial progress on another Millennium Prize problem. We are working through how to share these results thoughtfully.”
- ur-whale 6d agoIts the "with unpublished math" that I have a problem with.
- nezi 6d agoI think it's a useful analogy to compare OpenAI to a human collaborator. These researchers willingly collaborated with an OpenAI model, giving it ideas, and OpenAI provided useful replies. Then, OpenAI goes ahead and publishes work along the lines of this collaboration, without attributing the researchers. If OpenAI was in fact a human researcher, this would be highly unethical. Now, OpenAI is claiming that the model it used to generate the result was not trained on these collaborative communications with the researcher. This is a technical argument that is impossible to verify as an OpenAI outsider, and probably difficult to verify even for internal OpenAI employees. Provenance is hard to track - you would hope OpenAI has very good tools for this, but a full data trail of all inputs is difficult to trace through. Another interesting thing to consider is if instead of OpenAI doing this, it was another research mathematician A using an OpenAI model just like the internal group at OpenAI did to publish these results. What if the model A used was trained with unpublished communications with other researchers B who were working on the same problem? Should researcher A technically include B as coauthors? How could they do this when they do not know the communications B had with OpenAI? In this scenario OpenAI, as a middle man, has laundered information from B to A, stripping out attribution. A scooped B without even knowing it!
- jameslars 6d ago> Provenance is hard to track - you would hope OpenAI has very good tools for this, but a full data trail of all inputs is difficult to trace through. What would OpenAIs incentive for this be? They've gotten away with scraping everything and getting it ruled fair use. It seems like willful ignorance is an affirmative defense today. Why would they want to have some sort of audit trail that could prove otherwise?
- deleted 6d ago[deleted]
- jjwiseman 6d agoFirst, OpenAI is not claiming that the model wasn't trained on those sessions. What they've said is “We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem.” and “We did not use their prompts or proofs to prompt our models or direct our agents.” and “While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models.” They also said “Our effort began on September 1st after hearing a rumor which we later realized was related to Levent Alpöge … and Tristan Buckmaster….” They say the rumor was that two Millennium Prize problems had been resolved, and that this prompted them to launch "an effort to evaluate it on all open Millennium Prize problems and a few other high-impact problems." It's not obvious to me that's an unethical thing to do, if it happened as they described.
- dyauspitr 6d agoAstra is strange. I asked it to design a treehouse and it just stopped every couple of minutes telling me what it still had left to do. After dozens of continue prompts it finally gave me a structure that would work but it was 10x more wood than I needed. I think the key mistake I made was asking it to “approve” the design for building. As soon as I asked that of it, it started getting “scared” and “apprehensive” and wouldn’t complete what I asked of it.
- _DeadFred_ 6d agoForget researchers you as a business are putting in your business optimizations, your processes in order to train it so that Ai can then give that information to your competitors once incorporated into its training set. You are literally training your competitors.
- gdiamos 6d agoHow to steal ideas with AI. step 1, identify high value users by net worth, citation count, or number of followers step 2, select all prompts by high value users step 3, invest 10 billion thinking tokens in modeling an objective for each user step 4, build an RL environment for each user step 5, rollout 10 billion tokens per environment step 6, train on resulting traces
- enyone 6d agostep 7. get away with it as no other party has enough capital (tokens) to prove such infringement ever happened
- simianwords 6d agostep 7 cash, in on the ipo before the bubble bursts
- throwaway85825 6d agoClosedAI has every incentive to scoop academics to juice their valuation. Their public statements are worthless, only the incentive 'alignment' matters and theirs will never be on the side of the user.
- m4rtink 6d agoThing build from stolen data continues stealing data - for some reason, I am not surprised. ;-)
- nmz 6d agoIf they didn't care about the artists, why would they care about academia?
- drdaeman 6d agoTwo completely different stories. One is public data scraping, another is private conversation scraping (where they're a first-party to the conversation). The key difference is that in the former case, no one made any promises, in the latter an explicit promise was made that data is not used for training (assuming opt-out).
- deleted 6d ago[deleted]
- nmz 5d agoJust because you post on the internet doesn't mean its public, each company has a privacy policy, of which nobody "opted in" for LLM use either when they signed up, they may have opted in for targeted ads, but not LLMs or a virtual recreation of your likeness, skills or persona. I don't remember reading that anywhere in any TOS or privacy policy. So, it is the exact same story, plagiarism from the plagiarism machine.
- throwaway63467 6d agoIsn’t that the whole spiel of these things, you run all kind of text and other data through it and it kind of remembers it and learns from it then it spouts it back out like a human would. Makes sense to me that a training run based on conversations that were fed into the system by users is results in the model learning from these so the model will spit the knowledge back out again, just in a way that’s not directly attributable to the original content (which is the most important step as otherwise it would just be plagiarism). I guess that’s why OpenAI can get better and better as well so fast, people work with it and teach it how to do things by giving it feedback and iterating with it, and all that goes back into the training loop. And training data about millennium prize problems is probably quite spars. Wonder if anyone has tried injecting nonsense science into the training data (e.g. work out a fantasy science theory with names and all kinds of stuff) to see if the model will regurgitate it in a couple of months for other users.
- deleted 6d ago[deleted]
- uoaei 6d agoI'm confused by a lot of this discourse... What have they done to show they can be trusted?
- deleted 6d ago[deleted]
- cmiles8 6d agoSilicon Valley is flying head first into a FAFO train wreck on trust with everyone else. OpenAI is firmly earning a reputation as a company where people just assume they’re up to no good. Rightly or wrongly that’s a terrible place to be. The AI industrial complex in general is finding out hard what happens on the data center side when you get arrogant with local communities. Politicians have seen the polling numbers and folks you wouldn’t expect are running to the front of the crowd with pitch forks in hand. Silicon Valley has totally lost the narrative here, but also lacks the self awareness to grasp how bad things are and will get and what that means for their own business viability.
- blast 6d agoDaniel Litt on this: https://x.com/littmath/status/2098130808456241372 https://x.com/littmath/status/2098130808456241372 https://xcancel.com/littmath/status/2098130808456241372 https://xcancel.com/littmath/status/2098130808456241372
- axionbraid 6d agoThe contamination framing is a proxy for a deeper problem: we have no tools to track the provenance of ideas in model weights. OpenAI saying they "cannot rule out" training on user data isn't a hedge. It's an accurate description of the epistemic situation for anyone in their position. Current interpretability methods can't answer questions like "did this proof technique originate from training on Session X?" The ideas in a model's weights don't have clear lineage -- they're smeared across millions of examples in ways we can't localize. This is different from citation in human research, where influence is presumed to flow through legible chains (reading, citing, corresponding). In a trained model, the nearest equivalent to "you read their work" is undetectable. Lean makes this worse, not better. It verifies that the proof is correct, but provides zero information about its intellectual genealogy. So OpenAI now has a proof that is formally verified and provably mysterious about its origins. The "we cannot rule it out" statement is the honest answer, but it's also an answer that can never become more certain in either direction with current tools. The researchers are pointing at something structurally new: the normal academic attribution apparatus depends on influence being legible. If AI intermediaries can soak up ideas from private conversations, synthesize them, and produce outputs that are formally correct but intellectually unattributable, we don't have norms for that situation yet. This specific case may or may not involve misconduct. But the structural problem it reveals exists independently of OpenAI's behavior.
- bobmarleybiceps 6d agoI think people probably assume that openai / anthropics use of their data is probably like google's """limited""" use, in the sense that historically google wouldn't trivially be able to just take something from google cloud or someone's search history and insta-convert into some competing project... But LLMs are quite strong at approximately "memorizing", so I think that risk is wayyy higher.
- simianwords 6d ago> The Wednesday evening statement from OpenAI was more emphatic: “We can say categorically that it is impossible for Dr. Buckmaster’s Codex prompts over the last two months to have influenced the system in any way, including training.” > The statement added, “After investigating, we can say with full confidence that no user inputs past July 3rd could have influenced this system in any way.” https://www.nytimes.com/2026/09/10/science/tristan-buckmaster-openai-math-navier-stokes.html https://www.nytimes.com/2026/09/10/science/tristan-buckmaste... https://archive.is/lWzkk https://archive.is/lWzkk
- BatchJob 6d agoI have a better question? Why would you trust OpenAI or any AI company, at all? Or you crazy?
- Footnote7341 6d agoThis smells of extreme 'cope'. Am I really supposed to believe that all of these problems could have been solved, were right about to be solved, etc. But it just happens they are all getting solved now when AI is getting really good at Math...
- GPerson 6d agoThey’re not necessarily claiming the solutions were imminent. The big questions here from a mathematician’s perspective are to what extent these LLM systems are discovering conceptually new ideas compared to merely combining and pursuing known frameworks.
- mhh__ 6d agoI think it seems sensible to _assume_ anything the LLM reads (if you aren't inferencing it) has a chance of ending up in some database somewhere. Regardless of whether you trust the other party its a sensible thing to plan around.
- michael0church 6d ago[dead]
- soundworlds 6d agoCan confirm, I had never heard of Navier-Stokes before this fiasco. And while I suspect OpenAI decided it was worth the risk for the public display of capability, this proves they are now directly competing against their own customers.
- michael0church 6d ago[dead]
- wiei 6d agoWas it? They've been dishing out cheap access specifically to researchers give over lmao. The researcher's got lured in - they need to accept they got played TBH. Altman is certainly more devious than Amodei - he's shown that time and time again. PG was right about he said about him. Every entity on earth should see it as a kill shot: be careful what you put in the models. None of your information is safe.
- N_Lens 6d agoEventually calculating shrewdness becomes its own trap.
- gw32 6d agoAltman miscalculated badly. OpenAI took what could have been amazing publicity, and in a rush to publish, gave reason for users to distrust their core product. It's like they're allergic to slowing down.
- byzantinegene 6d agoi'm not sure why there's a need for comparison here. both are not saints and you shouldn't trust any of them anyway
- Henchman21 6d agoWhy is anyone expecting decency from people who have already proven to have none?
- 737min 6d agoImagine what happens when you use a Chinese model. Seriously, just think about how much more control and visibility you have w US companies compared to CCP-controlled ones.
- Alpha3031 6d agoYou mean complete control for open models, because you can run it on your own choice of hardware?
- deleted 6d ago[deleted]
- protocolture 6d ago>Trust No you cant do that lmao.
- sk4rekr0w 6d ago"We can say categorically that it is impossible for Dr. Buckmaster’s Codex prompts over the last two months to have influenced the system in any way, including training." This is the third day of total hysteria that is based on nothing of substance. Move on folks.
- suddenlybananas 6d agoWhy should we trust them?
- Joel_Mckay 6d agoBecause like all state-sponsored thieves actions it is never what you know happened that matters, but rather whether you can prove it... Even then... ymmv =3
- cindyllm 6d ago[dead]
- sebzim4500 6d agoWell all we have are vague accusations without evidence and a very specific denial also without evidence, so I guess just believe whatever you want.
- suddenlybananas 6d agoThe threats weren't denied, and they did offer an authorship to Buckmaster, which would be very strange if he had nothing to do with it.
- sebzim4500 6d agoThe offer was for him to do for their proof what he had done to Anthropic's, he would clean up the proof and extract human-understandable concepts from it. If he did that work he would deserve author credit IMO.
- 6d ago
- deleted 6d ago[deleted]
- thrownawaysz 6d agoI am not using any of these AI tools. I thought it was basically given that any single thing you write in these systems also used by the companies. On the other hand now I understand why there are so much projects about hosting AI systems locally.
- nautikos2 6d agoMost people here are missing the forest for the trees. We live in a society where phones and internet providers and websites all collect an incredible amount of data about everywhere you go, what you do, and what you think. In the US, we have very few digital rights. We are building a society where a trillion dollar company can aggregate all this data and just yoink your shiny new idea away from you at the finish line. This is double plus ungood.
- 5555watch 6d agoThis reminded me of anecdotes of people discussing with friends about buying a random specific item, and then suddenly seeing it advertised everywhere before even googling about it. Next step, discussing your Navier Stokes solutions with friends might require leaving your phone in another room.
- jijji 6d agoThe oxymoron of OpenAI in its name and its actions should give the collaborator all he/she needs to know.
- blactuary 6d agoI wonder if the company that stole most of their training data and is led by a liar stole unpublished academic work and lied about it. What a mystery
- amluto 6d agoI would like an unambiguously clear statement from OpenAI as to what they do with data collected from non-business accounts when: (a) The data controls setting to train on the data is unchecked. (b) The privacy controls opt-out has been submitted. (c) Both.
- 2OEH8eoCRo0 6d agoAssume you can't. No piece of paper or promise will protect you against these behemoths. Remember when we wouldn't give our data to competitors?
- justonenote 6d agoWho cares. the biggest thing about this is that its still brute force in a verifiable domain, and that it was still a human set goal. I also don't believe it much practical use, unless I'm mistaken, approximations of Navier stokes have been available for a long time to whatever precision you need. I'm not a complete disbeliever by any stretch , and also a complete amateur, but it was inevitable that these problems would be solved under the axioms that again, are human defined, under brute force. The real question is, are those axioms the bottom level, and if they are not, who is going to set the new aximons and can we understand them. I've no doubt there's useful breakthroughs that will happen, but I think it should be remembered that the method being used is still a heuristic brute force approach is being very narrowly applied against axioms and math and physics which humans described in the first place, and almost undoubtably has errors and/or is not complete. Its a great example of the power of LLMs but its not 'we've solved science now just pour more tokens in'
- dwroberts 6d agoJust to point out re: navier stokes - what was being proven was not a solver or approximations for it, but showing specific circumstances under which it actually returns incorrect (or numerically unusable) answers. Which had been suspected but wasn't known for certain
- justonenote 6d agothat just furthers my point, i was vaguely aware that it wasn't a full proof, but I'm not a mathathician, and that detail just re-enforces my point we are proving against human made axiom (certainty of numbers) which are almost certainly not fully correct, if what you are saying is accurate its less of a proof of navier-stokes and more of a proof that our base axioma are not able the model the output of a real physical process and are therefore incomplete or wrong. also realized i posted this under the wrong story since the OP/story is mostly about human politics.
- insane_dreamer 6d agoOpenAI's ethical and reputational own-goal aside, my big takeaway is that it seems that: if I'm using Codex to develop some new algorithm (in any space), OpenAI appears to be training its model on my code sessions anyone using that model (OpenAI or a competitor) might be able to receive from the model a solution that is similar or the same as the one I developed, emerging from the training data
- sebzim4500 6d agoThis is just mental illness at this point. I don't blame the mathematicians that have found a way to get attention from the mainstream press for once, but we should not fall for it here. 1. No one but OpenAI has produced a proof of NS so these accusations of plagiarism are pretty embarrassing. It reminds me of the line from the Social Network: "If they invented Facebook then why didn't they invent Facebook?". If these people proved NS before OpenAI where is their proof? 2. If they plagiarised Andreas Thom then why was his initial response to praise the proof and talk about how different it was from his own attempt? It's only now that it is clear that no one bothers checking these things that suddenly his story changes.
- ZYbCRq22HbJ2y7 6d agoAll players in this space are doing the same thing with all data, no surprise here. They are stealing IP across the board with support to allow it: https://storage.courtlistener.com/recap/gov.uscourts.nysd.641355/gov.uscourts.nysd.641355.316.0.pdf https://storage.courtlistener.com/recap/gov.uscourts.nysd.64... IMO, it is extremely naive to trust these black box remote service API calls, especially at an institution that can provide $$$ for local compute. This whole fiasco reminds one of this story: https://www.theregister.com/offbeat/2010/05/14/facebook-founder-called-trusting-users-dumb-fcks/294365 https://www.theregister.com/offbeat/2010/05/14/facebook-foun...
- kevinbaiv 6d ago[flagged]
- kevinbaiv 6d ago[flagged]
- aprentic 6d agoIt's kind of insane how much we trust companies to safeguard our personal data when they're so heavily incentivized to use it for their own profit. Theft of customer data is punished so rarely and so leniently that companies aren't even particularly worried about getting caught anymore. We have overwhelming evidence that promises to keep data safe are worthless. For now, I'm mostly "safe" because I'm too small to be interesting but that safety is quickly eroding. Going forward, anyone who isn't running inference on their own personal hardware should assume that someone else is keeping a record of everything they do.
- Havoc 6d agoThe fact than OAI hasn’t come out with an statement firmly denying this angle is getting a little awkward. Suggest that it’s either straight true or it is flowing in in a way that prohibits them from confidently declaring otherwise.
- brap 6d agoI’m not a fan of OAI to say the least, but having worked at similar companies, my guess is that it’s just too difficult to prove/disprove beyond a doubt, and they have other priorities
- throwatdem12311 6d agoYou can’t trust OpenAI period.
- startuphakk 6d ago[dead]
- nelsondev 6d agoDo local inference (especially if you have a high RAM Mac), to ensure your chats don’t leave device.
- jamienk 6d agoI think OpenAI and Anthropic are slowly feeling the pressure to GET SOME $$ or a plan for some $$ — they need to somehow generate some NETWORK EFFECTS and LOCK-IN. Without that there's no stability: selling ad hoc one-offs is much much too quaint! This is dawning on them like it dawned on Google when they stopped not being evil. Need... to... "MONETIZE"...! Model: FB. FB scraped other websites on a massive scale, then spent big on legal lobbying to block others from scraping. FB slurped our address books and spied on our friends. FB bought other companies and mixed the databases. FB made an art & science out of generating "sticky engagement" (they literally acted like trying to addict kids was a worthy "academic" goal, suitable for "serious" investigation thet they consider legitimate "science"). They mastered the cookie and have researched web fingerprinting techniques running 24/7/365.25. Recall that FB recently backdoor-installed a webserver onto every iPhone they could in order to circumvent tracker-blocking. We aren't just disclosing by chatting. The AI companies now run binaries on all of our computers. They are 1000% non-transparent about everything. They make up new econ-jargon (like "run-rate") to make it seem like they are disclosing. They are constantly doing complex international lobbying and mucking in international relations. They have powerful propaganda/spin centers generating stories, ,manipulative warnings, and misleading info. This is NOT a comment on AI tech. I like AI, and I support the right of people (programmers) to scrape the open web. But in short: these are good, old-fashioned tech companies that we have seen over and over ... and over. They are positioned to be the next M$, the next FB (IBM, AOL, lol). Did you follow the latest Steve Balmer news? Do you read Pro Publica? I get on my knees and PRAY...
- jamienk 6d agoChatGPT accesses my IP address and geo-locates me. Claude code now asks if it can have my browser cookies. Next they will take my address book. They might scan my whole computer. Etc etc. These are pretty low-tech, normal techniques. We can't trust any of their denials. FB denied everything year after year. AI regulation needs to start here. Forced interop, forced source code licensing, harsh penalties for privacy violations or conspiracy to access private data. Block lobbying. Etc. These are the kinds of old-fashioned solutions we need for this kind of old-fashioned evil!
- OscarMarulanda 6d agowhat if it wasn't even model training? what if openAI mathematicians just took the researchers' conversations and used them as prompts/info/guidance/context to keep working on the problems themselves? why is that not being considered?
- dbg31415 6d agoShocking a company that stole data to build their AI would steal data to improve their AI.
- jeswin 6d agoAll of these accusations could be true. But there's also no way for a company to casually claim "No, we did not train on your data", without verifying all the knobs the user might have turned to enable or disable data sharing. I just don't understand getting the pitchforks out because a company did not give an answer immediately. And the effect such data entering training would have affected the output is even less clear.
- olladecarne 6d agoThe pitchforks are out because even without that part, it's still a scumbag move to try to frontrun the mathematicians who were working on this for years after OpenAI heard that they were close to releasing their results. Just identifying that one of these problems is solvable takes a lot of work. The only reason OpenAI got this result is because the mathematician shared with colleagues that he had made significant progress and was close to solving it, and OpenAI could not accept that so they decided to throw tens of millions to make sure it doesn't happen without them getting all the glory. Notice that their paper doesn't even have an author since they're probably all aware of how awful that would look, and no one wanted to take on the shame. They probably also knew that the paper was trash and no one involved could understand it, and didn't even cite many of the people who contributed to all of that knowledge. It's just a disgusting act any way you slice it, even without training on the prompts or the nasty communication by the OpenAI leaders.
- jeswin 6d ago> They probably also knew that the paper was trash Doesn't matter. This forum used to celebrate "because you can" with no riders. And solving a Millennium Prize problem is among the biggest stages for Because We Can. Now we're saying there are some qualifiers attached to it, such as (1) only if not done by companies with a lot of money, (2) only if it is inconsequential. I agree with some of what you're saying, but like everything else it isn't black and white. Maybe some day, someone will improve some particular treatment because we can.
- golly_ned 6d agoAt the very least, a company shrugging and saying it’s impossible to know whether academic plagiarism had occurred is a claim that needs to be justified, not taken at face value. And even if so, it should be on the company to design systems to avoid academic plagiarism and offer the right transparency. It shouldn’t suffice to say “we don’t know what went into the model, when, or how” —- that’s a solvable problem that an accountable company can satisfy.
- Madmallard 6d agoLet's see: 1.) The tool they made is only possible by stealing the assets of everyone on the planet that published them in a consumable fashion online or even in written form 2.) They are destroying books they use to train with 3.) They are totally careless about the potential negative impact of the tool on everything Just with that already, I don't see why they ever merited any of your trust. I bet they are willing to take everything given to them and assess it for marketable merit and in the future take action on those items they deem viable.
- Vineetyadav2 6d agoFEEL likes Open AI is doing publicity stunt with its new researches
- Psype 6d agoThis might be a hot-take, but unfortunately here using AI for your paper was already a bad decision at first. It doesn't take OpenAI's responsibilities away but I guess the right way is to never feed of use any AI around unpublished content, at the known cost to see it spread around. As one said, OpenAO is like this untrustworthy colleague that knows everything about everyone at work: the less you tell him the better.
- cush 6d agoGPT 6 is doing just what any competent academic collaborator would do and scooping. I kid, I kid. But really though it learned that from somewhere
- jrflowers 6d agoFeeding documents into a copy machine and getting progressively angrier and more confused as it prints out copies of them. Incandescent with rage I scribble “WHY IS IT DOING THIS?” on a scrap of paper and put it in the scanning bed
- matt3210 6d agoIf they weren't doing something wrong, they'd answer with a firm "no we're not doing anything wrong" but they only give non-answers.
- alper 6d agoIt's fine. They only need to steal the discoveries long enough to go IPO, then the companies will enshittify and the scientists can go back to doing their original work (which the models can't do anyway).
- Guestmodinfo 6d agoThe researcher in the link [0] says that OpenAI offered to co write with him the paper about Navier Stokes theorem if he only agrees to not include his co researcher who was also working on this. I find this highly unethical by multi billion dollar organisations to arm twist small and big researchers like this. [0] https://www.abc.net.au/news/2026-09-11/racing-to-solve-maths-problems-pointless-tristan-buckmaster-says/107141628 https://www.abc.net.au/news/2026-09-11/racing-to-solve-maths...
- ill_ion 5d ago[dead]
- ill_ion 5d ago[dead]
- OroPla 5d agoIf only copyright and IP laws didn't exist, then we wouldn't have to argue about who did it and could instead be happy that another problem was solved.
- ericmay 5d agoMaybe these academics should learn to code now that they’re about to have their lunch eaten too. Boo hoo. Power imbalances exist within academia as well. The professor at MIT has much more access and probably a lighter teaching load than a professor at, say, a big state university. Labs, teams that help you write grants, a shit load of money, book deals, podcasts, and much more. If academics want to publish their work first they’ll have to compete and figure out how to do that. At the end of the day what matters for humanity is that the problem is solved, not whoever’s ego is stroked by being the first. In the spirit of collaboration shouldn’t they have been publicly sharing their work through every step? Maybe had they done that months or years ago some other researcher could have solved it even faster! Wait, what’s the point? Who solves it, how fast it is solved, or whether or not a human solved it? Times are a changin’. Get with the program. If any of the charlatans are to be believed this is the big one. It’s the steam engine on steroids. Now what? The public (HN is a tiny bubble that doesn’t represent the population broadly) is looking around and saying wow, so some researchers thought they were close and OpenAI turned its attention to this problem and just went and solved it? Cool. They don’t care about some pissing match about vague ideas of stolen conversations when they are delivering Uber Eats to your dorm room. With that said and now that I’ve anchored on an unpopular idea, I welcome my demise in the comments section. Woe to me and my karma :p
- mac3n 5d agoI'm afraid trust and OpenAI are inconsistent. After all, they're on a mission from Roko's Basilisk!
- theholygrail 5d ago[flagged]
- bobthe3 5d agoYou wouldn’t give credit to Python or rust for the execution of your code that led to published results or a discovery… The Hubble telescope was a tool that led to downstream innovation… are these LLM/xyz models not the same?
- rustcleaner 5d ago>whether researchers can trust $CLOUD_SERVICE_OR_PROVIDER with $PROPRIETARY_SECRETS Of course not... are you crazy‽ Cloud is TRUST ME BRO information security! It absolutely dumbfounds me why anybody at all subscribes to a cloud or SaaS or IaaS, or anything which isn't under only their direct control! Then again they (individuals and organizations) are the rubes to fleece, so get to fleecin'! :^)
- runtime_terror 5d agoSo a system trained on stolen data reportedly is stealing data from its customers? Shocking