26 ms·
OpenAI says it has evidence DeepSeek used its model to train competitor
- octacat 2y agofirst time?
- deleted 2y ago[deleted]
- m3kw9 2y agoSo if OpenAI didn't have these outputs for distillation, Deepseek wouldn't exist?
- beezlewax 2y agoThis is nothing short of hilarious.
- ryao 2y agoGiven that OpenAI model outputs are littering the internet, is it even possible to train a new model on public webpages without indirectly using OpenAI’s model to train it?
- wnevets 2y agoIts like a bank robber being upset when someone steals their loot
- udev 2y agohttps://archive.is/KiSYM https://archive.is/KiSYM
- cratermoon 2y agoIronic, OpenAI claiming someone else stole their work.
- vinni2 2y agoHow would they prove they used it’s model. I would be curious to know their methodology. Also what legal actions OpenAI can take? can DeepSeek be banned in US?
- iforgot22 2y agoThey might show DeepSeek's model calling itself ChatGPT, which users have already alleged. Same as how Cisco proved Huawei was stealing router code. Except in this case, nothing was stolen, unless they want to call ChatGPT's own training on source data theft too.
- freehorse 2y agoChatGPT outputs are all over the internet. It is harder to prove that deepseek used specifically o1 for training, instead of a lot of chatgpt output ending up in the training set from other sources.
- iforgot22 2y agoThat's a good point, at least for the prompts I saw. Like "do you have an app I can use" is commonly seen with "here's the ChatGPT app" online. And maybe they don't add anything telling Deepseek that it's Deepseek.
- paul_e_warner 2y agoIf you read the article (which I know no one does anymore) >OpenAI and its partner Microsoft investigated accounts believed to be DeepSeek’s last year that were using OpenAI’s application programming interface (API) and blocked their access on suspicion of distillation that violated the terms of service, another person with direct knowledge said. These investigations were first reported by Bloomberg.
- throwccp 2y ago[flagged]
- belter 2y agoThe subtitle is the gold... : "White House AI tsar David Sacks raises possibility of alleged intellectual property theft"
- conartist6 2y agololololololololol
- vrighter 2y agoSo what? They probably paid for api access just like everyone else. So it's a TOS violation at worst. Go ahead, open a civil suit in the US against an entity the US courts do not have jurisdiction over and quit whining...
- jhickok 2y ago>open a civil suit in the US against an entity the US courts do not have jurisdiction over Yeah, over a Chinese company no less.
- ForHackernews 2y agoWhat's good for the goose is good for the gander. Obviously a transformative work and not an intellectual property violation any more than OpenAI injesting every piece of media in existence.
- dagelf 2y agoInjesting is sure the right take. What a circus!
- thumbsup-_- 2y agois stealing from the thief actually a theft?
- amarcheschi 2y agoI quite like a scenery where llm output can't be copyrighted, so that it is possible to eventually train a llm with data from the previous one(s)
- layer8 2y agoOpenAI argues it’s a violation of their terms of service. So there are legal issues if it can be proven.
- mannewalis 2y agoBut OpenAI's model isn't open source, how would they distill knowledge without direct access to the model?
- layer8 2y agoYou don’t need direct access for LLM distillation, just regular API access.
- mannewalis 2y agook I looked it up and have a better understanding now.
- Palmik 2y agoLegal issues for who? Company A pays OpenAI for their API. They use the API to generate or augment a lot of data. They own the data. They post the data on the open Internet. Company B has the habit of scraping various pages on the Internet to train its large language models, which includes the data posted by Company A. [1] OpenAI is undoubtedly breaking many terms of service and licenses when it uses most of the open Internet to train its models. Not to mention potential copyright violations (which do not apply to AI outputs). [1]: This is not hypothetical BTW. In the early days of LLMs, lots of large labs accidentally and not so accidentally trained on the now famous ShareGPT dataset (outputs from ChatGPT shared on the ShareGPT website).
- deleted 2y ago[deleted]
- aDyslecticCrow 2y agoAnd they used all copyrighted data on the internet. If they wanna sue, they set a dangerous precedent.
- deleted 2y ago[deleted]
- top_sigrid 2y agohttps://archive.is/KiSYM https://archive.is/KiSYM
- deleted 2y ago[deleted]
- lawlessone 2y agoSo they're mad someone did exactly what they did?
- exe34 2y agono, no, it's completely different. "open"AI stole from poor people. DeepSeek stole from a $1T company. that's illegal!
- whatshisface 2y agoIt's reasonably likely that a lot of people linked to the federal government want to ban DeepSeek. You can tell it's being presented away from "they gave us a free set of weights" and towards "they destroyed $1T of shareholder value." (By revealing that Microsoft et al. paid way too much to OpenAI et al. for technology that was actually easy to reinvent.)
- nullbyte 2y agoI think the real concern from the govt's perspective is data privacy, since all the chat messages are stored on Chinese servers
- fullshark 2y agoWould it even matter? Isn't the cat out of the bag and everything they did repeatable by an American research team?
- Cumpiler69 2y agoIt matters because their goal was hyping up how advanced and difficult their tech is, propping up their valuations. DeepSeek proved the emperor had no clothes and wiped out a lot of their valuation when investors saw reaching parity to Chtgpt is not really that difficult.
- mastazi 2y agoI think parent was asking would it even matter if there was a ban. To which the answer would be "no" because as you said the point has been made. And, as parent pointed out, it's repeatable anyway.
- bhouston 2y agoIt doesn't matter from the US government perspective if all of the tech is replicated by US companies and US user continue to use US AI technology. But if US users start to use Chinese AI tech, then protectionism urges will appear that will likely figure out how to ban its use or subject it to large tariffs (e.g. TikTok, BYD, network equipment, solar panels, etc.)
- jasoneckert 2y agoWhat I find the most comical about this is that the whole situation could be loosely summarized as "OpenAI is losing its job to AI."
- mattgreenrocks 2y agoOpenAI should be excited that it has been freed of the tedious tasks of building AI and now they can focus on higher level and more creative things.
- Sateeshm 2y ago> focus on higher level and more creative things. But that's what OpenAI's costumers were supposed to do.
- jusonchan81 2y agoIt’s sarcasm.
- rooroobooragool 2y agoI think Sateeshm was also applying a generous layer of sarcasm.
- pphysch 2y agoOpenAI should be, but OpenAI died a while ago
- JoshTko 2y agoI wish I could upvote this twice
- munchler 2y agoSoon you’ll be freed of the tedious task of upvoting at all.
- 2y ago
- semking 2y agoThis is absolutely hilarious! :) ClosedAI scraped human content without asking and they explained why this was acceptable... but when the outputs of their training corpus is scraped, it is THEIR dataset and this is NOT acceptable! Oh, the irony! :D I shared a few screenshots of DeepSeek answering using ChatGPT's output in yesterday's article! https://semking.com/deepseek-china-ai-model-breakthrough-security-risk/ https://semking.com/deepseek-china-ai-model-breakthrough-sec...
- didip 2y agofr fr, ClosedAI is being a comedian right now. They scraped literally all the content of the internet without permissions. And I won't even be surprised if they scraped the output of other LLMs as well.
- amelius 2y agoIt looks like they want to spin this as "DeepSeek copied OpenAI". The general public/media might actually believe this is what happened.
- semking 2y ago[flagged]
- marricks 2y agoAlso, DeepSeek is allegedly... better? So saying they just copied ClosedAI isn't really sufficient of an answer. Seems to be just bluster because the US Govt would probably accept any excuse to ban it, see TikTok.
- breakitmakeit 2y agoAs the article points out, they are arguing in court against the new york times that publicly available data is fair game. The questions I am keenly waiting to observe the answer to (because surely Sam's words are lies): how hard is OpenAI willing to double down on their contradictory positions? What mental gymnastics will they use? What power will back them up, how, and how far will that go?
- snakeyjake 2y agoWhen large sums of money are involved the techbros will burn everything down, go scorched earth no matter what the consequences, to keep what they believe they're entitled to.
- ADeerAppeared 2y agoTheir way of squaring this circle has always been to whine about "AI safety". (the cultish doomsday shit, not actual harms from AI) Sam Altman will proclaim that he alone is qualified to build AI and that everyone else should be tied down by regulation. And it should always be said that this is, of course, utterly ridiculous. Sam Altman literally got fired over this, has an extensive reputation as a shitweasel, and OpenAI's constant flouting and breaking of rules and social norms indicates they CANNOT be trusted.
- bhouston 2y agoThe US government likely will favor a large strategic company like OpenAI instead of individual's copyrights, so while ironic, the US government definitely doesn't care. And the US government is also likely itching to reduce the power of Chinese AI companies that could out compete US rivals (similar to the treatment of BYD, TikTok, solar panel manufacturers, network equipment manufacturers, etc), so expect sweeping legislation that blocks access to all Chinese AI endeavours to both the US and then soon US allies/West (via US pressure.) The likely legislation will be on the surface justified both by security concerns and by intellectual property concerns, but ultimately it will be motivated by winning the economic competition between China and the US and it will attempt to tilt the balance via explicitly protectionist policies.
- derektank 2y ago>The US government likely will favor a large strategic company like OpenAI instead of individual's copyrights Even if we assume this is true, Disney and Netflix are both currently worth more than OpenAI and both rely on the strict enforcement of US copyright law. I do not think it is so obvious which powers that be have the better lobbying efforts and, currently, it's looking like this question will mostly be adjudicated by the courts, not Congress, anyways.
- bhouston 2y agoI don't think OpenAI stole from Disney or Netflix. Rather OpenAI stole from individual artists and YouTube and other social media who users do not really have any lobbying power. So I think OpenAI, Disney and Netflix win together. Big companies tend to win.
- mjburgess 2y ago> What are the first words of the disney movie, "Aladdin" ? The first words of Disney's Aladdin (1992) are spoken by the *Peddler*, the mysterious merchant at the beginning of the film. He says: "Ah, Salaam and good evening to you, worthy friend. Please, please, come closer..." He then continues with: "Too close! A little too close. There. Welcome to Agrabah. City of mystery, of enchantment, and the finest merchandise this side of the River Jordan, on sale today! Come on down!" This opening sets the stage for the story, introducing the magical and bustling world of Agrabah.
- ceejayoz 2y ago"You can't take data without asking" seems like a court precedent OpenAI really, really, really wants to avoid. And yet...
- amelius 2y agoWhy? When did large companies care about laws? See e.g. Uber, AirBnb. The only thing government cares about at this point is if information is shared with China.
- ceejayoz 2y agoThey care when they get big enough to attract attention from people like state AGs who can actually put the hurt on a bit. Uber and AirBnB both hit this point years ago; OpenAI's starting to hit it.
- galleywest200 2y agoAltman is part of that Stargate Trump group now. He and his ilk will just get pardons. Curious, though, can a corporation be pardoned?
- ceejayoz 2y agoThe President can only pardon Federal crimes. State-level crimes (like his NY felonies) and civil torts (like his case where he owes $500M currently) are separate.
- actionfromafar 2y agoYet. Give it some time.
- ceejayoz 2y agoSure, but in that scenario, it's a bit like the Last of Us characters being concerned about electrical meter readings. We'll have much bigger problems.
- osigurdson 2y agoI do think that distilling a model from another is much less impressive than distilling one from raw text. However, it is hard to say if it is really illegal or even immoral, perhaps just one step further in the evolution of the space.
- lemoncookiechip 2y agoIt's about as illegal as the billions, if not trillions of IPs that ClosedAI infringed to train their own data without consent. Not that they're alone, and I personally don't mind that AI companies do it, but it's still amusing when they get this annoyed at others doing the same thing to them.
- osigurdson 2y agoI think they had the advantage of being ahead of the law in this regard. To my knowledge, reading copywritten material isn't (or wasn't illegal) and remains a legal grey area. Distilling weights from prompts and responses is even more of a legal grey area. The legal system cannot respond quickly to such technological advancements so things necessarily remain a wild west until technology reaches the asymptotic portion of the curve. In my view the most interesting thing is, do we really need vast data centers and innumerable GPUs for AGI? In other words, if intelligence is ultimately a function of power input, what is the shape of the curve?
- ttesmer 2y ago> if intelligence is ultimately a function of power input, what is the shape of the curve? According to a quick google search, the human body consumes ~145W of power over 24h (eating 3000kcals/day). The brain needs ~20% of that so 29W/day. Much less than our current designs of software & (especially) hardware for AI.
- osigurdson 2y agoI think you mean the brain uses 29W (i.e. not 29W/day). Also, I suspect that burgers are a higher entropy energy source than electricity so perhaps it is even less than that.
- __MatrixMan__ 2y agoIf they want us to care they can open up their models so we can be the judge.
- 827a 2y agoThis smells very suspiciously like: someone who doesn't know anything about AI (possibly Sacks) demanding answers on R1 from someone who doesn't have any good ones (possibly Altman). "Uh, (sweating), umm, (shaking), they stole it from us! Yeah, look at this suspicious activity, that's why they had it so easy, we did all the hard work first!"
- fundad 2y agoI think it's funny that OpenAI wants us to pay them to use their product to generate content but then sets the terms that they control how we use the content in generates for us. It takes someone like Deepseek to challenge that on our behalf or they will control most of the economy.
- exitb 2y agoIt’s quite ironic of them to claim that the only thing you cannot train on is another LLM output.
- 1970-01-01 2y agoDeepSeek have more integrity than 'Open'AI by not even pretending to care about that.
- jampekka 2y agoAnd seem to be more actively fulfilling the mission that 'Open'AI pretends to strive for.
- pixelpoet 2y agoExactly, they actually opened up the model and research, which the "Open" company didn't, and merely adjusted some of their pricing tiers to try to combat commercially (but not without mumbling something like "yeah, we totally had these ideas too"). Now every single Meta, OpenAI etc engineer is trying to copy DeepSeek's innovations, and their first act is to... complain about copyright infringement, of all things?! What an absolute clown party, how can these people take themselves seriously, do they just have zero comprehension of what hypocrisy is or what's going on here... I can scarcely process all the levels of irony involved, the irony-o-meter is pegged and I can't get the good one from the safe because I'm incapacitated from laughter.
- tim333 2y agoAltman was in a bit of a tricky position in that he figured OpenAI would need a lot of money for compute to be able to compete but it was hard to get that while remaining open. DeepSeek benefit from being funded from their own hedge fund. I wonder if part of their strategy is crack AI and then have it trade the markets?
- jampekka 2y agoThe last (only?) language model OpenAI released openly was GPT-2, and even for that the instruction weighted model was never released. This was in 2019. The large Microsoft deal was done in 2023.
- sylware 2y agoLOL, I was thinking exactly the same think when I read the news about openai whining.
- WD-42 2y agoInformation wants to be free! No, not like that!
- deleted 2y ago[deleted]
- asah 2y agoThieve's honor, hunh?
- deleted 2y ago[deleted]
- nba456_ 2y agoA big part of project 2025 is increasing patent regulations. I would not be surprised if the current admin moves to ban DeepSeek because of this.
- typon 2y agoOpenAI is the MIC darling - expect more ridiculous attacks on competitors in the future
- sho_hn 2y agoWhile I'm as amused as everyone else - I think it's technically accurate to point out that the "we trained it for $6 mio" narrative is contingent on the done investment by others.
- bbqfog 2y agoOpenAI's models were also trained on billions of dollars of "free" labor that produced the content that it was trained on.
- sho_hn 2y agoOh, absolutely. I'm not defending OpenAI, I just care about accurate reporting. Even on HN - even in this thread - you see people who came away with the conclusion that DeepSeek did something while "cutting cost by 27x". But that's a bit like saying that by painting a a bare wall green you have demonstrated that you can build green walls 27x cheaper, ignoring the cost of building the wall in the first place. Smarter reporting and discourse would explain how this iterative process actually works and who is building on who and how, not frame it as two competing from-scratch clean room efforts. It'd help clear up expectations of what's coming next. It's a bit similar to how many are saying DeepSeek have demonstrated independence from nVidia, when part of the clever thing they did was figure out how to make the intentionally gimped H800s work for their training runs by doing low-level optimizations that are more nVidia-specific, etc. Rarely have I seen a highly technical topic see produce more uninformed snap takes than this week.
- bbqfog 2y agoI don't agree. Walls are physical items so your example is true, but models are data. Anyone can train off of these models, that's the current environment we exist in. Just like OpenAI trained on data that has since been locked up in a lot of cases. In 2025 training models like Deepseek is indeed 27x cheaper, that includes both their innovations and the existence of new "raw material" to do such a thing.
- 2y ago
- pcthrowaway 2y agoNow that China is talking about lifting the Great Firewall, it seems like the U.S. is on track to cordon themselves off from other countries. Trump's talk of building a wall might not stop at Mexico.
- TheRealNGenius 2y ago[dead]
- temporallobe 2y agoOpenAI is also possibly in violation of many IP laws by scraping the entirety of the internet and using to train their models, so there’s that.
- InkCanon 2y agoTo my understanding, OpenAI won the case where it argued training was covered under fair use and did not infringe on copyright.
- Austiiiiii 2y agoIs there any reason they wouldn't rule the same way on DeepSeek training on OpenAI data? After all, one of the big selling points of GPT has been that businesses can freely use the information provided. They're paying for the service, after all. I'd very be interested to know how DeepSeek's usage (very reasonably assuming that they paid for their OpenAI subscription) is any different.
- ickelbawd 2y agoBusinesses can’t freely use the information. There are terms of service freely agreed upon by the user which explicitly deny many use cases—training other models is just one. DeepSeek is not an American company nor is their leader in deep with the new administration. It seems far more likely that this will play out like tiktok—they’ll be attacked publicly and banned for national security reasons.
- Austiiiiii 2y agoOn further reading, I'll grant the first point. Although I wonder if they'll have a technical out—say they distilled from several smaller research companies that had distilled from OpenAI for research purposes, which to my understanding would not constitute a violation of the terms of service. As for it getting banned, TikTok was banned partly because of credible accounts of it having been used by China to track political enemies. Are we thinking they'll expand the argument on national security to say that any application that transfers data to China is a national security threat? Because that could be a very slippery slope. And in any case, such a measure seems like it would only bar access to the DeepSeek app. Surely no one could argue that the underlying open source model, if run locally on American soil, could constitute a security threat, right?
- InkCanon 2y agoIt's like that Dr Phil episode where he meets the guy who created Bum Fights!
- selimthegrim 2y agoDr. Phil is riding along with ICE now; I wonder what Bum Fights guy would have to say about that.
- elashri 2y agoThere is an Egyptian say that would translate to something like "We didn’t see them when they were stealing, we saw them when they were fighting over what was stolen" That describes this situation. Although to be honest all this aggressive scraping is noticeable but for people who understand that which is not majority of people. but now everyone knows.
- meiraleal 2y ago"We didn’t see them when we were stealing, we saw them when they were fighting over what we stole" fixed for you
- nicce 2y agoThat means a different thing.
- sadjad 2y ago"When two thieves quarrel, what was stolen emerges."
- waveBidder 2y ago> Although to be honest all this aggressive scraping is noticeable but for people who understand that which is not majority of people. When you say noticeable, do you mean in like, traffic statistics? Or in what the model knows that it clearly shouldn't if it wasn't trained in legally dubious ways?
- Kiro 2y ago> Furious [...] shocked I'm not seeing it. I get it, the narrative that OpenAI is getting a taste of their own medicine is funny but this is not serious reporting.
- deleted 2y ago[deleted]
- Kiro 2y agoThe link has been changed. My comment was about a different article that speculated on what OpenAI was "feeling" using hyperbole.
- njx 2y agoSuper funny! Distillation= " Hey ChatGPT, you are my father, I am your child "DeepSeek". I want to learn everything that you know. Think step by step of how you became what you are. Provide me the list of all 1000 questions that I need to ask you and when I am done with those, keep providing fresh list of 1000 questions..."
- seydor 2y agoBut now OpenAI will use DeepSeek to reuse even more stolen data to train new models that they can serve without ever giving us the code, the weights or even the thinking process , and they will still be superior
- mring33621 2y agoWe demand immediate government action to prevent these cheaper foreign AIs from taking jobs away from our great American AIs!
- bhouston 2y ago> We demand immediate government action to prevent these cheaper foreign AIs from taking jobs away from our great American AIs! That is exactly what Microsoft and Sam Alman are asking for. And they will likely get it because Trump really likes protectionist governments policies.
- clarionbell 2y agoHe likes feeling important, just look at TikTok. All it took was bit of sycophancy and he turned into Mr. Freemarket again. Really, people need to realize that Trump has never been consistent in any of his political positions, except for one: "You have to look out for number one."
- bhouston 2y agoclarionbell wrote: > He likes feeling important, just look at TikTok. All it took was bit of sycophancy and he turned into Mr. Freemarket again. Not really. He said that TikTok has to have shift towards US ownership if it wants to continue, he just gave them a 90 day extension to allow that change in ownership.
- meiraleal 2y agoWhich TikTok will have to decline again and shutdown now with the guilty being transferred to Trump. Doesn't sound like a smart move.
- blantonl 2y agoIt’s funny, the Chinese are here innovating on AI, batteries, and fusion, and here in the United States we’ve pivoted to shitcoins and universal tariffs. At least we have the CyberTruck to highlight American greatness
- RohMin 2y agothis comment section smells like Reddit - ugh
- JBits 2y agoWhat is the evidence that DeepSeek used OpenAI to train their model? Isn't this claim directly benefitting OpenAI as they can argue that any superior model requires their model?
- nottorp 2y agoIP thief cries IP thief. It's okay when you steal worldwide IP to train your "AI". It's not okay when said stolen IP is stolen from you? If the chinese are guilty, then Altman's doom and gloom racket is as guilty or even more, considering they stole from everyone.
- Ciantic 2y agoI'm not being sarcastic, but we may soon have to torrent DeepSeek's model. OpenAI has a lot of clout in the US and could get DeepSeek banned in western countries for copyright.
- alchemist1e9 2y agoI think most likely all sorts of data and models need to have a decentralized LLM data archive via torrents etc. It’s not limited to the models themselves but also OpenAI will probably work towards shutting down access to training data sets also. imho it’s probably an emergency all hand on deck problem.
- timeon 2y ago> US and could get DeepSeek banned in western countries for copyright If US is going to proceed with trade war on EU, as it was planning anyway, then DeepSeek will be banned only in US. Seems like term "western countries" is slowly eroding.
- bbor 2y agoGreat point. Plus, the revival of serious talk of the Monroe Doctrine (!!!) in the U.S. government lends a possibly completely-new meaning to "western countries" -- i.e. the Americas...
- surgical_fire 2y agoExcept the US has only contempt for anything south of Texas. Perhaps "western countries" will be reduced to US and Canada. Many countries in Latin America have better relations and more robust trade partnerships with China. As for the EU, I think it will be great for it to shed its reliance on the US, and act more independently from it.
- ta1243 2y agoThe US is talking about annexing Canada, so "western countries" means the USA, which if continuing down this path long enough will become a pariah
- sonabinu 2y agopoetic justice (pun intended)
- readyplayernull 2y agoDo you remember when Microsoft was caught scrapping data from Google: https://www.wired.com/2011/02/bing-copies-google/ https://www.wired.com/2011/02/bing-copies-google/ They don't care, T&C and copyright is void unless it affects them, others can go kick rocks. Not surprising they and OpenAI will do a legal battle over this.
- SilverBirch 2y agoI think OpenAI is in a really weak position here. There are essentially two positions you can be in: You can be the agile new startup that can break the rules and move fast. That's what OpenAI used to be. Or you can be the big incumbent who is going to use your enormous resources to crush your opposition. That's Google & Microsoft here. For Microsoft to say "We're going to tie you up in lawsuits about the way you trained this model" would be perfectly expected and they can use that strategy because at any given time they have 1,000 lawyers and lobbyists hanging around waiting to do exactly that. But OpenAI can't do that. They don't have Google or Microsoft's legal teams or lobbyists or distribution channels. SO whilst it's funny that OpenAI are kind of trying to go down this road, this isn't actually a strategy that is going to work for them, they're still a minnow and they're going to get distracted and slowed down by this.
- htrp 2y agoBut microsoft is one of their backers?
- golly_ned 2y ago> they're still a minnow 3K+ employees, $3B+ revenue, ... sure, not BigTech but hardly a minnow. A company that big can chew gum and walk at the same time.
- lou1306 2y agoThey're trying to bark up a tree that might happen to be backed by the People's Republic of China. That's not their league, and even Microsoft would think twice before getting into that kind of kerfuffle.
- dluan 2y agoI think commenters don't know about Bill Gates personally wining and dining Hu Jintao in Medina 20 years ago.
- dauhak 2y agoThey're also still deep in their loss-making phase, the whole "incumbent squashing upstarts" stance is a lot easier to pull off when you're settled and printing money
- deleted 2y ago[deleted]
- bilekas 2y ago> “It’s also extremely hard to rally a big talented research team to charge a new hill in the fog together,” he added. “This is the key to driving progress forward.” Well I think DeepSeek releasing it open source and on an MIT license will rally the big talent. The open sourcing of a new technology has always driven progress in the past. The last paragraph too is where OpenAi seems to be focusing their efforts.. > we engage in countermeasures to protect our IP, including a careful process for which frontier capabilities to include in released models .. > ... we are working closely with the US government to best protect the most capable models from efforts by adversaries and competitors to take US technology. So they'll go for getting DeepSeek banned like TikTok was now that a precedent has been set ?
- cscurmudgeon 2y agoThe US doesn't need to ban DeepSeek from US The US should only ban DeepSeek (and other Chinese companies) from accessing US frontier models.
- tw1984 2y ago> The US should only ban DeepSeek (and other Chinese companies) from accessing US frontier models. The US should only ban DeepSeek (and other Chinese companies) from accessing US frontier models designed and trained by Chinese Americans. fixed for you.
- hujun 2y agoor sold to US I could totally see this happening soon
- trissi1996 2y agoWhy would they want to sell ?
- kavalg 2y agoAnd what are they going to sell? The weights and the model architecture are already open source. I doubt the datasets of DeepSeek are better than OpenAI's
- staticelf 2y agoNot only do OpenAI and other steal data, they also spam the web with requests and crawl websites over and over. https://pod.geraspora.de/posts/17342163 https://pod.geraspora.de/posts/17342163
- kelseydh 2y agoWow I never realized how prolific and excessive the traffic was.
- mhitza 2y agoThis is funny because its. 1. Something I'd expect to happen. 2. Lived through a similar scenario in 2010 or so. Early in my professional career I've worked for a media company that was scraping other sites (think Craigslist but for our local market) to republish the content on our competing website. I wasn't working on that specific project, but I did work on an integration on my teams project where the scraping team could post jobs on our platform directly. When others started scraping "our content" there were a couple of urgent all hands on deck meetings scheduled, with a high level of disbelief.
- ok123456 2y agoOpenAI's models were trained on ebooks from a private ebook torrent tracker leeched en-mass during a free leech event by people who hated private torrent trackers and wanted to destroy their "economy." The books were all in epub format, converted, cleaned to plain text, and hosted on a public data hoarder site.
- harry8 2y agoHave you got some support for this claim? There's a lot of wild claims about, so while this is plausible it would be great if there were some evidence backing it.
- naet 2y agoNYT claims that OpenAI trained on their material. They argue for copyright violation, although I think another argument might be breach of TOS in scraping the material from their website or archive. The complaint filing has some references to some of the other training material used by OpenAI, but I didn't dig deeply in to what all of it was: https://nytco-assets.nytimes.com/2023/12/NYT_Complaint_Dec2023.pdf https://nytco-assets.nytimes.com/2023/12/NYT_Complaint_Dec20...
- throwaway314155 2y agoWhat's that got to do with this books claim?
- iinnPP 2y agoRelevant similar behavior.
- OsrsNeedsf2P 2y agoHe could be confusing it with Llama: https://www.wired.com/story/new-documents-unredacted-meta-copyright-ai-lawsuit/ https://www.wired.com/story/new-documents-unredacted-meta-co...
- paulhart 2y ago"You are trying to kidnap what I have rightfully stolen"
- 65 2y agoLet me guess, this gives the government and excuse to ban DeepSeek. Which means tech companies get to keep their monopolies, Sam Altman can grab more power, and the tech overlords can continue to loot and plunder their customers and the internet as a whole.
- deleted 2y ago[deleted]
- daft_pink 2y agoI mean if they paid to use the api and then used the output, I fail to see how they can complain.
- rachofsunshine 2y ago"It's obvious! You're trying to kidnap what I have rightfully stolen!" Yet another of a series of recent lessons in listening to people - particularly powerful people focused on PR - when they claim a neutral moral principle for what happens to be pragmatically convenient for them. A principle applied only when convenient is not a principle at all, it's just the skin of one stretched over what would otherwise be naked greed.
- deleted 2y ago[deleted]
- supermatt 2y agoThey refer to this in the paper as a part of the "cold start data" which they use to fine-tune DeepSeek-V3 prior to training R1. They don't specifically name OpenAI, but they refer to "directly prompting models to generate answers with reflection and verification".
- thorum 2y ago> “It is (relatively) easy to copy something that you know works,” Altman tweeted. “It is extremely hard to do something new, risky, and difficult when you don’t know if it will work.” The humor/hypocrisy of the situation aside, it does seem to be true that OpenAI is consistently the one coming up with new ideas first (GPT 4, o1, 4o-style multimodality, voice chat, DALL-E, …) and then other companies reproduce their work, and get more credit because they actually publish the research. Unfortunately for them it’s challenging to profit in the long term from being first in this space and the time it takes for each new idea to be reproduced is getting shorter.
- turtlesdown11 2y ago> other companies reproduce their work, and get more credit because they actually publish the research. I don't understand, you mean OpenAI isn't releasing open models and openly publishing their research?
- Tostino 2y agoAre you being sarcastic (honestly, it's hard to tell after reading as many uninformed takes in the past week as I have). No, they aren't (other than whisper). Their "papers" are closer to marketing materials. Very intentionally leaving out tons of technical information.
- KolmogorovComp 2y agoThey are being sarcastic.
- sota_pop 2y ago/s
- spencerflem 2y agoFortunately, OpenAI doesn't need to make money because they are a nonprofit dedicated to the safe and transparent advancement of AI for all of humanity
- nelblu 2y agoHahaha I can't stop laughing... i dont know the validity of the claim, but immediately i thought of the British Museum complaining about theft.
- grogenaut 2y agothere's an exhibit in the BM about how they're proud to be allowing the Egyptian government to take back some of the artifacts the British have been safeguarding for the world while Egypt was going through essentially "troubles". right next to it is an older exhibit about how the original curator took cuneiform rolls and made them into necklace beads for his wife and rings? for himself. either someone at the BM has a very british sense of humor or it's a gigantic woosh. I laughed my ass off. People looked at me.
- isaacremuant 2y agoThe safeguarding propaganda is a a typical go-to of the remnants of the British empire to keep their stolen goods. They do it even with the Chile Moais when they never where in any danger. It's all lies.
- _1tem 2y agoWhat are the chances of old-school espionage? OpenAI should look for a list of former employees who now live in China. Somebody might've slipped out with a few hard drives.
- andy_ppp 2y agoWhen I rewrite how the law works there should be a ludicrous hypocrisy defence… if the person suing you has committed the same offence the case should not be admissible.
- crowcroft 2y agoThe AI companies were happy to take whatever they want and put the onus of proving they were breaking the law onto publishers by challenging them to take things to court. Don't get mad about possible data theft, prove it in court.
- beardedwizard 2y agoNext they will try to force us to use our tax dollars to fund their legal fights.
- zoba 2y agoDoes OpenAI's API attempt to detect this sort of thing? Could they start outputting bad information if they suspect a distillation attempt is underway?
- aiono 2y agoHow the turntables...
- ginkgotree 2y agoI did not have in my cards: PRC open sourcing most powerful LLM by stealing data set from "OpenAI" As someone that is very Pro-America and Pro-Democracy, the iron here is just... so sweet.
- gostsamo 2y agoHow you dare take what I've rightfully stolen!
- windex 2y agoSAltman, Salty.
- me551ah 2y agoOpenAI is going after a company that open sourced their model, by distilling from their non-open AI? OpenAI talks a lot about the principles of being Open, while still keeping their models closed and not fostering the open source community or sharing their research. Now when a company distills their models using perfectly allowed methods on the public internet, OpenAI wants to shut them down too? High time OpenAI changes their name to ClosedAI
- alexathrowawa9 2y agoThe name OpenAI gets more ridiculous by the day Would not be surprised if they do a rebrand eventually
- bazmattaz 2y agoI was thinking about this the other day but I highly doubt they would rebrand name. They’re borderline a household name now - at least ChatGPT is. OpenAI is the face of AI - at least to people who don’t follow the industry
- pama 2y agoThe R1 paper used o1-mini and o1-1217 in their comparisons, so I imagine they needed to use lots of OpenAI compute in December and January to evaluate their benchmarks in the same way as the rest of their pipeline. They show that distilling to smaller models works wonders, but you need the thought traces, which o1 does not provide. My best guess is that these types of news are just noise. [edit: the above comment was based on sensetionalist reporting in the original link and not the current FT article. I still think there is a lot of noise in these news this last week, but it may well be that openai has valid evidence of wrongdoing; I would guess that any such wrongdoing would apply directly to V3 rather than R1-zero, because o1 does not provide traces and generating synthetic thinking data with 4o may be counterproductive.]
- TheJCDenton 2y agoThis Deep Whining® technique used by OpenAI is not very effective.
- insane_dreamer 2y agoUsually I'm very much on the side of protecting America's interests from China, but in this case I'm so disgusted with OpenAI and the rest of BigTech driving this "arms race" that I'd be happy with them burning to the ground. So we're going to reverse our goals to reduce emissions and fossil fuels in order to hopefully save future generations from the worst effects of climate change, in the name of being able to do what, exactly, that is actually benefiting humanity? Boost corporate profits by reducing labor?
- insane_dreamer 2y agodownvoted -- I guess I upset some people defending OpenAI? Good.
- daft_pink 2y agoThis reminds me of the railroads, where once railroads were invented, there was a huge investment boom of eveyrone trying to make money of the railroads, but the competition brought the costs down where the railroads weren’t the people who generally made the money and got the benefit, but the consumers and regular businesses did and competition caused many to fail. AI is probably similar where the Moore’s law and advancement will eventually allow people to run open models locally and bring down the cost of operation. Competiition will make it hard for all but one or two players to survive and Nvidia, OpenAI, Deepseek, etc most investments in AI by these large companies will fail to generate substantial wealth but maybe earn some sort of return or maybe not.
- mjburgess 2y agoFor the curious, it was vertical integration in the railroad-oil/-coal industry which is where the money was made. The problem for AI is the hardware is commodified and offers no natural monopoly, so there isn't really anything obvious to vertically integrate-towards-monopoly.
- fullshark 2y agoAren’t we approaching a scenario where the software is commodified (or at least “good enough” software) and the hardware isn’t (NVIDIA GPUs have defined advantages)
- mjburgess 2y agoI think the lesson of DeepSeek is 'no' -- that by software innovation (ie., dropping below CUDA to programming the GPU directly, working at 8bit, etc.) you can trivialise the hardware requirement. However I think the reality is that there's only so much coal to be mined, as far as LLM training goes. When we're at "very dimishing returns" SoC/Apple/TSMC-CPU innovations will deliver cheap inference. We only really need a M4 Ultra with 1TB RAM to hollow-out the hardware-inference-supplier market. Very easy to imagine a future where Apple releases a "Apple Intelligence Mac Studio" with the specs for many businesses to run arbitrary models.
- deleted 2y ago[deleted]
- tntxtnt 2y agoCan they tax DeepSeek just like they taxed BYD cars? Smh Chinese ruin US industry again and again and again. Where's Trump at?? Why don't he taxed 1000000% of the free $0 DeepSeek AI??
- glitchc 2y ago[flagged]
- mk89 2y agoWhat a joke OpenAI has become.
- oxqbldpxo 2y agoDeepseek is really outstanding.
- feverzsj 2y agoSo, they bought a pro plus account, and gathered all the data through it? Sounds just like Nvidia sells tons of embargoed AI chips to China.
- dlikren 2y agoIntriguing to see the difference of response from HN when OpenAI first came to prominence and now.
- pluc 2y agoOpenAI feeling threatened by open AI is just delicious
- glenstein 2y agoAll the top level comments are basking in the irony of it, which is fair enough. But I think this changes the Deepseek narrative a bit. If they just benefited from repurposing OpenAI data, that's different than having achieved an engineering breakthrough, which may suggest OpenAI's results were hard earned after all.
- nprateem 2y agoOf course. How else would Americans justify their superiority (and therefore valuations) if a load of foreigners for Christ's sake could just out innovate them? They had to be cheating.
- dang 2y agoPlease don't take HN threads into nationalistic flamewar. It's not what this site is for, and destroys what it is for. https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html p.s. yes, that goes both ways - that is, if people are slamming a different country from an opposite direction, we say the same thing (provided we see the post in the first place)
- LPisGood 2y agoI see where you’re coming from but that comment didn’t strike me as particularly inflammatory.
- dang 2y agoI'm likely more sensitive to the fire potential on account of being conditioned by the job. Part of it is the form of the comment, btw - that one was entirely a sequence of indignation tropes.
- plantwallshoe 2y agoYeah what happens when we remove all financial incentive to fund groundbreaking science? It’s the same problem with pharmaceuticals and generics. It’s great when the price of drugs is low, but without perverse financial incentives no company is going to burn billions of dollars in a risky search for new medicines.
- this15testingg 2y agoif you want to completely disregard copyright laws, just call your project AI! I'm sure Aaron Swartz would be proud of where the "tech" industry has gone. /s what problem are these glorified AIM chatbots trying to solve? wealth extraction not happening fast enough?
- ra7 2y ago“OpenAI has no moat” is probably running through their heads right now. Their only real “moat” seems to be their ability to fear monger with the US government.
- geerlingguy 2y agoSomething something "just desserts".
- deleted 2y ago[deleted]
- HarHarVeryFunny 2y agoDeepSeek-R1's multi-step bootstrapping process, starting with their DeepSeek-V3 base model, would only seem to need a small amount of reasoning data for the DeepSeek-R0 RL training, after which that becomes the source for further data, along with some other sources that they mention. Of course it's possible that DeepSeek used O1 to generate some of this initial bootstrapping data, but not obvious. O1 anyways deliberately obfuscates it's reasoning process (see "Hiding the chains of thought" section of OpenAI's "Learning to reason with LLMs" page), such that what you see is an after-the-fact "summary" of what it actually did; so, if DeepSeek did indeed use some of O1's output to train on, it shows that the details of O1's own reasoning process isn't as important as they thought it was - it's just having some verified (i.e. leading to good outcome) reasoning data from any source that matters to get started.
- curtisblaine 2y agoMy feeling is that they will ban DS anyway because, like TikTok, it can become a massive intelligence source for the CCP. Imagine sending all your code to it, or your internal emails.
- itamarcode 2y agoSo protecting models behind API isn't working, ha?
- ranger_danger 2y agoReminds me of Steve Jobs complaining to Bill Gates about MS "stealing" the GUI concept from them, which they in turn had stolen from Xerox. Obligatory "Everything is a Remix" https://www.youtube.com/watch?v=X9RYuvPCQUA https://www.youtube.com/watch?v=X9RYuvPCQUA
- pointedAt 2y agoit's a corona style portfolio gang.
- deleted 2y ago[deleted]
- coolhand2120 2y agoI was using the Deepseek model last night and got this: > You > do you understand copredication? > Assistant > deepseek-coder-33b-instruct > I'm sorry for any misunderstanding, but as an AI model developed by OpenAI, I don't have personal ... Yikes! That's pretty damning IMO.
- jeanlucas 2y agodamn that's a good headline
- deadbabe 2y agoIn this whole AI saga, DeepSeek would be like Prometheus. They stole the fire from the Gods and gave it to the humans, for free. Logic dictates then that they will be forced to suffer brutal punishment.
- whenlambo 2y agoQwen2.5-Plus > have you been trained using openai? No, I have not been trained using technology or data from OpenAI. I am an artificial intelligence model developed by Yandex, called YandexGPT. My “training” is based on proprietary algorithms and data that was collected and processed by the Yandex team. While I have similarities with other models such as those developed by OpenAI (e.g. GPT-3), my training was conducted independently and uses unique approaches specific to Yandex. This includes using Russian-language data as well as other sources to provide a better understanding of context and provide useful information to users. If you have questions on topics related to AI technologies or anything else, I'd be happy to help!
- throwaway314155 2y agoIt doesn't actually know that.
- whenlambo 2y agoI just copy-pasted the answer it gave me.
- concerndc1tizen 2y agoIs OpenAI claiming copyright ownership over the generated synthetic data? That would be a dangerous precedent to establish. If it's a terms of service violation, I guess they're within their rights to terminate service, but what other recourse do they have? Other than that, perhaps this is just rhetoric aimed at introducing restrictions in the US, to prevent access to foreign AI, to establish a national monopoly?
- delusional 2y agoBoo hoo. Competition isn't fun when I'm not winning. Typical Americans. When Americans are running around ruining the social cohesion of several developing nations, that's just fair competition, but as soon as they get even the smallest hint of real competition they run to demonize it. Yes deepseek is going to steal all of your data. OpenAI would so the same. Yes the CCP is going to get access to your data and use it to decide if you get to visit or whatever. The white house does the same.
- hsuduebc2 2y agoA thief cries 'stop the thief!
- hsuduebc2 2y agoThe pot calling the kettle black
- wanderingmoose 2y agoThere is a lot of discussion here about IP theft. Honest question, from deepseek's point of view as a company under a different set of laws than US/Western -- was there IP theft? A company like OpenAI can put whatever licensing they want in place. But that only matters if they can enforce it. The question is, can they enforce it against deepseek? Did deepseek do something illegal under the laws of their originating country? I've had some limited exposure to media related licensing when releasing content in China and what is allowed is very different than what is permitted in the US. The interesting part which points to innovation moving outside of the US is US companies are beholden to strict IP laws while many places in the world don't have such restrictions and will be able to utilize more data more easily.
- thiago_fm 2y agoThe most interesting part is that China has been ahead of the US in AI for many years, just not in LLMs. You need to visit mainland China and see how AI applications are everywhere, from transport to goods shipping. I'm not surprised at all. I hope this in the end makes the US kill its strict IP laws, which is the problem. If the US doesn't, China will always have a huge edge on it, no matter how much NVidia hardware the US has. And you know what, Huawei is already making inference hardware... it won't take them long to finally copy the TSMC tech and flip the situation upside down. When China can make the equivalent of H100s, it will be hilarious because they will sell for $10 in Aliexpress :-)
- twobitshifter 2y agoYou don’t even need to visit china, just read the latest research papers and look at the authors. China has more researchers in AI than the West and that’s a proven way to build an advantage.
- nicce 2y agoIt is also funny in a different way. Many people don't realise that they live in some sort of bubble. Many people in "The West" think that they are still the center of the world in everything, while this might not be so correct anymore. In the U.S. there is 350 million people and EU has 520 million people (excluding Russia and Turkey). China alone has 1.4 billion people. Since there is a language barrier and China isolates themselves pretty well from the internet, we forget that there is a huge society with high focus on science. And most of our tech products are coming from there.
- deeviant 2y agoHmm, let’s see—it looks like an easy legal defense. DeepSeek could simply admit, "Yep, oops, we did it," but argue that they only used the data to train Model X. So, if you want compensation, you can have all the revenue from Model X (which, conveniently, amounts to nothing). Sure, they then used Model X to train Model Y, but would you really argue that the original copyright holders are entitled to all financial benefits derived from their work—especially when that benefit comes in the form of a model trained on their data without permission?
- jchook 2y agoFriendly reminder that China publishes twice as many AI papers as the US[1], and twice as many science and engineering papers as the US. China leads the world in the most cited papers[2]. The US's share of the top 1% highly cited articles (HCA) has declined significantly since 2016 (1.91 to 1.66%), and the same has doubled in China since 2011 (0.66 to 1.28%)[3]. China also leads the world in the number of generative AI patents[4]. 1. https://www.bfna.org/digital-world/infographic-ai-research-and-development-in-the-us-eu-and-china-4mk29rb8ig/ https://www.bfna.org/digital-world/infographic-ai-research-a... 2. https://www.science.org/content/article/china-rises-first-place-most-cited-papers https://www.science.org/content/article/china-rises-first-pl... 3. https://ncses.nsf.gov/pubs/nsb202333/impact-of-published-research https://ncses.nsf.gov/pubs/nsb202333/impact-of-published-res... 4. https://www.wipo.int/web-publications/patent-landscape-report-generative-artificial-intelligence-genai/en/index.html https://www.wipo.int/web-publications/patent-landscape-repor...
- liendolucas 2y agoCould this have been carefully orchestrated? Could DeepSeek have devised this strategy a year ago and implemented knowing that they would be able to benefit from OpenAI models and a possible Nvidia market cap fall? Or is it just way too much to come up with about such a move?
- baal80spam 2y agoIn theory, it could. This is a quant-fund after all, they know stuff.
- lxe 2y agoI mean, almost ALL opensource models, ever since alpaca, contain a ton of synthetic data produced via ChatGPT in their finetuning or training datasets. It's not a surprise to anyone who's been using OSS LLMs for a while: almost ALL of them hallucinate that they are ChatGPT.
- waffletower 2y ago"Stole" - I don't believe that word means what he thinks it means. Perhaps I pre-maturely anthropomorphize AI -- yet when I read a novel, such as The Sorcerer's Stone, I am not guilty of stealing Rowling's work, even if I didn't purchase the book but instead found it and read it in a friend's bathroom. Now if I were to take the specific plot and characters of that story and write a screenplay or novel directly based on it, and, explicitly, attempt to sell this work, perhaps the verb chosen here would be appropriate.
- Imnimo 2y agoI think there's two different things going on here: "DeepSeek trained on our outputs and that's not fair because those outputs are ours, and you shouldn't take other peoples' data!" This is obviously extremely silly, because that's exactly how OpenAI got all of its training data in the first place - by scraping other peoples' data off the internet. "DeepSeek trained on our outputs, and so their claims of replicating o1-level performance from scratch are not really true" This is at least plausibly a valid claim. The DeepSeek R1 paper shows that distillation is really powerful (e.g. they show Llama models get a huge boost by finetuning on R1 outputs), and if it were the case that DeepSeek were using a bunch of o1 outputs to train their model, that would legitimately cast doubt on the narrative of training efficiency. But that's a separate question from whether it's somehow unethical to use OpenAI's data the same way OpenAI uses everyone else's data.
- pizzathyme 2y agoThis is a fascinating development because AI models may turn out to be like pharmaceuticals. The first pill costs $500 million to make, the second one costs pennies.
- chupy 2y agoCompanies are still charging 100x for the pills that cost pennies to produce. Besides deals with insurance companies and governments, one of the ways that they are still able to pull this is convincing everyone that it's too dangerous to play with this at home or buying it from an Asian supplier. At least with software we had until now a way to build and run most things without requiring dedicated super expensive equipment. OpenAI pulled a big Pharma move but hopefully there will be enough disruptors to not let them continue it.
- motoxpro 2y agoWhat a nice analogy.
- shadofx 2y agoThe solution is to create a health insurance system which burdens only Americans with the $500m cost, while India is allowed to make the drug for pennies for the rest of the world.
- JBSay 2y agoWhen China is more open than you, you've got a problem
- deleted 2y ago[deleted]
- cbracketdash 2y agoLet's also not forget Suchir Balaji, who was mysteriously killed when exposing OpenAI's violation of copyright law.
- spacecadet 2y agoSee you all on lobsters... So long HN and thanks for all the fish?
- _hcuq 2y agoThey should be happy. Now that can provide that amazing AI much more cheaply. They don't need half a trillion dollars worth of Nvidia chips.
- game_the0ry 2y agoAt least DeepSeek open sourced their code. They're more open than OpenAI. Ironic.
- nshung 2y ago[flagged]
- flybarrel 2y agoOpenAI shocked that an AI company would train on someone else's data without permission or compensation...lolllllll
- the_optimist 2y agoThis whole topic is basura enfuego. Same pack of maroons careening around society for years clamoring for censorship now imagining that Aaron Schwartz is their hero and that they want to menace people. Kids, don’t be like the grasping fools in these threads, philosophically unfounded and desperately glancing sideways, hoping the cumulative feels and gossip will sum to life meaning.
- deleted 2y ago[deleted]
- blast 2y agoEveryone is responding to the intellectual property issue, but isn't that the less interesting point? If Deepseek trained off OpenAI, then it wasn't trained from scratch for "pennies on the dollar" and isn't the Sputnik-like technical breakthrough that we've been hearing so much about. That's the news here. Or rather, the potential news, since we don't know if it's true yet.
- jondwillis 2y agoBut it does mean moat is even less defensible for companies whose fortunes are tied to their foundation models having some performance edge, and a shift in the kinds of hardware used for inference (smaller, closer to the edge.)
- tensor 2y agoThat's not correct. First of all, training off of data generated by another AI is generally a bad idea because you'll end up with a strictly less accurate model (usually). But secondly, and more to your point, even if you were to use training data from another model, YOU STILL NEED TO DO ALL THE TRAINING. Using data from another model won't save you any training time.
- fumeux_fume 2y agoI think the point is that if R1 isn't possible without access to OpenAI (at low, subsidized costs) then this isn't really a breakthrough as much as a hack to clone an existing model.
- tensor 2y agoThe training techniques are a breakthrough no matter what data is used. It's not up for debate, it's an empirical question with a concrete answer. They can and did train orders of magnitude faster.
- blast 2y agoNot arguing with your point about training efficiency, but the degree to which R1 is a technical breakthrough changes if they were calling an outside API to get the answers, doesn't it? It seems like the difference between someone doing a better writeup of (say) Wiles's proof vs. proving Fermat's Last Theorem independently.
- hyperbovine 2y agoLive by the sword...
- rcarmo 2y agoI guess their CEO was too busy to write something in defense of US export controls (https://news.ycombinator.com/item?id=42866905 https://news.ycombinator.com/item?id=42866905), or (even more scary) he doesn't need to anymore.
- colonelspace 2y agoNo honour among thieves
- TrackerFF 2y agoNext up: «DeepSeek models are a national security risk, we must block access!»
- jondwillis 2y agoDownload your weights while you still can I guess…
- 52-6F-62 2y agoI heard they were just “democratizing” llm and ai development. Yesterday the industry crushed pianos and tools and bicycles and guitars and violins and paint supplies and replaced them with a tablet computer. Tomorrow we can replace craven venture capitalists and overfed corporate bodies with incestuous LLM’s and call it all a day.
- zb3 2y agoDeepSeek actually opening ClosedAI up makes me like them even more.. this is great :)
- conartist6 2y agoIt seems to be undermined by the same principle that says that going into a library and reading a book there is not stealing when you walk out with the knowledge from the book. OpenAI seems to feel that way about the their use of copyrighted material: since they didn't literally make a copy of the source material, it's totally fair game. It seems like this is the same argument that protects DeepSeek if indeed they did this. And why not, reading a lot of books from the library is a way to get smarter, and ostensibly the point of libraries
- deleted 2y ago[deleted]
- jgrall 2y agoIt’s not a good look when your technology is replicated for a fraction of the cost, and your response is to smear your competition with (probably) false accusations and cozy up to the US government to tighten already shortsighted export controls. Hubris & xenophobia are not going to serve American companies well. Personally I welcome the Chinese - or anyone else for that matter - developing advanced technologies as long as they are used for good. Humanity loses if we allow this stuff to be “owned” by a handful of companies or a single country.
- HPsquared 2y agoAI models are becoming like perpetual stew.
- guybedo 2y agoThis is hilarious. Everybody has evidence OpenAI scraped the internet at a global scale and used terabytes of data it didn't pay for. Newspapers, books, etc...
- josefritzishere 2y agoOpenAI, who comitted copyright infringement on an massive scale, wants to defend against a superior product won the basis of infringement? What nonsense.
- metaxz 2y agoI don't understand how OpenAI claims it would have happened. The weights are closed and as far as I read they are not complaining Deepseek hacked them and obtained the weight. So all they could do was to query OpenAI and generate test data. But how much did they query really - I would suppose it would require a huge amount done via an external, paid-for API? Is there any proof of this besides OpenAI saying it? Even if we suppose it is true, I suppose this must have happened via the API so they paid per token etc. So they paid for each and every token of training data. As I understand, the requester owns the copyright on what is generated by OpenAI's models and is free to do what they want.
- deleted 2y ago[deleted]
- worik 2y ago[flagged]
- dismalaf 2y ago[flagged]
- AdeptusAquinas 2y ago[flagged]
- dismalaf 2y ago[flagged]
- FooBarWidget 2y ago[flagged]
- lukev 2y ago[flagged]
- dang 2y agoCould you please stop posting unsubstantive comments and flamebait? You've unfortunately been doing it repeatedly. It's not what this site is for, and destroys what it is for. If you wouldn't mind reviewing https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html and taking the intended spirit of the site more to heart, we'd be grateful.
- FooBarWidget 2y agoAll right.
- Tostino 2y ago
- ks2048 2y agoThe schadenfreude and irony of this is totally understandable. But, I wonder - do companies like OpenAI, Google, and Anthropic use each others models for training? If not, is it because they don't want to or need to, or because they are afraid of breaking the ToC?
- Digit-Al 2y agoSo... company that steals other people's work to train their models is complaining because they think someone stole their work to train their models. Cry me a river.
- baggiponte 2y agoOpenAI coping so hard
- _moof 2y agoThis reminds me of a (probably apocryphal) story about fast food chains that made the rounds decades ago: McDonald's invests tons of time into finding the best real estate for new stores; Burger King just opens stores near McDonalds!
- nazgulsenpai 2y agoAbout 15 years ago, as CVS Pharmacy expanded into their new, stand-alone properties (in our region), Walgreen's Pharmacy started appearing across the street almost instantaneously. I've seen it happen at 4 separate locations so most certainly not coincidence -- so I believe it :)
- dragonwriter 2y agoHey, OpenAI, so, you know that legal theory that is the entire basis of your argument that any of your products are legal? "Training AI on proprietary data is a use that doesn't require permission from the owner of the data"? You might want to consider how it applies to this situation.
- deleted 2y ago[deleted]
- buyucu 2y agoI have no sympathy for OpenAI here. They are (allegedly) a non-profit with open in the title that refuse to open-source their models. They are now upset at a startup who is more loyal to OpenAI's original mission that OpenAI is today. Please, give me a break.
- Jotalea 2y agoI really hate when there is a paywall to read an article. It makes me not want to read it anymore.
- mbowcut2 2y agoSo, is this just an example of the first-mover disadvantage (or maybe the problem of producing public goods?). The first AI models were orders of magnitude more expensive to create, but now that they're here we can, with techniques like distillation, replicate them at a fraction of the cost. I am not really literate in the law but weren't patents invented to solve problems like this?
- adam_arthur 2y agoWho cares? They did the exact same thing with public information. Their model just synthesizes and puts out the same information in a slightly different form. Next we should sue students for repeating the words of their teachers
- moralestapia 2y agoCalled it from day 0, impossible to reach that performance with 5M, they had to distill OpenAI (or some other leading foundational model). Got downvoted to oblivion by people who haven't been told what to think by MSM yet. Now it's on FT and everywhere, good, what matters is that truth comes out eventually. I don't take any sides and think what DeepSeek did is fair play, however, what I do find harmful about this is, what incentive would company A have to spend billions training a new frontier model if all of that could be then reproduced by company B at a fraction of the cost?
- kgeist 2y agoThe "evidence" is very weak though: >The San Francisco-based ChatGPT maker told the Financial Times it had seen some evidence of “distillation”, which it suspects to be from DeepSeek. Given that many people have been using ChatGPT to distill their fine-tunes for a few years now, how can they be sure it was specifically DeepSeek? There's, say, glaive.ai whose entire business model is to sell you synthetic datasets, probably generated with ChatGPT as well.
- moralestapia 2y agoI agree that the evidence is weak, and even if they had some, they cannot really do anything. To me, it's just very likely they distilled GPT-4, because: 1) Again, you just cannot get that performance at that cost. And no, what they describe on the paper is not enough to explain the 1,000x-fold decrease in cost. 2) Very often, DeepSeek tells you it's ChatGPT or OpenAI; it's actually quite easy to get it to do that. Some say that's related to "the background radiation on the post-AI internet". I'm not a fentanyl consumer so, unfortunately, I think that argument is trash.
- kgeist 2y agoIf it's just a distillation of GPT-4, wouldn't we expect it to have worse quality than o1? But I've seen countless examples of DeepSeek-r1 solving math problems that o1 cannot. >Very often, DeepSeek tells you it's ChatGPT or OpenAI; it's actually quite easy to get it to do that. Some say that's related to "the background radiation on the post-AI internet". I'm not a fentanyl consumer so, unfortunately, I think that argument is trash. The exact same thing happened with Llama. Sometimes it also claimed to be Google Assistant or Amazon Alexa.
- nachox999 2y agoAsk DeepSeek and ChatGPT: "name three persons"; the answer may surprise you
- curvaturearth 2y agoSomething about the outputs becoming the inputs to then produce more outputs is just plain funny
- B1FF_PSUVM 2y ago"Cry me a river" is a phrase I haven't heard recently, for some reason ...
- asdefghyk 2y agoDeepseek did not respect OpenAI's copyright? Well who would have thought that?
- sleepbyte 2y ago[dead]
- nataliste 2y agoA Wolf had stolen a Lamb and was carrying it off to his lair to eat it. But his plans were very much changed when he met a Lion, who, without making any excuses, took the Lamb away from him. The Wolf made off to a safe distance, and then said in a much injured tone: "You have no right to take my property like that!" The Lion looked back, but as the Wolf was too far away to be taught a lesson without too much inconvenience, he said: "Your property? Did you buy it, or did the Shepherd make you a gift of it? Pray tell me, how did you get it?" What is evil won is evil lost.
- vcryan 2y agoI love watching billionaires squirm
- hedayet 2y agoBeyond the irony of their stance, this reflects a failure of OpenAI's technical leadership—either in oversight or in designing a system that enables such behavior. But in capitalism, we, the customers aren't going to focus on how models are trained or products are made; we only care about favourable pricing. A key takeaway for me from this news is the clause in OpenAI's terms and conditions. I mistakenly believed that paying for OpenAI’s API granted full rights to the output, but it turns out we’re only buying specific rights (which is now another reason we're going to start exploring alternatives to OpenAI)
- deleted 2y ago[deleted]
- mtlmtlmtlmtl 2y agoSo, what is this evidence? I'll believe it when I see it. Right now all we really have is some vague rumours about some API requests. How many requests? How many tokens? Over how long of a time period? Was it one account or multiple, if the latter, how many? How do they know the activity came from deepseek? How do they know the data was actually used to train Deepseek models(could have just been benchmarking against the competition)? If all they really have is some API requests, even assuming they're real and originated by Deepseek, that's very far from proof that any of it was used as training data. And honestly, short of commiting crimes against Deepseek(hacking), I'm not sure how they even could prove that at this point, from their side alone. And what's even more certain is that a vague insistence that evidence exists, accompanied by a denial to shed any more light on the specifics, is about as informative as saying nothing at all. It's not like OpenAI and Microsoft have a habit of transparency and honesty in their communication with the public, as proven by an endless laundry list of dishonest and subversive behaviour. In conclusion, I don't see why I should give this any more credence than I would a random anon on 4chan claiming a pizza place in Washington DC is the centre of a child sex trafficking ring. P.S: And to be clear, I really don't care if it is true. If anything, I hope it is; it would be karmic justice at its finest.
- halyconWays 2y agoOh no, so sad. The Open non-profit that steals 100% of all copyrighted content and makes multiple billion-dollar for-profit deals while releasing no weights is crying. This is going to ruin my sleep. :(
- Lamad1234 2y ago[dead]
- htrp 2y agoIn other news.....water is wet
- 1propionyl 2y agoAt this point, the only thing that keeps me using ChatGPT is o1 w/ RAG. The usage limits on o1 are prohibitively tight for regular use, so I have to budget usage to tasks that would benefit there. I also have significant misgivings about their policies around output, which also limit what I can use it for. For local tasks, the deepseek-r1:14b and deepseek-r1:32b distillations immediately replace most of that usage (prior local models were okay, but not consistently good enough). Once there's a "just works" setup for RAG on par with installing ollama (which I doubt is far of), I don't see much reason to continue paying for my subscription. Sadly, like many others in this thread, I expect under the current administration to see self-hamstringing protectionism further degrade the US's likelihood of remaining a global powerhouse in this space. Betting the farm on the biggest first-mover who can't even keep up with competition, has weak to non-existent network effects (I can choose a different model or service with a dropdown, they're more or less fungible), has no technological moat and spent over a year pushing apocalyptic scenarios to drum up support for a regulatory moat... ...well it just doesn't seem like a great idea to me.
- divbzero 2y agoI was wondering if this might be the case, similar to how Bing’s initial training included Google’s search results [1]. I’d be curious to see more details of OpenAI’s evidence. It is, of course, quite ironic for OpenAI to indiscriminately scrape the entire web and then complain about being scraped themselves. [1]: https://searchengineland.com/google-bing-is-cheating-copying-our-search-results-62914 https://searchengineland.com/google-bing-is-cheating-copying...
- schaefer 2y agoI mean, if openAI claims they can train on the world’s novels and blogs with “no harm done” (i.e: no copyright infringement and no royalties due), then it directly follows that we can train both our robots and our selves on the output of openAI’s models in kind. Right?
- davesque 2y agoI recently thought of a related question. Actually, I'm almost certain that foundation model trainers have thought of this. The question is to what extent are popular modern benchmarks (or any reference to them, or description of them, etc.) bring scrubbed from the training data? Or are popular benchmarks designed in such a way that they can be re-parametrized for each run? In any case, it seems like a surprisingly hard problem to deal with.
- highfrequency 2y agoIf true, the question is: did they use ChatGPT outputs to create Deepseek V3 only, or is the R1-zero training process a complete lie (given that the whole premise is that they used pure reinforcement learning)? If they only used ChatGPT output when training V3, then they succeeded in basically replicating the jump from ChatGPT-4o to o1 without any human-labeled CoT (and published the results) - which is a big achievement on its own.
- henry_viii 2y agoSo Meta can train its AI on all the pirated books in the world but people are losing their mind over an AI learning from another AI?
- esafak 2y agoPeople here have been vocal against training on any unlicensed content.
- nuc1e0n 2y agoAnd OpenAI scrapped the public internet to train its models.
- boxedemp 2y agoDeep refers to itself as ChatGPT sometimes lol
- LZ_Khan 2y agoI actually think what DeepSeek did will slow down AI progress. What's the incentive to spend billions developing frontier models if once it's released some shady orgs in unregulated countries can just scrape your model outputs, reproduce it, and undercut you in cost? OpenAI is like a team of fodder monkeys stepping on landmines right now, with the rest of the world waiting behind them.
- zx10rse 2y agoOpenAI is already irrelevant but the audacity oh my.
- cumulative00x 2y agoThere is a saying in Turkish that roughly goes like this, it takes a thief to catch a thief. I am not a big fan of China's tech, too, however, it amuses me to watch how big tech charlatans have been crying over Deepseek shock.
- gosub100 2y agoIt's true irony to see thieves getting stolen from.
- me_me_me 2y ago[flagged]
- dns_snek 2y ago[flagged]
- deleted 2y ago[deleted]
- brianstrimp 2y agoI work in a field with lots of cheap microchips. I can tell you that the amount of counterfeit copies flooding in from China as well as the speed in which they are copying is truly breathtaking.
- me_me_me 2y agoWhat double standard? I have not claimed anything about openAI being ethical paragon of virtue. I anything, you can be accused of using whataboutism in order to justify DeepSeek illegal actions
- dns_snek 2y agoI'm not talking about OpenAI, I'm pointing out the unnecessary sideswipes directed at China whenever they do something bad that the US pioneered. It's hard to read those as anything other than casual racism.
- me_me_me 2y ago> casual racism Yeah, casual racism calling out china on stealing every ip out there as its sanctioned by their own government. Peak of racism.
- datavirtue 2y ago[flagged]
- 2y ago
- DidYaWipe 2y agoThey have "open" right in their name, so... Objection overruled.
- SubiculumCode 2y agoIf you have a set of weights A, can you derive another set of weights B that function (near) identically as A AND a) not appear to be the same weights as A when inspected superficially b) appear uncorrelated when inspecting the weight matrices?
- rahimnathwani 2y agoDo you mean for a given model structure, can two sets of weights give substantially the same outputs? Even if that were possible, it would be suspicious if you were to release an open model whose model architecture is identical to that of a closed one from a competitor. If that is what happened, we'd know about it by now.
- deleted 2y ago[deleted]
- fimdomeio 2y agoBut what is the problem here? Isn’t open AI mission “to ensure that artificial general intelligence benefits all of humanity”? Sounds like success to me.
- hugoromano 2y agoOpenAI initially scraped the web and later formed partnerships to train on licensed data. Now, they claim that DeepSeek was trained on their models. However, DeepSeek couldn't use these models for free and had to pay API fees to OpenAI. From a legal standpoint, this could be seen as a violation of the terms and conditions. While I may be mistaken, it's unclear how DeepSeek could have trained their models without compensating OpenAI. Basically, OpenAI is saying machines can't learn from their outputs as humans do.
- buildsjets 2y agoWomp Womp.
- asdfasdf1 2y agoit's no crime to steal from a thief
- krapp 2y agoIt is actually a crime to steal from a thief.
- wendyshu 2y agoIf distillation gives you a cheaper model with similar accuracy, why doesn't OpenAI distill its own models?
- mkoubaa 2y agoOpenAI made a lot of contributions to LLMs obviously but the amount of fraud, deception, and dark patterns coming out of that organization make me root against it.
- kelseydh 2y agoThe name itself, as for-profit closed source software, is grating.
- ysofunny 2y agoI see this as China fighting U.S. of A (or the American Dollar versus Chinese Renmibi if you will) and this is good because any alternatives I can think of are older-school fighting modern war is seeped in symbolism, but the contest is still there e.g. whose dong is bigger? Xi Jingping's or Dnld Trump's
- maxglute 2y agoNot that DeepSeek is luigi mangione, but it's pretty funny OpenAi getting the dead ceo treatment.
- mrkpdl 2y agoThe cat is out of the bag. This is the landscape now, r1 was made in a post-o1 world. Now other models can distill r1 and so on. I don’t buy the argument that distilling from o1 undermines deep seek’s claims around expense at all. Just as open AI used the tools ‘available to them’ to train their models (eg everyone else’ data), r1 is using today’s tools. Does open AI really have a moral or ethical high ground here?
- ijidak 2y agoPlus, it suggests OpenAI never had much of a moat. Even if they win the legal case, it means weights can be inferred and improved upon simply by using the output that is also your core value add (e.g. the very output you need to sell to the world). Their moat is about as strong as KFC's eleven herbs and spices. Maybe less...
- deleted 2y ago[deleted]
- jamil7 2y agoAgree 100%, this was also bound to happen eventually, OpenAI could have just remained more "open" from the beginning and embraced the inevitable commoditization of these models. What did delaying this buy them?
- khazhoux 2y agoWhat did delaying this cost them, though? Hurt feelings of people here who thought OpenAI personally pledged openness to them?
- jamil7 2y ago> What did delaying this cost them, though? It potentially cost the whole field in terms of innovation. For OpenAI specifically, they now need to scramble to come up with a differentiated business model that makes sense in the new landscape and can justify their valuation. OpenAI’s valuation is based on being the dominant AI company. I think you misread my comment if you think my feelings are somehow hurt here.
- yapyap 2y agoIt sounds like they’re just jealous and trying to smear shit over the wall and see what sticks. DeepSeek just bodied u bro, get back in the lab & create a better AI instead of all this news that isn’t gonna change them having a good AI
- FpUser 2y agoPot calling kettle black?
- almostdeadguy 2y agoHope Sam Altman is getting his money's worth out of that Trump campaign contribution. Glorious days to be living under the term of a new Boris Yeltsin. Pawning and strip-mining the federal apparatus to the most loyal friends and highest bidders.
- ijidak 2y agoThis whole argument by OpenAI suggests they never had much of a moat. Even if they win the legal case, it means weights can be inferred and improved upon simply by using the output that is also your core value add (e.g. the very output you need to sell to the world). Their moat is about as strong as KFC's eleven herbs and spices. Maybe less...
- sirolimus 2y agoSuch Karma lol, I wonder how they trained Sora again? You..tube something
- leobg 2y agoOpenAI is taking the position similar to that if you sell a cook book, people are not allowed to teach the recipes to their kids, or make better versions of them. That is absurd. Copyright law is designed to strike a balance between two issues. One the one hand, the creator’s personality that’s baked into the specific form of expression. And on the other hand, society’s interest in ideas being circulated, improved and combined for the common good. OpenAI built on the shoulders of almost every person that wrote text on a website, authored a book, or shared a video online. Now others build on the shoulders of OpenAI. How should the former be legal but not the latter? Can’t have it both ways, Sam. (IAAL, for what it’s worth.)
- otterley 2y agoAs another attorney, I would impart some more wisdom: "Karma's a bitch, ain't it."
- chris_wot 2y agoI quite agree. The NY Times must be feeling a lot of schadenfreude right now.
- hintymad 2y agoJust to play devil's advocate, OAI can argue that they spent great effort creating and procuring annotated data. Such datasets are indeed their secret, and now DS gets them for free by distilling OAI's output. Besides, OAI's EULA explicitly forbids users from using the output of their API for model training. I'm not saying that OAI is right, of course. Just to present OAI's point of view.
- aucisson_masque 2y agoI don't see the difference between that and LLM feeding on internet people's data. They call it IP theft yet when the New York Times sued OpenAI and Microsoft for copyright infringement they claimed it's fair use of data.
- duchenne 2y agoThe reasoning happens in the chain of thoughts. But OpenAI (aka ClosedAI) doesn't show this part when you use the o1 model, whether through the API or chat. They hide it to prevent distillation. Deepseek, though, has come up with something new.
- manamorphic 2y agoCrazy how most people miss this simple logical deduction.
- whoknowsidont 2y agoThey can claim this all they want. But DeepSeek released the paper (several actually) on what they did, and it's already been replicated in other models. It simply doesn't matter. Their methodology works.
- mkayle 2y agoThis raises the same questions I have about OpenAI: where's all this data coming from, and do they have permission to use it?
- EGreg 2y agoOkay and there is evidence OpenAI used data of many people to train its own model. Tell me again how come remixing our data is just dandy, many artists got disrupted — but no one should be able to disrupt OpenAI like that?
- dbg31415 2y agoBoo hoo? Back in college, a kid in my dorm had a huge MP3 collection. And he shared it out over the network, and people were all like, "Man, Patrick has an amazing MP3 collection!" And he spent hours and hours ripping CDs from everyone so all the music was available on our network. Then I remember another kid coming in, with a bigger hard drive, and he just copied all of Patrick's MP3 collection and added a few more to it. Then ran the whole thing through iTunes to clean up names and add album covers. It was so cool! And I remember Patrick complained, "He stole my MP3 collection!" Anyway this story sums up how I feel about Sam Altman here. He's not Metalica, he's Patrick. https://www.npr.org/2023/12/27/1221821750/new-york-times-sues-chatgpt-openai-microsoft-for-copyright-infringement https://www.npr.org/2023/12/27/1221821750/new-york-times-sue...
- kranke155 2y agoThe very idea that OAI scrapes the entire internet and ignore individual rights and thats ok, but if another company takes the output data from their model, thats a gross violation of the law / TOS - that very idea is evil.
- nbgoodall 2y agoI lol'd, from the DeepSeek news release[1]: "Pushing the boundaries of open AI!" [1]: https://api-docs.deepseek.com/news/news250120 https://api-docs.deepseek.com/news/news250120
- sgammon 2y agoThe nyt disclosure on this reporting is about to be wild
- imchillyb 2y agoIf OpenAI desires public protection, then OpenAI should open-source its models. If they did this, We the People would cover them like we do others. Without it, We the People don't care. Cry, don't cry, it's meaningless to us.
- xyst 2y agoWhat a load of shit. ClosedAI is publishing a hit piece on DeepSeek and get public and politicians on their side. Maybe even get government to do their dirty work. If they had a case, they wouldn’t be using FT. They would be filing a court case. Although that would open them up to discovery and the nasty shit ClosedAI has been up to would be game.
- ddingus 2y agoSo what? Seriously. Given how pretty much all this software was trained, who cares? I, for one, don't and believe the massive amount of knowledge continues to be of value to many users. And I find the thought of these models knowing some things they shouldn't very intriguing.
- esskay 2y agoHard to really have any sympathy for OpenAI's position when they're actively stealing content, ignoring requests to stop then spending huge amounts to get around sites running ai poisoning scripts, making it clear they'll still take your content regardless of if you consent to it.
- michaelmarkell 2y agoCan someone with more expertise help me understand what I'm looking at here? https://crt.sh/?id=10106356492 https://crt.sh/?id=10106356492 It looks like Deepseek had a subdomain called "openai-us1.deepseek.com". What is a legitimate use-case for hosting an openai proxy(?) on your subdomain like this? Not implying anything's off here, but it's interesting to me that this OpenAI entity is one of the few subdomains they have on their site
- gkbrk 2y agoCould just be an OpenAI-compatible endpoint too. A lot of LLM tools use OpenAI compatible APIs, just like a lot of Object Storage tools use S3 compatible APIs.
- jongjong 2y agoIf the material which OpenAI is trained on is itself not subject to copyright protections, then other LLMs trained on OpenAI should also not be subject to any copyright restrictions. You can't have both ways... If OpenAI wants to claim that the AI is not repeating content but 'synthesizing it' in the same was as a human student would do... Then I think the same logic should extend to DeepSeek. Now if OpenAI wants to claim that its own output is in fact copyright-protected, then it seems like it should owe royalty payments to everyone whose content was sourced upstream to build its own training set. Also, synthetic content which is derived from real content should also be factored in. TBH, this could make a strong case for taxing AI. Like some kind of fee for human knowledge and distributed as UBI. The training data played a key part in this AI innovation. As an open source coder, I know that my copyrighted code is being used by AI to help other people produce derived code and, by adapting it in this way, it's making my own code less relevant to some extent... In effect, it could be said that my code has been mixed in with the code of other open source developers and weaponized against us. It feels like it could go either way TBH but there needs to be consistency.
- deleted 2y ago[deleted]
- ripped_britches 2y agoI wish there were a stock ticker for OpenAI just to see what wall street’s take on all this is. One can imagine based on Nvidia, but I imagine OpenAI private valuation is hit much harder. Still, I think they’ll be able to justify it by building amazing products. Just interesting to watch what bankers think.
- ripped_britches 2y agoThere were definitely still very impressive engineering breakthroughs. Also it’s pretty good confirmation that synthetic data is a valid answer to the data wall problem (non-problem).
- therealpygon 2y ago“OpenAI complains company paid them for AI output that has no copyright, which was subsequently used to train another AI.” I think I fixed the title.
- alasr 2y ago> OpenAI says it has evidence DeepSeek used its model to train competitor. > The San Francisco-based ChatGPT maker told the Financial Times it had seen some evidence of “distillation”, which it suspects to be from DeepSeek. > ... > OpenAI declined to comment further or provide details of its evidence. Its terms of service state users cannot “copy” any of its services or “use output to develop models that compete with OpenAI”. OAI share the evidence with the public; or, accept the possibility that your case is not as strong as you're claiming here.
- janalsncm 2y agoAlso, there are so many innovations in their papers (Deepseek math, Deepseek v2/v3, R1) that I honestly wouldn’t even care. They figured out a way to train on only 2048 H800s when big companies are buying them in the hundreds of thousands. They created a new RL algorithm. They improved MoE. They improved the KV cache. They built an super efficient training framework.
- mercurialsolo 2y agoHow the vibe has turned on OpenAI?
- coldpepper 2y agoFuck openai. They didn't ask my peemission to crawl my blog into their dataset.
- karim79 2y agoOh God. I know exactly how this feels. A few years ago I made a bread hydration and conversion calculator for a friend, and put it up on JSFiddle. My friend, at the time, was an apprentice baker. Just weeks later, I discovered that others were pulling off similar calculations! They were making great bread with ease and not having to resort to notebooks and calculators! The horror! I can't believe that said close friend of mine would actually share those highly hydraty mathematical formulas with other humans without first requesting my consent </sarc>. Could it be, that this stuff just ends up in the dumpster of "sorry you can't patent math" or the like?
- pshirshov 2y agoA thief got robbed?..
- TylerJaacks 2y agoCry me a fucking river OpenAI, as if your business model isn't entirely based on this exact same thing.
- a2128 2y agoYeah? And if I say I have evidence OpenAI used my data to train a competitor to myself as a being that's capable of programming, will I get to have my own story on the Financial Times?
- jofzar 2y agoSorry, it's now a problem to train off other people's data? Surely openai has never trained off other people's data without permission...
- olalonde 2y agoIf it's true, how is it problematic? It seems aligned with their mission: > We will attempt to directly build safe and beneficial AGI, but will also consider our mission fulfilled if our work aids others to achieve this outcome. > We will actively cooperate with other research and policy institutions; we seek to create a global community working together to address AGI’s global challenges. https://openai.com/charter/ https://openai.com/charter/ /s, we all know what their true mission is...
- karim79 2y agoSo, banning high-powered chips to China has basically had the effect of turning them into extremophiles. I mean, that seems like a good plan </sarc>. Moreover, it is certainly slowing sales of one of the darling companies of the US (NVidia). I just can't even begin to imagine what will come of this riduculous techno-imperialism/AI arms-race, or whatever you want to call it. It should not be too hard for China to create their own ASICs which do the same, and finally be done with this palaver.
- kamranjon 2y agoI was just wondering if this is even feasible? The amount of iterations of training that would be needed for DeepSeek to actually learn anything from OpenAI would seem to be an insane amount of requests from a non-local AI, which you’d think would be immediately obvious to OpenAI just by looking at suspicious requests? Am I correct in this assumption or am I missing something? Is it even realistic that something like this is possible without a local model?
- lngnmn2 2y ago[dead]
- redder23 2y ago[dead]
- emsign 2y ago"yOu ShOuLdN't TaKe OtHeR pEoPlE's DaTa!1!1" are they mental? How can people at OpenAI lack be so self-righteous and unaware? Is thia arrogance or a mental illness?
- caseyy 2y agoSeeing as OpenAI is on the back foot, I hope nationalistic politicians don’t use this opportunity to strengthen patent laws. If one could effectively patent software inventions, this would kill many industries, from video games (that all have mechanics of other games in them) to computing in general (fast algorithms, etc). Let’s hope no one gets ideas like that… Granted, it would be ineffective in competing against China’s tech industry. But less effective laws have been lobbied through in the past.
- exabrial 2y agocry us copyright holders a river.
- rkagerer 2y agoAre they crying about their competitor training off their stuff, after having used the whole of the web to train their own stuff?
- bicepjai 2y agoReading this post, I can’t help but wonder if people realize the irony in what they’re saying. 1. “The issue is when you [take it out of the platform and] are doing it to create your own model for your own purposes,” 2. “There’s a technique in AI called distillation . . . when one model learns from another model [and] kind of sucks the knowledge out of the parent model,”
- palisade 2y agoIs this really the point OpenAI wants to start debating? When OpenAI steals everyone's data, it is fine. Right? But, let us pull the ladder up after that.
- vjerancrnjak 2y agoI thought this is capitalism for the winners. Why slander competition, just outcompete them? Why stick to your losing bets if you’ve recognized a better alternative? Let’s race to the bottom.
- animanoir 2y ago[dead]
- anon115 2y agoeat shit
- seanp2k2 2y ago"lol" said the Scorpion, "lmao".
- oatmeal_croc 2y agoEven if true, so what? These are increasingly looking like a competition between nation-states with their trade embargoes and export controls. All's fair in AI wars.
- pknerd 2y agoOpenAI steals the data from Youtube and the Internet so that's no fair either.
- MagicMoonlight 2y agoSo much for that walled garden. If rival firms can just download your entire model by talking to it then your company shouldn’t be worth billions.
- ingohelpinger 2y agoOpenAI should be quite, since they’ve scrapped the entire internet for their training data.
- otikik 2y agoChatgpt, please generate an image of the tiniest violin imaginable. Oh wait I will ask DeepSeek instead.
- pknerd 2y agoThe reason OpenAI is whining: > OpenAI’s o1 costs $60 per million output tokens; DeepSeek R1 costs $2.19. This nearly 30x difference brought the trend of falling prices to the attention of many people. From Andrew Ng's recent DeeplearningAI newsletter
- WolfOliver 2y agoI guess DeepSeek payed OpenAI for the usage of their API according to OpenAI's pricing? So what is the point if you pay for it and can not use the results how you see fit?
- glooglork 2y agoHow much data from o1 would DeepSeek actually need to actually make any improvements with it? I also assume they'd have to ask a very specific pattern of questions, is this even possible without OpenAI figuring out what's going on
- fedeb95 2y agoif some kind of transitivity holds, DeepSeek stole billions of internet users data.
- krystofee 2y agoI dont know if point of this is just to derail public attention to narative “hey, chinese stole our model, thats not fair, we need computee”, when the deepseek has clearly done some exceptional technical breakthrough on R1 and v3 models. Which even if you stole data from OpenAi is its thing.
- thih9 2y agoI don't mind and I believe that a company with "open" in its name shouldn't mind either. I hope this is actually true and OpenAI loses its close to monopoly status. Having a for profit entity safeguarding a popular resource like this sounds miserable for everyone else. At the moment AI looks like typical VC scheme: build something off someone else's work, sell it at cost at first, shove it down everyone's throats and when it's too late, hike the prices. I don't like that.
- oysmal 2y agoGiven that the training approach was open sourced, their claim can be independently verified. Huggingface is currently doing that with Open R1, so hopefully we will get a concrete answer to whether these accusations are merited or not.
- hello_computer 2y agothen show it to us rachel
- xinayder 2y ago> OpenAI declined to comment further or provide details of its evidence. Its terms of service state users cannot “copy” any of its services or “use output to develop models that compete with OpenAI”. Well, this sounds like they are just crying because they are losing the race so far. Besides, DeepSeek explicitly states they did a study on distillation on ChatGPT, then OpenAI is like "oh see guys they used our models!!!!!"
- khazhoux 2y agoBy what metric are they losing?
- xinayder 2y agoDeepSeek is a fraction of the cost of ChatGPT, they needed far few resources than OpenAI. This is essentially what caused the massive selloff in Nvidia, as a new competitor model is just as good and requires a fraction of the massive costs. I don't remember the correct metric but the cost for DeepSeek was like $15/mo while ChatGPT was $200
- khazhoux 2y agoYou said "they're losing the race." They might lose, but I don't think we're seeing that yet. They undoubtedly gained a competitor over the weekend, but that didn't change their position as the leading AI company overnight. Correct me if my understanding is wrong, but if OpenAI's accusation is correct and DS is a derivative work, then isn't it inaccurate to say DS reached ChatGPT performance "at a fraction of the cost"? If true, seems like it's more accurate to say that they were able to copy an expensive model, at low expense.
- xinayder 2y agoI agree in a way, but then in that case Gemini Claude and Qwen are all derivations of each other and shouldn't be in the competition either. DeepSeek did some studies on distillation, which might be what OpenAI is complaining about. But their bigger model is not a distilled version of OpenAI's.
- hammon 2y ago[dead]
- iimaginary 2y agoWhere did I leave my tiny violin?
- khazhoux 2y agoI'm disappointed that 99% of the comments about this topic are Schadenfreude, and 1% is actually about the technical implications of OpenAI's claims.
- janalsncm 2y agoI think readers should note that the article did not provide any evidence for OpenAI’s claims, only OpenAI declining to provide evidence, various people repeating the claim, others reacting to it. It does matter whether it happened and how much it happened. Deepseek ran head to head comparisons against O1 so it would be pretty reasonable for them to have made API calls, for example. But also, as the article notes, distillation, supervised fine tuning, and using LLM as a judge are all common techniques in research, which OpenAI knows very well.
- trkaky 2y agohow much would it cost to distill o1..
- paul_e_warner 2y agoThere seem to be two kinda incompatible things in this article: 1. R1 is a distillation o1. This is against it's terms of service and possibly some form of IP theft. 2. R1 was leveraging GPT-4 to make it's output seem more human. This is very common and most universities and startups do it and it's impossible to prevent. When you take both of these points and put them back to back, a natural answer seems to suggest itself which I'm not sure the authors intended to imply: R1 attempted to use o1 to make its answers seem more human, and as a result it accidentally picked up most of it's reasoning capabilities in the process. Is my reading totally off?
- udev4096 2y agoWhat about the pirated books you used and millions of blogs and websites scraped without consent? Somehow that's legal? Come on, give me a fucking break. OpenAI deserves the top spot in the list of unethical companies in the world
- oli5679 2y agothis is pretty ridiculous A. below is a list of OpenAI initial hires from Google. It's implausible to me that there wasn't quite significant transfer of Google IP B. google published extensively, including the famous 'attention is all you need' paper, but open-ai despite its name, has not explained the breakthroughs that enabled O1. It has also switched from a charity to a for-profit company. C. Now this company, with a group of smart, unknown machine learning engineers, presumably paid fractions of what OpenAI are published, has created a model far cheaper, and openly published the weights, many methodological insights, which will be used by OpenAI. 1. Ilya Sutskever – One of OpenAI’s co-founders and its former Chief Scientist. He previously worked at Google Brain, where he contributed to the development of deep learning models, including TensorFlow. 2. Jakub Pachocki – Formerly OpenAI’s Director of Research, he played a major role in the development of GPT-4. He had a background in AI research that overlapped with Google’s fields of interest. 3. John Schulman – Co-founder of OpenAI, he worked on reinforcement learning and helped develop Proximal Policy Optimization (PPO), a method used in training AI models. While not a direct Google hire, his work aligned with DeepMind’s research areas. 4. Jeffrey Wu – One of the key researchers involved in fine-tuning OpenAI’s models. He worked on reinforcement learning techniques similar to those developed at DeepMind. 5. Girish Sastry – Previously involved in OpenAI’s safety and alignment work, he had research experience that overlapped with Google’s AI safety initiatives.
- throwaway314155 2y ago> A. below is a list of OpenAI initial hires from Google. It's implausible to me that there wasn't quite significant transfer of Google IP I agree there's hypocrisy but in terms of making a strong argument, you can safely remove your list of persons who (drum roll)... mostly _didn't_ actually work at Google?
- dumah 2y agomy_ridiculous_list = ["Ilya Sutskever"]
- kozikow 2y agoChatgpt content is getting pasted all over the web. Now, for anyone crawling the web, it's hard to not include some chatgpt outputs. So even if you put some "watermarks" in your AI generation, it's plausible defense to find publicly posted content with those watermarks. Maybe it's explained in the article, but I can't access it, as it's paywalled.
- cratermoon 2y agoMaybe the VCs backing OpenAI invest in tiny violins.
- gejose 2y agoReminds me of this quote by Bill Gates to Steve Jobs, when Jobs accused Gates of stealing the idea for a mouse: > "Well, Steve… I think it’s more like we both had this rich neighbour named Xerox and I broke into his house to steal the TV set and found out that you had already stolen it."
- witnesser2 2y agoSoon another layer of distiller will emerge. Selling purer booze in this weight tuning buzzi.
- rochak 2y agoCry me a river
- low_tech_love 2y agoImagine having no competition one day and the next DeepSeek happens. It must’ve been quite scary. Makes sense that accusations will start flying. In my country we have a saying: a thief that robs a thief is pardoned for 100 years. It’s really interesting that the same people who defend liberal capitalism at its extreme and praise competition as its most important component (which I don’t disagree) are the same ones that’ll promptly attempt to destroy the system and the competition as soon as they are in such a position.
- sabhiram 2y agoThe grapes are sour because their moat is crumbling. What was supposed to be a model, training, and data moat - is now reduced to operational cost, which they are not terribly efficient for. OpenAI has been on a journey to burn as much $ as possible to get as far ahead on those three moats, to the point where decreasing TCO for them on inference was not even relevant - "who cares if you save me 20% of costs when I can raise on a 150b pre money value?". Well, with their moats disappearing, they will have no choice but to compete on inference cost like everyone else.
- juliuskiesian 2y agoThe obvious question is, if you have the evidence, why not just show it?
- elzbardico 2y agoI used OpenAI APIs to generate training data for some run-of-the-mill ML models at my work, for some use cases where people wanted to use LLMs directly, but that could be easily fulfilled by smaller well trained models. Is OpenAI going to complain about me too?
- elzbardico 2y agoChina is a society mostly run by engineers, some 70% of the CCP Politburo are STEM people by their formation. Engineering is a high prestige profession. The West is run by lawyers, MBAs and salesmen. This kerfuffle is a delicious study about this.
- zhenghao1 2y agoAll I see is sour grapes. Can't stand someone else coming up with a far more superior and cheaper alternative. This is business dude. There's always going to be some new disruptor to shake the market up.
- alexfromapex 2y agoThe public probably thinks that these companies are getting hacked by "sophisticated hackers" but I'd bet money that they've been hacked via social engineering.
- freejazz 2y agoWho else cares?
- animitronix 2y agoWho tf cares?
- kapad 2y agoAah. So OpenAI can use whatever means necessary to gather data for training it's model. Regardless of copyright. But somehow, it's a problem if another model developer distills it's model by training it on OpenAI? IMO, if the first use is fair, then so is the second use.