6 ms·
Why would it cast any doubt? If you can use o1 output to build a better R1. Then use R1 output to build a better X1... then a better X2.. XN, that just shows a
by riantogo 2y ago
Why would it cast any doubt? If you can use o1 output to build a better R1. Then use R1 output to build a better X1... then a better X2.. XN, that just shows a method to create better systems for a fraction of the cost from where we stand. If it was that obvious OpenAI should have themselves done. But the disruptors did it. It hindsight it might sound obvious, but that is true for all innovations. It is all good stuff.
- rockemsockem 2y agoI think the prevailing narrative ATM is that DeepSeek's own innovation was done in isolation and they surpassed OpenAI. Even though in the paper they give a lot of credit to Llama for their techniques. The idea that they used o1's outputs for their distillation further shows that models like o1 are necessary. All of this should have been clear anyway from the start, but that's the Internet for you.
- aprilthird2021 2y ago> the prevailing narrative ATM is that DeepSeek's own innovation was done in isolation and they surpassed OpenAI I did not think this, nor did I think this was what others assumed. The narrative, I thought, was that there is little point in paying OpenAI for LLM usage when a much cheaper, similar / better version can be made and used for a fraction of the cost (whether it's on the back of existing LLM research doesn't factor in)
- aiono 2y agoThat's only the case if you don't need to use the output of a much more expensive model.
- TheGRS 2y agoYes, well the narrative that rocked the stock market is different. Its looking at what DeepSeek did and assuming they may have competitive advantage in this space and could outperform OpenAI at their own game. If the narrative is actually that DeepSeek can only reach whatever heights OpenAI has already gotten to with some new tricks, then markets will probably refocus on OpenAI's innovations and price things accordingly, even if the initial cost is huge. It also means OpenAI probably needs a better moat to protect its interests. I'm not sure where the reality is exactly, but market reactions so far have basically followed that initial narrative and now the rebuttal.
- addicted 2y agoThe idea that someone can easily replicate an OpenAI model based simply on OpenAI outputs is, I’d argue, immeasurably worse for OpenAI’s valuation than the idea that someone happened to come up with a few innovations that leapfrogged OpenAI. The latter could be a one time thing, and/or OpenAi Could still use their financial might to leverage those innovations and get even better with them. However, the former destroys their business model and no amount of intelligence and innovation from OpenAI protects them from being copied at a fraction of the cost.
- aprilthird2021 2y ago> Yes, well the narrative that rocked the stock market is different. How do you know this? > If the narrative is actually that DeepSeek can only reach whatever heights OpenAI has already gotten to with some new tricks, then markets will probably refocus on OpenAI's innovations and price things accordingly Why? If every innovation OpenAI is trying to keep as secret sauce becomes commoditized quickly and cheaply, then why would markets care about any innovations they have? They will be unable to monetize them.
- davrosthedalek 2y agoCouldn't OpenAI just put in their license that training off OpenAi output is not allowed? With shibboleth or API logs, this could be verifiable.
- aprilthird2021 2y agoWhy would it matter when Chinese deepseek is not going to abide by such rules or be forced to and will release their model open weights so anyone anywhere can host it? Also, scraping most of the websites they scrape is also not allowed, they do it anyways
- davrosthedalek 2y agoIf they can make the US and Europe block the use of Deepseek and derivatives, they would be able to protect most of their market.
- kelnos 2y ago> I did not think this, nor did I think this was what others assumed. That's what I thought and assumed. This is the narrative that's been running through all the major news outlets. It didn't even occur to me that DeepSeek could have been training their models using the output of other models until reading this article.
- bigfudge 2y agoFwiw I assumed they were using o1 to train. But it doesn’t matter: the big story here is that massive compute resources are unlikely to be as important in the future as we thought. It cuts the legs off stargate etc just as it’s announced. The CCP must be highly entertained by the timeline.
- paul_e_warner 2y agoThere were different narratives for different people. When I heard about r1, my first response was to dig into their paper and it's references to figure out how they did it.
- joe_the_user 2y agoThe idea that they used o1's outputs for their distillation further shows that models like o1 are necessary. Hmm, I think the narrative of the rise of LLMs is that once the output of humans has been distilled by the model, the human isn't necessary. As far as I know, DeepSeek adds only a little to the transformers model while o1/o3 added a special "reasoning component" - if DeepSeek is as good as o1/o3, even taking data from it, then it seems the reasoning component isn't needed.
- david-gpu 2y ago> I think the narrative of the rise of LLMs is that once the output of humans has been distilled by the model Distillation is a term of art in AI and it is fundamentally incorrect to talk about distilling human-created data. Only an AI model can be distilled. https://en.m.wikipedia.org/wiki/Knowledge_distillation#Methods https://en.m.wikipedia.org/wiki/Knowledge_distillation#Metho...
- joe_the_user 2y agoMeh, It seems clear that the term can be used informally to denote the boiling down of human knowledge, indeed it was used that way before AI appeared in the popular imagination.
- david-gpu 2y agoIn the context in which you said it, it matters a lot. >> The idea that they used o1's outputs for their distillation further shows that models like o1 are necessary. > Hmm, I think the narrative of the rise of LLMs is that once the output of humans has been distilled by the model, the human isn't necessary. If deepseek was produced through the distillation (term of art) of o1, then the cost of producing deepseek is strictly higher than the cost of producing o1, and can't be avoided. Continuing this argument, if the premise is true then deepseek can't be significantly improved without first producing a very expensive hypothetical o1-next model from which to distill better knowledge. That is the argument that is being made. Please avoid shallow dismissals. Edit: just to be clear, I doubt that deepseek was produced via distillation (term of art) of o1, since that would require access to o1's weights. It may have used some of o1's outputs to fine tune the model, which still would mean that the cost of training deepseek is strictly higher than training o1.
- hmmm-i-wonder 2y ago>shows that models like o1 are necessary. But HOW they are necessary is the change. They went from building blocks to stepping stones. From a business standpoint that's very damaging to OAI and other players.
- KingOfCoders 2y agoOpenAI couldn't do it, when the high cost of training and access to GPUs is their competitive advance against startups, they can't admit that it does not exist.
- gmd63 2y agoWhy not just copy and paste the model and change the name? That's an even more efficient form of distillation.
- wgjordan 2y agoEven assuming the model was somehow publicly available in a form that could be directly copied, that would be a more blatant form of copyright infringement. Distillation launders copyrighted material in a way that OpenAI specifically has argued falls under fair use.
- Imnimo 2y agoI think it would cast doubt on the narrative "you could have trained o1 with much less compute, and r1 is proof of that", if it turned out that in order to train r1 in the first place, you had to have access to bunch of outputs from o1. In other words, you had to do the really expensive o1 training in the first place. (with the caveat that all we have right now are accusations that DeepSeek made use of OpenAI data - it might just as well turn out that DeepSeek really did work independently, and you really could have gotten o1-like performance with much less compute)
- SpaceManNabs 2y agoMy question is if deepseek r1 is just a distilled o1, i wonder if you can build a fine tuned r1 through distillation without having to fine tune o1.
- MrLeap 2y agoo1 wouldn't exist without the combined compute of every mind that led to the training data they used in the first place. How many h100 equivalents are the rolling continuum of all of human history?
- dchichkov 2y agoIt should be possible to learn to reason from scratch. And the ability to reason in a long context seems to be very general.
- Nevermark 2y agoHow does one learn reasoning from scratch? Human reasoning, as it exists today, is the result of tens of thousands of years of intuition slowly distilled down to efficient abstract concepts like "numbers", "zero", "angles", "cause", "effect", "energy", "true", "false", ... I don't know what reasoning from scratch would look like without training on examples from other reasoning beings. As human children do.
- 2y ago
- iforgot22 2y ago"Then use R1 output to build a better X1" is the part I'm not sure about. Is X1 going to actually be better than R1?
- Sophira 2y agoHonestly, it's kind of silly that this technology is in the hands of companies whose only aim is to make money, IMO.
- goatlover 2y agoIt's because they're the ones who could raise the money to make those models. Academics don't have access to that kind of compute. But the free models exist.
- lenerdenator 2y agoWell, originally, OpenAI wasn't supposed to be that kind of organization. But if you leave someone in the tech industry of SV/SF long enough, they'll start to get high on their own supply and think they're entitled to insane amounts of value, so...
- qwertox 2y agoThey're standing on the shoulders of giants, not only in terms of re-using expensive computing power almost for free by using the outputs of expensive models. It's a bit of a tradition in that country, also in manufacturing.
- unreal37 2y agoI thought OpenAI GPT took Wikipedia and the content of every book as inputs to train their models? Everyone is standing on the shoulders of giants.
- qwertox 2y agoWhat I meant to say was that OpenAI did put a lot of money into extracting value out of the pile of (partially copyrighted) data, and that DeepSeek was freeloading on that investment without disclosing it, making them look more efficient than they truly are.
- deleted 2y ago[deleted]
- bigfudge 2y agoHow do you think manufacturing in the US got started? Everyone is on someone’s shoulders.
- dontreact 2y agoIs there any evidence R1 is better than O1? It seems like if they in fact distilled then what we have found is that you can create a worse copy of the model for ~5m dollars in compute by training on its outputs.
- ospray 2y agoThey did do that themselves it's called o3.
- dartos 2y agoWhat does “better” really even mean here? Better benchmark scores can be cooked
- herodoturtle 2y agoThanks for the insightful comment. I have a question (disclaimer: reinforcement learning noob here): Is there a risk of broken telephone with this? Kinda like repeatedly compressing an already compressed image eventually leads to a fuzzy blur. If that is the case then I’m curious how this is monitored and / or mitigated.
- anothernewdude 2y agoIf they're training R1 on o1 output on the benchmarks - then I don't trust those benchmarks results for R1. It means the model is liable to be brittle, and they need to prove otherwise.
- patcon 2y agoAre we it rediscovering the evolutionary benefit of progeny (from an information theoretic lens)? And is this related to the lottery ticket hypothesis? https://arxiv.org/pdf/1803.03635.pdf https://arxiv.org/pdf/1803.03635.pdf
- indymike 2y agoBad things happen in tech when you don't do the disrupting yourself.
- RHSman2 2y agoWhen will over training happen on the melange of models at scale? And will AGI only ever be an extension of this concept? That is where artificial intelligence is going. Copy things from other things. Will there be a AI Eureka moment where it deviates and knows where and why the reason it is wrong?