6 ms·
Meta to release open-source commercial AI model
- Jeff_Brown 3y agoWhat's the monetization model here? Is this a closed-source version of their open-source model? (That's suggested by the phrase in the article, "a commercial version of LLaMA, its open-source large language model".)
- pueblito 3y agoI think the strategy is more to prevent competitors from monetizing
- fullshark 3y agoThat's a huge reason to do it also, but it also makes sense if you have researchers + developers improving the engine of something that powers your product. The moat / competitive advantage at FB is their network, not so much the proprietary underlying tech.
- mtillman 3y agoPeople often say this but having interviewed ~200 facebook engineers over the years, their scaling tech around both software and hardware is pretty impressive.
- fullshark 3y agoYeah I guess it's a competitive advantage when a competitor (Twitter) is showing to have technical problems operating at global scale with a smaller team. Their scale is not trivial by any means. But people aren't going to go to FB because they have the best LLM, makes sense to offload that development to the open source community.
- dpflan 3y agoYes, commoditize the competition.
- treprinum 3y agoYou still need to build real-time serving infrastructure on top of LLaMA/Vicuna/Alpaca in order to compete with ChatGPT/OpenAI so it's not going to be done by that many companies and OpenAI already has a mindshare/first mover advantage.
- staticman2 3y agoWhen you use ChatGPT you are leasing their GPU infrastructure and their proprietary model, this opens the possibility of leasing GPU infrastructure from another company and using an open source model. You don't necessarily need to do the hard parts yourself, you can hire it out to competing companies.
- treprinum 3y agoSure, but it's extra work slowing you down as your competitor is surfing the wave at full speed. Moreover, you are relying on an old LLM whereas OpenAI is developing newer versions of theirs, keeping their competitive advantage. Even Google who has the infra has a ridiculously bad LLM to compete.
- CharlesW 3y ago> What's the monetization model here? There doesn't necessarily have to be one. Facebook's goal may be to help commoditize its complements. https://gwern.net/complement https://gwern.net/complement
- RobotToaster 3y agoMaybe they just want the de facto standard LLM to be one that only says nice things about facebook and Zuckerberg?
- deleted 3y ago[deleted]
- elorant 3y agoYou could pay to customize it and/or retrain it for your use case. Or you pay a subscription and every few months you receive updated weights.
- zpeti 3y agoFree AI models mean more free content, which is exactly what drives facebooks moat.
- dpflan 3y agoI don’t know, maybe they don’t need to monetize their model? I don’t know if they have to, they need their models to be the best and to support their core business of ads, anything that keeps users on their platform for any reason is their goal. They need their models to be an industry standard and one upon which other things are built.
- valine 3y agoLLaMA isn’t licensed for commercial use. It’s probably an update to the licensing. Facebook benefits heavily from the open source development done on LLaMA. There was a report I saw that facebook has started using llama.cpp internally for inference. Updates to the licensing will cement facebook as the go to choice for open source language models.
- fnands 3y agoThere is an open source version available, with weights that were leaked, but licensed as "for academic purposes only". This seems they will release the weights under some license that allows commercial usage. How they monetise it (which I assume they will try and do?) is an interesting question. Maybe some variant of paying a licencing fee?
- thatguymike 3y agoCommoditize Your Complement, it’s a strategic play since Meta is behind on LLMs.
- Roark66 3y agoWell, if they really released it as open source, I guess depending on the exact license a company that modifies(fine tunes) it and wants to make money on that modified version would have to distribute the weights and/or disclose the details about how they fine tuned it. On what data etc. By offering a commercial license , the buyer can do anything they want.
- dundun 3y agoGoogle opensourced Tensorflow because they believed it would help with the hiring process: if researchers could use the same framework to do their PhDs as Google used in their production systems, that was seen as an advantage. Maybe that's Meta's play here? Maybe the idea is that the ecosystem around a model could be as valuable or more valuable than the model itself too, so an OSS model could benefit Meta a lot more by gaining more of the ecosystem mind share? Or Maybe Yann LeCun is just a hippie that dreams of free love, hard drugs and open-source models?
- andy99 3y agoLLaMA is already in it's way to becoming an industry standard (in my opinion, look at llama.cpp plus everything build on LLaMA). There are benefits to being able to set direction like that. Same as pytorch for example, it's not just about direct revenue it's about everyone building on and contributing to your platform. They might have done well to make gg an offer he couldn't refuse and take on ggml and llama.cpp as an open source project.
- TechnicolorByte 3y agoLike others said it’s probably to commoditize their competition. The models don’t matter so much as ownership of the platform and critical data. Which is why OpenAI is in a tricky position (although I guess they’re partnered with Microsoft). It seems like the existing large platforms of today—Microsoft’s enterprise moat, Google’s ads and internet services, Meta’s social networks, Apple’s consumer and mobile products—will remain the primary platforms of the future. So having models that can operate exclusively on those platforms via integration to their key products and date will only continue this trend. If you’re an outsider with an AI model, you’ll have a harder time getting access to critical data and your standalone AI product (e.g., ChatGPT) won’t be as useful. More broadly speaking, I believe the days where the top X largest companies in the stock company would be displaced by newer companies every decade or so is over. The FAANGs just control so many major platforms in so many aspects of our lives.
- isaacremuant 3y ago> More broadly speaking, I believe the days where the top X largest companies in the stock company would be displaced by newer companies every decade or so is over. The FAANGs just control so many major platforms in so many aspects of our lives. It also helps that they buy or otherwise cooperate to destroy their competition in questionable ways while heavily lobbying the gov to favor them over others in a quid-pro-quo that benefits politicians and not their constituents.
- sangnoir 3y ago> More broadly speaking, I believe the days where the top X largest companies in the stock company would be displaced by newer companies every decade or so is over. I disagree: I think big tech is hard to disrupt ATM because the companies are still young and nimble. In the last cycle, the companies being displaced were ancient (by tech standards). When Google and Facebook are 30 years old, their DNA will get in the way of adopting to a new paradigm that will change the world. A paradigm that may be to the Metaverse what the smartphone was to the Apple Newton
- jerrygenser 3y agoBased on the podcast with Lex Friedman and Mark Zuckerberg, see ~minute 30. My hypothesis based on the context of Mark discussing the release is that it's going to be completely open source and can licensed to be used commercially. Not that Meta is going to add a whole new revenue side of business to compete with OpenAI. i.e. "Here is model, with commercially permissive licensing" not "Here is model that you can use commercially but must pay me" https://www.youtube.com/watch?v=Ff4fRgnuFgQ&ab_channel=LexFridman https://www.youtube.com/watch?v=Ff4fRgnuFgQ&ab_channel=LexFr...
- hospitalJail 3y agoAnother hypothesis is that they are trying to rehab their brand. They can even write it as 'good will' on their financial statements. It kind of is working.
- justapassenger 3y agoMeta has been one of the major open source contributors for about a decade now. They open source/contribute to a lot of tech, as their business isn’t about tech, but products.
- smoldesu 3y agoThis isn't some recent revelation or anything. Facebook's AI team (FAIR) open sourced their major technology in 2017 with Pytorch. In 2018 they published Pytext in an age when most people didn't know what a Large Language Model even meant. Seeing LLaMA get made should not be a surprise to anyone who is familiar with the history of AI research. It's like hearing people call CUDA an "unfair advantage" while ignoring billions of Nvidia R&D dollars getting spent in the AI sector over the course of a decade. It might feel like "brand rehab" or "good will" as a consumer, but a lot of this work was put in motion a while ago.
- discmonkey 3y agoMeta is a company that makes money off of users endlessly browsing content. It would follow that making it easier/faster to generate content would benefit Meta.
- ekojs 3y agoSeems that the source is a FT article that was discussed yesterday: https://news.ycombinator.com/item?id=36712168 https://news.ycombinator.com/item?id=36712168 From the FT article: '“The goal is to diminish the current dominance of OpenAI,” said one person with knowledge of high-level strategy at Meta.'
- zargon 3y agoIt's not open source, it's freeware or something like that. Weights aren't the source code of LLMs, they're the binaries.
- whimsicalism 3y agoStrong disagree - I think OSS is fine framing of this. Weights are a third category, you can 'fork' them in an a way that you can't with standard binaries.
- l33t233372 3y agoYou can add hooks to functions and “fork” binaries, which is a pretty similar effort to adding training data to given model weights.
- IshKebab 3y agoNobody does that because if you only have binaries you probably don't have permission to do that. Plus it's impractical to make any significant changes that way.
- l33t233372 3y agoIf you have binaries you almost always have “permission” to do that — you can do whatever you’d like with files on your own system.
- deleted 3y ago[deleted]
- hoofedear 3y agoThank you for succinctly explaining the difference, I learned something today
- StackOverlord 3y agoCompiling source code doesn't cost million of dollars though
- 40yearoldman 3y agoIs the title an oxymoron? Open-source commercial?
- satvikpendem 3y agoNo, you can sell open source software commercially. That being said, I'm wondering if the license will truly be open source or more like Stable Diffusion's license which is not really open source.
- schleck8 3y agoBecause deep learning weights aren't source code. https://huggingface.co/blog/open_rail https://huggingface.co/blog/open_rail
- gpm 3y agoCommercial presumably as opposed to non-commercial licensing (e.g. the CC BY-NC license, or the weird situation LLaMa is in). If you listen to the definition the Open Source Initiative would have applied to the term open source had they succeeded in acquiring rights to the term, then commercial is redundant with open source, not the opposite of it.
- RobotToaster 3y agoIf anything it's a tautology, open source by definition allows commercial use.
- isaacremuant 3y agoI think you could've googled that one and founds years of knowledge on that one. Free as in beer Vs free as in speech and the whole thing.
- greatpostman 3y agoMeta is going to ruin open ais moat on purpose. Great business strategy and good for everyone but metas competitors
- deleted 3y ago[deleted]
- jonnat 3y agoQuite the opposite, this is great for Meta's competitors. Meta is not trying to get market share with this strategy, it's trying to commoditize their complements (https://www.joelonsoftware.com/2002/06/12/strategy-letter-v/ https://www.joelonsoftware.com/2002/06/12/strategy-letter-v/) Content is a complement to a social network: the cheaper it is to create content, the more content is available, the easier it is to optimize a feed, the larger the time people spend in the platform, the higher the revenue. GenAI is just a method to drive the cost of content creation to zero.
- strikelaserclaw 3y agokind of a dystopian nightmare world in which large corporations utilize AI to create low cost, infinite content that humans engage with (mostly content catering to the human tendency for tribalism, prestige, sexual desires etc...), sounds like we are creating a world similar to the Matrix.
- roody15 3y agoI think we may have already entered it. Infinite scroll based feeds like TikTok, Instagram, and Threads (and possible Reddit these days) … just AI algorithm deciding what you should find “entertaining” or “important”. It’s really the ultimate nightmare with the internet becoming just TV 3.0 in which content is controlled and curated … you just consume mindlessly. Any attempts to create a Reddit clone.. or system in which people freely communicate is now “regulated” for “hate” speech or “terrorism”. The days of open discourse … appear to be numbered. Even email will be analyzed by AI to look for “trends” or “optimize” employee efficiency. It really is time for a new internet.
- obblekk 3y agoMaybe they've solved the fingerprinting problem and can identify text generated from their model, and this is a way of discovering the market they can sell more advanced models to directly. B2B leadgen...
- sva_ 3y agoI mean you could probably just train it on some sequence s.t. the model identifies itself, would be hard to detect that
- sebzim4500 3y agoThat would prbably work to detect if e.g. OpenAI or Anthropic start using their weights directly. It wouldn't detect whether e.g. a blog was generated with their model or not.
- vlovich123 3y agoI don't think so because I believe you can train AI models against other AI models. I believe you can fingerprint a family of models, but that's not going to tell whether you just used the general approach outlined in the academic papers.
- foob 3y agoFrom the recent story about the Sarah Silverman lawsuit: The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private tracker.” Bibliotik and the other “shadow libraries” listed, says the lawsuit, are “flagrantly illegal.” IANAL, but this basically sounds like LLaMa was trained on illegally obtained books by Meta's own admission. It's an exciting development that Meta is releasing a commercial-use version of the model, but I wonder if this is going to cause issues down the road. It's not like Meta can remove these books from the training set without retraining from scratch (or at least the last checkpoint before they were used). [1] https://news.ycombinator.com/item?id=36657540 https://news.ycombinator.com/item?id=36657540
- _Algernon_ 3y agoHow is it different than training from random blogs, or stack overflow or in general "The Internet"?
- schleck8 3y agoReally, really bad look for Eleuther if this is true. I did not expect them do something like this and not even see the issue with it.
- 0cf8612b2e1e 3y agoFrom my quick skim I could not find a date. Any idea when this might happen?
- DueDilligence 3y ago[dead]
- forgingahead 3y agoZuck is a total killer. What better way to fight Google and Microsoft than to effectively spawn thousands of potential competitors to their AI businesses with this (and other) releases. There will be a mad scramble over the released weights to develop new tech, these startups will raise tons of money, and then fight the larger incumbents. This is not charity, this is a shrewd business move.
- amelius 3y ago"Commoditize your complement"
- stale2002 3y agoI remember arguing with people who honest to god thought that LLAMA was some sort of secret ploy, to trick startups into using it, so that meta could sue them for using it commercially. Well now there is a commerical release. I guess it wasn't some corporate plot after all! Some people just can't admit when a corporation does a good thing. (In this case, the good thing is being done to obsolete their competitors, but it is good none the less, that a commerical LLM is available for people to use for free)
- pmarreck 3y agoI have a 128 core Threadripper, a 2080 Ti and a 3080 Ti. How can I play with open source LLM's locally?
- brucethemoose2 3y agoKobold.cpp is your best bet. You can leverage those big CPUs while still loading both GPUs with a 65B model. ... If you are feeling extra nice, you should set that up as an AI horde worker whenever you run koboldcpp to play with models. It will run API requests for others in the background whenever its not crunching your own requests, in return allowing you priority access to models other hosts are running: https://aihorde.net/ https://aihorde.net/
- pmarreck 3y agooooh, this is a great idea
- brucethemoose2 3y agoAlso, I would suggest this model as one to play with: https://huggingface.co/ycros/airoboros-65b-gpt4-1.4.1-PI-8192-GGML https://huggingface.co/ycros/airoboros-65b-gpt4-1.4.1-PI-819... Check the prompting syntax here, it has a huge effect on the output: https://huggingface.co/jondurbin/airoboros-65b-gpt4-1.4 https://huggingface.co/jondurbin/airoboros-65b-gpt4-1.4
- estreeper 3y agoIf you're just looking to play with something locally for the first time, this is the simplest project I've found and has a simple web UI: https://github.com/cocktailpeanut/dalai https://github.com/cocktailpeanut/dalai It works for 7B/13B/30B/65B LLaMA and Alpaca (fine-tuned LLaMA which definitely works better). The smaller models at least should run on pretty much any computer.
- brucethemoose2 3y agoThat project seems unmaintained, which is a problem because llama.cpp is changing extremely rapidly. Also, it has no "1 click" exe release like kobold.
- rvz 3y agoSee. They don't care about the LLaMA model leak. It turns out that it was OpenAI that cares because it ruins their moat. It costs Meta nothing to release a better open-source or freely available version of LLaMA again. Still waiting for the 'Meta is dying' and 'Fire Mark Zuckerberg' calls from last year. A year later, where are they now?
- sebzim4500 3y agoTo be fair to those commenters in the past, I don't think anyone could have forseen that Zuckerberg would turn out to be the "the good guy".
- whimsicalism 3y agoIf you read past the title, this article is not at all clear if they are referring to a commercial offering (ie. license our model for $$) or an open-source license with commercial usage (Apache, etc.) My guess is still the latter because that's what I've heard the rumors about, but this article is pretty unclear on this fact.
- brucethemoose2 3y agoFalcon 40B was released as a "free with royalties above a certain amount of profit" license, and got roasted for it. It was so bad that they changed the license to Apache. I don't think any business would run such a "licensed" model over MPT 30B or Falcon 40B, unless its way better than LLaMA 65b.
- whimsicalism 3y agoI think it is supposed to be better than LLaMA 65B. Plenty of businesses are paying for OAI API access.
- loufe 3y agoI'm surprised nobody here has brought up the sensorship in this model. Listening to Mark Zuckerberg on Lex Friedman's podcast talk about it, it sounds like the model will be significantly blunted vs its "research" version release.
- sagebird 3y agorepeat after me: hardware is the only moat If you want to live the good life before you are exquisitely extinguished, spend every other day figuring out how to buy more NVDA, the other days exercising outside, being human.
- sifar 3y agountil better algorithms or newer paradigm obviate the need for large memory/computations
- bilsbie 3y agoIs it possible to do further training on the weights they release?
- brucethemoose2 3y agoYes, and there are a sea of finetunes. See: https://huggingface.co/models?sort=modified&search=Ggml https://huggingface.co/models?sort=modified&search=Ggml QLORA is the most cost effective method so far. Some people also do finetuning on Google TPUs
- TheBengaluruGuy 3y agoThis conversation triggers a thought. Does it mean that any blogs that I wrote from my own insights, will automatically be trained on the model… without my permission? As an author, it feels like it’s stealing the knowledge and insight without appropriate attribution.
- anaganisk 3y agoI think we are at a looking where we just have to let go unless we are Disney, with an army of lawyers. May be it's time for the change in thinking. Having said that. Attribution allows a person to trace the source, it's not a success marker anymore. Probably, if enough negative statements generated by AI get popular, that could potentially piss of countries/people for example some LLM recognizing Taiwan as independent country you can bet China will push for attribution to sources. We have bills pending in multiple countries that want access to personal of encrypted messages to trace the source.