8 ms·
Adding GPT4 to anything now increases marketing x4. So many AI news coming out lately that not adding it risks drowning in a sea of info.. even in the case of a
by collaborative 3y ago
Adding GPT4 to anything now increases marketing x4. So many AI news coming out lately that not adding it risks drowning in a sea of info.. even in the case of a good project
- Uehreka 3y agoThe word for this is “trademark infringement”. You are specifically not allowed to capitalize on the marketing of another entity’s product to bolster yours by implying through your name that you are somehow related. This is why “DALL-E Mini” had to change their name to craiyon.
- HarHarVeryFunny 3y agoIt's also just (deliberately) misleading. It's based on the 13B Vicuna/Llama model, not 175B GPT-3 or 1T GPT-4. There is zero justification for calling it MiniGPT-4. A more honest name would be Visual-Vicuna or Son-of-BLIP.
- sebzim4500 3y agoI don't see how it's misleading. MiniGPT-4 makes it sound like a smaller alternative to GPT-4, if it was based on GPT-4 there would be nothing 'mini' about it.
- HarHarVeryFunny 3y agoIt has more in common with GPT-3 than GPT-4 in terms of size, but in reality it's based on Vicuna/Llama which is 10x smaller than either, so as far as the LLM part of it goes its not mini-anything - it's just straight-up Vicuna 13B. The model as a whole is just BLIP-2 with a larger linear layer, and using Vicuna as the LLM. If you look at their code it's literally using the entire BLIP-2 encoder (Salesforce code). https://arxiv.org/pdf/2301.12597.pdf https://arxiv.org/pdf/2301.12597.pdf
- fire 3y agovicuna was done with sharegpt transcripts, right? did they ever say if those transcripts were while users were using gpt3.5 or gpt4.0?
- HarHarVeryFunny 3y agoI haven't read the details of how they created the training data.
- deleted 3y ago[deleted]
- Tepix 3y ago> 1T GPT-4 The number of parameters used for GPT-4 is unknown.
- HarHarVeryFunny 3y agoI got the 1T GPT-4 number from here - this is the video that goes with the Microsoft "Sparks of AGI" paper, by a Microsoft researcher that had early access to GPT-4 as part of their relationship with OpenAI. https://www.youtube.com/watch?v=qbIk7-JPB2c https://www.youtube.com/watch?v=qbIk7-JPB2c
- sandkoan 3y agoBubeck has clarified that the "1 trillion" number he was throwing around was just a hypothetical metaphorical—it was in no way shape or form implying that GPT-4 has 1 trillion parameters [0]. [0] https://twitter.com/SebastienBubeck/status/1644151579723825154 https://twitter.com/SebastienBubeck/status/16441515797238251...
- HarHarVeryFunny 3y agoOK - thanks! So we're back to guessing ... A couple of years ago Altman claimed that GPT-4 wouldn't be much bigger than GPT-3 although it would use a lot more compute. https://news.knowledia.com/US/en/articles/sam-altman-q-and-a-gpt-and-agi-lesswrong-aa6293dde7b95b537f5f50e37c861645fd9b4dbb https://news.knowledia.com/US/en/articles/sam-altman-q-and-a... OTOH, given the massive performance gains scaling from GPT-2 to GPT-3, it's hard to imagine them not wanting to increase the parameter count at least by a factor of 2, even if they were expecting most of the performance gain to come from elsewhere (context size, number of training tokens, data quality). So in 0.5-1T range, perhaps ?
- HarHarVeryFunny 3y agoFWIW, Stephen Gou, Manager of ML at Cohere, is currently doing a Reddit AMA, and is also guessing at 1T params for GPT-4. https://www.reddit.com/r/IAmA/comments/12rvede/im_stephen_gou_manager_of_ml_founding_engineer_at/ https://www.reddit.com/r/IAmA/comments/12rvede/im_stephen_go...
- EntrePrescott 3y ago> Son-of-BLIP maybe even add an "a" for extra spice: Son-of-a-BLIP
- collaborative 3y agoAt this point the letters GPT make more sense than "AI" or "LLM" in many peoples minds
- Uehreka 3y agoHard disagree. Outside of the brand name ChatGPT, lay members of the general public are way more likely to call these chatbots (like Bard and Bing) “AIs” than “GPTs”. And although GPT could technically refer to any model that uses a Generative Pre-trained Transformer approach (although it probably wouldn’t be an open-and-shut case), the mark “GPT-4” definitely is associated with OpenAI and their product, and you can’t just use it without their permission.
- collaborative 3y agoSo OpenAI ostensibly owns "GPT4" according to your argument. But does it own "MiniGPT4"? I hope you see the absurdity of this. Let's not discuss the amount of copyright licenses OpenAI has already infringed, too
- Uehreka 3y agoI’ll put it this way: At Brewer’s Art in Baltimore, MD they just released a beer called GPT (Green Peppercorn Tripel)[1]. They’re likely allowed to do that because a reasonable consumer would probably not actually think they had collaborated with OpenAI, because OpenAI does not make beer. OP is releasing a model called “MiniGPT-4”. A reasonable consumer could look at that name and become confused about the origin of the product, thinking it was from OpenAI. This would be understandable, since OpenAI also makes large language models and has a well known one that they’ve been promoting whose brand name is “GPT-4”. If MiniGPT-4 does not meet that consumer’s expectation of quality (which has been built up through using and hearing about GPT-4) it may cause them to think something like “Wow, I guess OpenAI is going downhill”. Trademark cases are generally decided on a “reasonable consumer” basis. So yeah, they can seem a little arbitrary. But it’s important for consumers to be able to distinguish the origin of the goods they are consuming and for creators to be able to benefit from their investment in advertising and product development. [1] https://www.thebrewersart.com/bottles-cans https://www.thebrewersart.com/bottles-cans
- nashashmi 3y agoThey can always say GPT-like. Or miniaturized GPT-like LLM.