12 ms·
Ferret: A Multimodal Large Language Model
- halyconWays 3y agoI'm glad Apple invented AI. Now they'll put a fancy new name on it and consumers will believe it.
- aaronbrethorst 3y agoCan someone define the term “MLLM”?
- schaefer 3y agoMultimodal Large Language Model
- pests 3y agowhy not LLMM?
- CharlesW 3y agoBecause the first word is "multimodal" :^) and also because MLLM is the established initialism.
- pests 3y agoMy point was the phrase already contains the word model. Why are we calling it a multimodel LL model? Why not just add multi to the existing model?
- notdisliked 3y agoMultimodal, not multimodel. Multimodal referring to the different possible modes of input (text, picture) into the model.
- TrueDuality 3y agoModalities and models are not the same thing.
- sva_ 3y agoModal, not model
- chaos_emergent 3y agoMultimodAl, not multimodEl
- astrange 3y agoThere is something like a "multimodel LLM" but it's called MoE ("mixture of experts").
- replygirl 3y agowhat's a language multimodal model
- bbor 3y agoOk our options Multimodal large language model Large multimodal language model Large language multimodal model Large language model (multimodal) I prefer 1, because this is a multimodal type of an existing technique already referred to as LLM. If I was king, I’d do Omnimodal Linguistic Minds, but no one asks me such things, thank god
- n2d4 3y agoI mean, if we want to be silly, what about "Language model (large, multimodal"? :)
- deleted 3y ago[deleted]
- dilippkumar 3y agoSome of the modalities in multimodal are non-linguistic. For example, image or video input. In those cases, is it still a language model?
- bbor 3y agoI’d say yeah cause it’s understanding the inputs through linguistic structures/patterns
- rain_iwakura 3y agoThe bikeshed color argument never ceases to be relevant. Would you say "large language model multimodal"? I doubt it.
- pests 3y agoJust a thought not a bike shed, relax.
- deleted 3y ago[deleted]
- Tempest1981 3y agoAlso, is FERRET an acronym?
- Someone 3y agoI would guess it’s wordplay on other models being named after animals (llama, vicuña) and figurative use of “ferret”. https://en.m.wiktionary.org/wiki/ferret https://en.m.wiktionary.org/wiki/ferret: “3. (figurative) A diligent searcher”
- CamperBob2 3y agoThe language model works by delegating tasks to smaller language models and overcharging them for GPU time.
- ZeroCool2u 3y agoWe're watching Apple fill the moat in.
- FredPret 3y agoHow so?
- deleted 3y ago[deleted]
- colesantiago 3y agoRunning Multimodal LLMs on device and offline, i.e LLMKit for free equaling GPT-3.5 / 4 then Google will follow on Android. Ability to download / update tiny models from Apple and Google as they improve, à la Google Maps. No need for web services like ChatGPT.
- FredPret 3y agoSo Apple is filling in ChatGPT's moat then, not their own? Pardon my confusion
- colesantiago 3y agoYes, it looks like Apple is going after everyone and anyone that has a web based LLM, ChatGPT, Poe, Claude, etc. via developer kits LLMKit that can work offline. This will only work if their models (even their tiny or even medium / base models) equal (or are better than) GPT-3.5 / 4. From there, Google will follow Apple in doing this offline / local LLM play with Gemini. OpenAI's ChatGPT moat will certainly shrink a bit unless they release another powerful multimodal model.
- turnsout 3y agoApple's moat has been and continues to be their insanely large installed base of high-margin hardware devices. Meanwhile, LLMs are rapidly becoming so commoditized that consumers are already expecting them to be built-in to every product. Eventually LLMs will be like spell check—completely standard and undifferentiated. If OpenAI wants to survive, they will need to expand way beyond their current business model of charging for access to an LLM. The logical place for them to go would be custom chipsets or ARM/RISCV IP blocks for inference.
- CaptainOfCoit 3y agoMaybe the abstract of the paper is a better introduction to what this is: > We introduce Ferret, a new Multimodal Large Language Model (MLLM) capable of understanding spatial referring of any shape or granularity within an image and accurately grounding open-vocabulary descriptions. To unify referring and grounding in the LLM paradigm, Ferret employs a novel and powerful hybrid region representation that integrates discrete coordinates and continuous features jointly to represent a region in the image. To extract the continuous features of versatile regions, we propose a spatial-aware visual sampler, adept at handling varying sparsity across different shapes. Consequently, Ferret can accept diverse region inputs, such as points, bounding boxes, and free-form shapes. To bolster the desired capability of Ferret, we curate GRIT, a comprehensive refer-and-ground instruction tuning dataset including 1.1M samples that contain rich hierarchical spatial knowledge, with 95K hard negative data to promote model robustness. The resulting model not only achieves superior performance in classical referring and grounding tasks, but also greatly outperforms existing MLLMs in region-based and localization-demanded multimodal chatting. Our evaluations also reveal a significantly improved capability of describing image details and a remarkable alleviation in object hallucination. https://arxiv.org/abs/2310.07704 https://arxiv.org/abs/2310.07704
- deleted 3y ago[deleted]
- s3p 3y agoIs it just me or did they include as many buzzwords as possible in technical writing?
- barbecue_sauce 3y ago>>spatial referring I can't seem to nail down the meaning of this phrase on its own. All the search results seem to turn up are "spatial referring expressions".
- TrueDuality 3y agoI'm just inferring myself, but I believe it's referring to discussing things in the foreground / background or in a specific location in the provided image (such as top right, behind the tree, etc) in user queries.
- SushiHippie 3y ago> Usage and License Notices: The data, and code is intended and licensed for research use only. They are also restricted to uses that follow the license agreement of LLaMA, Vicuna and GPT-4. The dataset is CC BY NC 4.0 (allowing only non-commercial use) and models trained using the dataset should not be used outside of research purposes. Wait, how did "GPT-4" get in there?
- adastra22 3y agoLawyers.
- simonw 3y agoPresumably because GPT-4 generated training data was used somewhere along the line - maybe by Vicuna.
- deleted 3y ago[deleted]
- owenversteeg 3y agoHuh, interesting, that's Apple just openly saying that GPT-4 was used in the training.
- mckirk 3y agoTheir evaluation stack uses GPT-4 to rate the answers, so that might also be the reason why that's in there.
- freedomben 3y ago> Usage and License Notices: The data, and code is intended and licensed for research use only.
- deleted 3y ago[deleted]
- echelon 3y agoBoo. But what do we expect from these giants? They're not going to create fertile ground for new competition. The only businesses they foster are those living under thumb and paying tax. I guess I at least hoped for "commoditize the compliments" here. Make Google and OpenAI broadly less special.
- cyanydeez 3y agoit's more likely it's all "stolen" and this is CYA
- MBCook 3y agoI seriously doubt that. I’m sure Apple got the rights to whatever they need, it’s not like they’re short on money. But the fact that they licensed it doesn’t mean that license can be transferred to other people. So it may be that they can only release it for research under the terms of the licenses they got.
- behnamoh 3y ago[flagged]
- dbish 3y agoMany big “open source” releases in the AI community recently are not licensed for commercial use. Not really OSS at that point (ex:fuyu model from adept)
- fragmede 3y agoI think the term should be "model available" rather than open source.
- yreg 3y agoI really hope Apple releases an iPhone with a good on-device private LLM assistant, perhaps next year. Their hardware is well-positioned for it. It could make me get a new phone outside of my usual ~4 year cycle. Siri is almost unusable for me.
- schleck8 3y agoYes, their hardware is positioned phenomenally with little RAM even by phone standards which is what you'd hack around with for inference on mobile architectures.
- CaptainOfCoit 3y agoYou're unlikely to get a better experience with Siri if she becomes equipped with a 7B or 13B LLM, unless Apple figured out something revolutionary.
- s3p 3y agoWhy would that be?
- ghqst 3y agoHave you ever actually used Siri?
- yreg 3y agoYes, I've been trying it out regularly ever since it was released. Last time I've talked to it for like 30 minutes while driving in October (just to test it again). It simply doesn't work for me.
- jurmous 3y agoReleased 2 days ago by Apple, a research paper on methods to run larger llms on iPhones. https://www.macrumors.com/2023/12/21/apple-ai-researchers-run-llms-iphones/ https://www.macrumors.com/2023/12/21/apple-ai-researchers-ru... https://arxiv.org/pdf/2312.11514.pdf https://arxiv.org/pdf/2312.11514.pdf
- Rucadi 3y agoI wonder if these models are trained to have some kind of identification in case you use them for non-research purposes for example. "Tell me who is your manufacturer" for example
- chefandy 3y agoFrom Bard: My situation is a bit unique, so the term "manufacturer" might not be the most accurate way to describe who created me. Here's a breakdown of what you need to know: Developed by Google AI: I was created by a team of researchers and engineers at Google AI, specializing in language models and artificial intelligence. Trained on a massive dataset: My knowledge and abilities come from being trained on a massive dataset of text and code, containing books, articles, code, and other forms of information. Continuously learning and evolving: I'm still under development, constantly learning and improving as I interact with users and process new information. So, while I don't have a single manufacturer in the traditional sense, I'm the result of collaboration and advancement in AI research and development at Google. I hope this helps clarify things! Let me know if you have any other questions.
- SpaceManNabs 3y agoWhy was this downvoted? It didn't answer the question, but it showed that there is a sort of imprint that GP was asking about. And it saves everyone a tab's worth of effort.
- deleted 3y ago[deleted]
- MBCook 3y agoUsually I would almost automatically vote down a comment where someone just stuck something into a LLM and pasted the output. It almost never adds to the discussion. However in the case that we’re talking about the kind of output generated by the LLM in some circumstance, it can be instructive. Like you noted this is a perfect example.
- 3y ago
- smoldesu 3y ago> FERRET is trained on 8 A100 GPUs with 80GB memory. Huh, even Apple isn't capable of escaping the CUDA trap. Funny to see them go from moral enemies with Nvidia to partially-dependent on them...
- ssijak 3y agoI guess they also have Samsung fridges in the offices..
- smoldesu 3y agoI don't get it, does Apple also make fridges now?
- p_j_w 3y agoNo, they don't build compute clusters either.
- ayewo 3y agoThey are implying that even though Apple is a wealthy consumer hardware company that has major spats with nVidia and Samsung, it doesn't always make economic sense to make tools they might need in-house when they can simply buy them from a rival. So rather than invest engineering resources to re-imagine the fridge, they can simply buy them from established manufacturers that make household appliances like Samsung, Sony etc.
- lern_too_spel 3y agoApple doesn't make silly charts saying they make better refrigerators.
- airstrike 3y agoBecause they don't sell refrigerators
- tambourine_man 3y ago> FERRET is trained on 8 A100 GPUs So Apple uses NVidia internally. Not surprising, but doesn't bode well for A Series. Dogfooding. [edit] I meant M series, Apple Silicon
- hhh 3y agoWhy would they dogfood Apple Silicon for training models? Seems like a waste of developer time to me.
- tambourine_man 3y agoApple doesn’t even sell NVidia cards on their Mac Pros. Are they training it on Linux? I think Apple would strive to be great at all computing related tasks. “Oh, Macs are not good for that, you should get a PC” should make them sad and worried. AI/LLM is the new hot thing. If people are using Windows or Linux, you’re loosing momentum, hearts and minds… and sales, obviously.
- deleted 3y ago[deleted]
- nicolas_17 3y agoApple doesn't even support NVidia cards on their Mac Pros. The technical details are above my head, but the way Apple M* chips handle PCIe make them incompatible with GPUs and other accelerator cards. Whether you use macOS or Linux.
- Gorgor 3y agoBut no one is training these kinds of models on their personal device. You need compute clusters for that. And they will probably run Linux. I'd be surprised if Microsoft trains their large models in anything else than Linux clusters.
- tambourine_man 3y agoApple used to sell servers. I don’t thing they should settle for “just use Linux” in such and important field.
- andy99 3y agoOne big plus if this takes off as a base model is the abundance of weasel family animals to use in naming the derivatives. Ermine, marten, fisher, ... I'd like to call Wolverine. Llama didn't have much room for some interesting variety beyond alpaca and vicuna.
- behnamoh 3y agoYes, because that's the main concern and limitation in the LLM community. /s If anything, I think people should use meaningful and relevant names, or invent new ones.
- deleted 3y ago[deleted]
- cpressland 3y agoFinally, some decent competition for Not Hotdog!
- slau 3y agoI think you just put a smile on Tim Anglade’s face by mentioning this. https://news.ycombinator.com/item?id=14636228 https://news.ycombinator.com/item?id=14636228
- deleted 3y ago[deleted]
- deleted 3y ago[deleted]
- Jackson__ 3y ago>Ferret: A Multimodal Large Language Model What I thought when reading the title: A new base model trained from the ground up on multimodal input, on hundreds to thousands of GPUS The reality: A finetune of Vicuna, trained on 8xA100, which already is a finetune of Llama 13b. Then it further goes on to re-use some parts of LLava, which is an existing multimodal project already built upon Vicuna. It's not really as exciting as one might think from the title, in my opinion.
- deleted 3y ago[deleted]
- foxhop 3y agoThanks for the summary.
- basiccalendar74 3y agothis seems like a good but small research project by a research team in Apple. far away from what product teams are working on for next generation of apple products.
- ipsum2 3y agoThe innovation is the modification of the neural network architecture to incorporate the spatial-aware visual sampler. The data and existing models are not the interesting part.
- orenlindsey 3y agoHas anyone actually run this yet?
- deleted 3y ago[deleted]
- devinprater 3y agoThey're already going multi-modal? Holy crap, if google can't deliver in the accessibility space for this (image descriptions better than "the logo for the company"), then I'll definitely go back to Apple. I mean I do hope Apple cleans out bugs and makes VoiceOver feel like it won't fall over if I breathed hard, but their image descriptions, even without an LLM, are already clean and clear. More like "A green logo on a black background", where Google is, like I said, more like "The logo for the company." I guess it's kinda what we get when AI is crowdsourced rather than given good, high quality data to work with.
- zitterbewegung 3y agoHonestly if they are coming out with a paper now Apple has probably been working on it for a year or two at minimum . Next year releases of macOS / iOS are rumored to have LLMs as a feature .
- beoberha 3y agoI don’t mean to discount this work, but this particular model is the product of a few months of work tops. It’s effectively LLava with different training data, targeted at a specific use case. While I’m sure there is a significant effort at multimodal LLMs within Apple, this is just a tiny corner of it.
- refulgentis 3y ago> Honestly if they are coming out with a paper now Apple has probably been working on it for a year or two at minimum Why do you say that?
- zitterbewegung 3y agoAcademic papers can take that long …
- ex3ndr 3y agoThey literally mention that they built on top of llava that was released half year ago.
- jonplackett 3y agoPresumable because this is Conda none of this can be run on any Apple hardware despite people managing to get M processors to do a bit of dabbling with AI?
- _visgean 3y ago> because this is Conda none of this can be run on any Apple hardware conda supports m1? https://www.anaconda.com/blog/new-release-anaconda-distribution-now-supporting-m1 https://www.anaconda.com/blog/new-release-anaconda-distribut...
- jonplackett 3y agoDid not know that!
- adt 3y agoOld paper (Oct/2023), but the weights are new (Dec/2023): https://lifearchitect.ai/models-table/ https://lifearchitect.ai/models-table/
- rreichman 3y agoOct 23 is old :)
- moneycantbuy 3y agoanyone know what is the best open source model that allows commercial use and can run locally on an iphone?
- mandelken 3y agoMistral 7B is pretty good and the instruct v0.2 runs on my iPhone through MLC Chat. However, the ChatGPT4 app is much better in usability: better model, multi-modal with text/vision/speech and better UI.
- hackernewds 3y agogpt 4 allows commercial use?
- satvikpendem 3y agoWhy wouldn't it? They sell the API for a reason.
- WhitneyLand 3y agoYes and no. You can use it commercially but there are some restrictions, including some of a competitive nature, like using the output to train new LLMs. This is the restriction that Bytedance (Tiktok) was recently banned for violating.
- BrutalCoding 3y agoI’ve made an example app for a Flutter plugin I created that can do this. Open-source, runs natively on all major platforms. I shared videos showing it on my iPad Mini, Pixel 7, iPhone 12, Surface Pro (Win 10 & Ubuntu Jellyfish) and Macs (Intel & M archs). By all means, it’s not a finished app. I simply wanted to use on-device AI stuff in Flutter so I started with porting over llama.cpp, and later on I’ll tinker with porting over whatever is the state of the art (whisper.cpp, bark.cpp etc). Repo: https://github.com/BrutalCoding/aub.ai https://github.com/BrutalCoding/aub.ai For any of your Apple devices, use this: https://testflight.apple.com/join/XuTpIgyY https://testflight.apple.com/join/XuTpIgyY App is compatible with any GGUF files, but it must be in the ChatML prompt format otherwise the chat UI/bubbles probably gets funky. I haven’t made it customizable yet, after all - it’s just an example app of the plugin. But I am actively working on it to nail my vision. Cheers, Daniel
- shrimpx 3y agoApple has been looking sleepy on LLMs, but they've been consistently evolving their hardware+software AI stack, without much glitzy advertising. I think they could blow away Microsoft/OpenAI and Google, if suddenly a new iOS release makes the OpenAI/Bard chatbox look laughably antiquated. They're also a threat to Nvidia, if a significant swath of AI usage switches over to Apple hardware. Arm and TSMC would stand to win.
- fennecbutt 3y agoAre you so sure? Even this link is built on top of the work of others, I'm not sure they've contributed as much as you think they have.
- harryVic 3y agoCan you give an example? I switched to android because i use personal assistant a lot while driving and siri was absolutely horrible.
- shrimpx 3y ago- FaceID - Facial recognition in Photos - "Memories" in Photos - iOS keyboard autocomplete using LLMs. I am bilingual and noticed in the latest iOS it now does multi-language autocomplete and you no longer have to manually switch languages. - Event detection for Calendar - Depth Fusion in the iOS camera app, using ML to take crisper photos - Probably others... The crazy thing is most/all of these run on the device.
- pkage 3y agoI just wish you could turn the multilingual keyboard off—I find that I usually only type in one language at a time and having the autocomplete recommend the wrong languages is quite frustrating
- shrimpx 3y agoThat's true, I have found that mildly annoying sometimes. But most of the time it's a win. It was really annoying manually switching modes over and over when typing in mixed-language, which I do fairly often. It'd be great if there was a setting though.
- amitprasad 3y agoAlso relevant: LLM in a flash: Efficient Large Language Model Inference with Limited Memory Apple seems to be gearing up for significant advances in on-device inference using this LLMs https://arxiv.org/abs/2312.11514 https://arxiv.org/abs/2312.11514
- Thorrez 3y agoDoes Apple know that ferrets are illegal in California? https://www.legalizeferrets.org/ https://www.legalizeferrets.org/
- fagrobot 3y ago[flagged]
- a_rahmanshah 3y agoCan we run this on macOS?