12 ms·
Brave Leo now uses Mixtral 8x7B as default
- rhdunn 3y agoIf you want to run Mixtral 8x7B locally you can use llama.cpp (including with any of the supporting libraries/interfaces such as text-generation-webui) with https://huggingface.co/TheBloke/Nous-Hermes-2-Mixtral-8x7B-SFT-GGUF https://huggingface.co/TheBloke/Nous-Hermes-2-Mixtral-8x7B-S.... The smallest quantized version (2bit) needs 20GB of RAM (which can be offloaded onto the VRAM of a decent 4090 GPU). The 4bit quantized versions are the largest models that can just about fit onto a 32GB system (29GB-31B). The 6bit (41GB) and 8bit (52GB) models need a 64GB system. You would need multiple GPUs with shared memory if you wanted to offload the higher precision models to VRAM. I've experimented with the 7B and 13B models, but haven't experimented with these models yet, nor other larger models.
- jodleif 3y agoI prefer koboldcpp over llama.cpp. It’s easy to spilt between gpu/cpu on models larger than VRAM
- DrSiemer 3y agoRuns in Oobabooga textUi as well, if you add the llama.cpp extension. Easier interface imo, plus fun stuff like coqui and whisper integration.
- rhdunn 3y agoThat's interesting. It also looks like koboldcpp works better with long interactions, as it only processes changed tokens. I'm using llama.cpp with text-generation-webui and its OpenAI compatible API. I'll have to look to see if I can use koboldcpp with it.
- sp332 3y agoLlama.cpp has an interactive mode, but I don't think text-generation-webui uses it. https://github.com/ggerganov/llama.cpp/blob/master/examples/main/README.md#interaction https://github.com/ggerganov/llama.cpp/blob/master/examples/...
- jodleif 3y agoIndeed. Koboldcpp works fine with other UIs than the bundled one.
- magicalhippo 3y agoI've got an aging 2080Ti and Ryzen 3800X with 96GB RAM, any point in trying to mess with the GPU or? Haven't really been able to justify upgrading to a 4090 or similar given I play so few new games these days.
- htsh 3y agoYes, offloading some layers to the GPU and VRAM should still help. And 11gb isn't bad. If you're on linux or wsl2, I would run oobabooga with --verbose. Load a GGUF, start with a small number of GPU layers and creep up, keeping an eye on VRAM usage. If you're on windows, you can try out LM Studio and fiddle with layers while you monitor VRAM usage, though windows may be doing some weird stuff sharing ram. Would be curious to see the diffs. Specifically if there's a complexity tax in offloading that makes the CPU-alone faster but in my experience with a 3060 and a mobile 3080, offloading what I can makes a big diff.
- macNchz 3y ago> Specifically if there's a complexity tax in offloading that makes the CPU-alone faster Anecdotal, but I played with a bunch of models recently on a machine with a 16GB AMD GPU and 64GB of system memory/12 core CPU. I found offloading to significantly speed things up when dealing with large models, but there was seemingly an inflection point as I tested models that approached the limits of the system, where offloading did seem to significantly slow things down vs just running on the CPU.
- baq 3y agoI had only cuda installed and it took 2 ollama shell commands in WSL2 from quite literally 0 local LLM experience to running mixtral fast enough on a 1070 and 12700k. Go for it.
- sp332 3y agoLlama.cpp has --n-gpu-layers that lets you set how much of the model to put on the GPU.
- attentive 3y agokobold bundles and runs llama.cpp. So it should be fairly the same with convenient defaults.
- DreamGen 3y agoWhen talking about memory requirements one also needs to mention the sequence length. In case of Mixtral, which supports 32000 tokens, this can be a significant chunk of the memory used.
- viraptor 3y agoAnd if you want better performance when talking about code, you can try the dolphin-mixtral fine tuning https://huggingface.co/TheBloke/dolphin-2.7-mixtral-8x7b-GGUF https://huggingface.co/TheBloke/dolphin-2.7-mixtral-8x7b-GGU...
- thriw63748 3y agoWhy not normal RAM? Ryzen 5600 with 128GB DDR4 is perfectly fine to run mixtral 8bit, and costs less than $1000. GPUs are only needed if you can not wait 5 minutes for an answer, or for training.
- rhdunn 3y agoThat was what I was referring to with the 32/64 GB systems.
- snowfield 3y agoOr if you want multiple sessions at the same time. Or if you want to do anything else with your machine while it's running. But realistically, 5 minutes is too long. It should be conversational, and for that you need at least 5 tokens per second. Which your Ryzen just can't do.
- MPSimmons 3y ago>It should be conversational, and for that you need at least 5 tokens per second. To be fair, a lot of people are using this for non-interactive work, like batching document analysis or offline processing of user generated content.
- Gracana 3y agoIn my experience, it takes some experimentation to figure out a good prompt. I don’t think I would have gotten very far off I had to wait that long for each result.
- Diti 3y agoThis particular thread we are commenting on is about Dolphin Mixtral, which is mostly used for offline code completion (à là Microsoft GitHub Copilot). You don’t want to have to wait 5 minutes at every keystroke to get code suggestions.
- irusensei 3y agoWhy not both? Llama.cpp allows layering GGUF models between GPU and CPU memory.
- tarruda 3y ago> You would need multiple GPUs with shared memory if you wanted to offload the higher precision models to VRAM. Or just a powerful apple silicon machine? I've tried dolphin mixtral 4bit on a 36gb ram MacBook m3, and inference is super fast.
- pennaMan 3y agoI can run 4bit on a beat up 1070 ti. GP talks about higher precision models
- mkesper 3y agoOr a Linux machine with a Ryzen using the internal GPU and the unified RAM (scroll down at llama.cpp and look for ROCm).
- chpatrick 3y ago`ollama run mixtral:8x7b-instruct-v0.1-q3_K_L` works fast on my 3090 locally
- beck5 3y agoDumb question, but how can a 32 bit number be converted to 2 bits and still be useful? It seems like magic.
- ttoinou 3y agoAll the 32 bits weren't necessarily used, and it's the whole network itself that has to be useful. It's a tradeoff. We started with very good precision to test the new method, now we can optimize some parts of it
- sp332 3y agoThe extra precision is more useful for training. Once the network is optimized, it's a statistical model and only needs enough precision to make good guesses. In fact, one of the big papers on this also pointed out that you can drop about 40% of the weights completely. I think people generally skip that part because sparse matrix operations are slower, so it doesn’t help here.
- viraptor 3y agoFor models with dropped weights, the keyword is "distilled". For example ssd-1b is a 50% size version of Stable Diffusion XL (https://huggingface.co/segmind/SSD-1B https://huggingface.co/segmind/SSD-1B)
- sp332 3y agoThat’s crazy, I’ve never seen one that dropped whole layers from a pre-trained model. I guess that avoids the sparse matrix math.
- beefield 3y agoI think of it with this kind of analogy: the original image is stored with 32 bit color scheme. You can reduce the color scheme to 16 bit accuracy and still figure out pretty well what the image is about. 2 bit is stretching this to a bit far, basically either pixel is white or it is black, but even if you lose lots of nuances in the image, in many images even that gives you some idea whats going on in the image.
- zaat 3y agoUsing TheBloke nous-hermes-2-mixtral-8x7b-dpo.Q5_K_M.gguf, llam.cpp with oobabooga, on Ryzen 9 7950x and RTX 2070 super, offloading 4 layers to the GPU. 27,631MiB system memory used. 7358/8192MiB usage on the GPU. I'm getting 6.32 tokens/second. You: You are Hunter S. Thompson, write a 450 words essay on the LHC project. The writing must be in Thompson's style AI: Well, buckle up, kiddos, because we're about to dive into the twisted world of particle physics and the Large Hadron Collider, or as I like to call it, the Doomsday Machine. This beast of a thing is buried deep beneath the Swiss countryside, a 17-mile-long ring of death and destruction that makes the atom bomb look like a goddamn firecracker. Now, let me tell you, when the Swiss government first proposed this monstrosity, I thought they'd gone completely batshit crazy. But then again, this is the land of Toblerone and Heidi, so maybe I shouldn't be too surprised. The LHC is essentially a giant, underground racetrack for subatomic particles, where they're accelerated to nearly the speed of light before being smashed into each other like a couple of drunken frat boys at a kegger. The goal of this madness? To recreate the conditions that existed just after the Big Bang, when the universe was nothing more than a seething cauldron of energy and chaos. And if that doesn't sound like a recipe for disaster, I don't know what does. But let's not forget the real reason we're all here: the Higgs boson, or as it's more commonly known, the God Particle. This little bugger is what gives other particles their mass, and without it, we'd all be floating around like a bunch of cosmic tumbleweeds. The LHC was designed to find this elusive particle, and in 2012, they finally did it. Sort of. You see, the Higgs boson they found was a bit of a disappointment. It was too light, too wimpy, and didn't behave quite the way they expected it to. So now, the physicists are back at the drawing board, trying to figure out what the hell went wrong. And in the meantime, the LHC continues to churn out more and more data, like some kind of unholy particle-smashing machine. But let's not forget the dark side of the LHC. There are those who believe that this thing could very well be the end of the world as we know it. They fear that the LHC could create a black hole, or even a miniature Big Bang, right here on Earth. And while the scientists assure us that these fears are unfounded, I can't help but wonder if they're just whistling past the graveyard. So, what have we learned from all this? Well, for one thing, we've discovered that the universe is a much stranger place than we ever imagined. We've also learned that even the most brilliant minds can be blinded by their own ambition, and that sometimes, the pursuit of knowledge can lead us down some very dark paths. In the end, the LHC is a testament to the insatiable curiosity of the human race, and our unquenchable thirst for knowledge. It's a monument to the power of science, and the potential consequences of playing God. And as we continue to push the boundaries of what's possible, we must always remember that with great power comes great responsibility. Or, as I like to say, "When the going gets weird, the weird turn pro."
- bearjaws 3y ago2bit is pretty damn terrible, I don't recommend it for anything serious.
- ignoramous 3y agoAt that level of quantization / distillation, smaller models like phi-2 (q&a) and wavecoder-6.7b (code-gen) might be preferable over QLoRAd ones: https://huggingface.co/microsoft/phi-2 https://huggingface.co/microsoft/phi-2 > 2bit is pretty damn terrible Wait till you go hybrid [0] or even 1bit [1] [0] https://github.com/efeslab/Atom https://github.com/efeslab/Atom [1] https://github.com/IST-DASLab/qmoe https://github.com/IST-DASLab/qmoe
- EVa5I7bHFq9mnYK 3y agoFaraday.dev has it in its selection of models now. Good for us clueless Windows folks. Runs decently fast with 16gb mobile 3080 gpu. Results seem better than any other free option.
- MuffinFlavored 3y agoWhat differences would I measurably notice running the 2-bit version vs the 4-bit version vs the 6-bit vs the 8-bit?
- davikr 3y agoIt's nice using Brave because you have Chromium's better performance, without having to worry about Manifest V2 dying and taking adblocking down with it. I have uBlock Origin enabled, but it has barely caught anything that slipped past the browser filters.
- charcircuit 3y agoMV3 doesn't prevent adblockers from existing.
- rpastuszak 3y agoIt makes them almost useless in practice.
- charcircuit 3y agoThat is a baseless statement. It doesn't make them useless as they can still block ads.
- HeatrayEnjoyer 3y agoBecause the filter list is capped, right? Is there a reason the Brave team cannot just remove or increase the cap?
- gkbrk 3y agoNot just because of the filter list cap. It also reduces ad blockers to static filter lists instead of powerful dynamic filters. MV3 makes it impossible for ad-blockers to inspect requests with code and then allow/deny dynamically.
- charcircuit 3y ago>It also reduces ad blockers to static filter lists instead of powerful dynamic filters. This is very outdated information and borderline misinformation by representing it as how it currently works. It allows for 30,000 dynamic rules and 5,000 session rules (session rules only persist until the browser is closed). >MV3 makes it impossible for ad-blockers to inspect requests with code and then allow/deny dynamically. Giving this ability to extensions can slow down the browser for the user. These ads can still be blocked through other means.
- firtoz 3y agoWhat are good API providers that serve mixtral? I know only octo ai which seems decent but will be good to know alternatives too
- jug 3y agoYou also have Replicate: https://replicate.com/mistralai/mixtral-8x7b-instruct-v0.1 https://replicate.com/mistralai/mixtral-8x7b-instruct-v0.1
- htsh 3y agoopenrouter, fireworks, together. we use openrouter but have had some inconsistency with speed. i hear fireworks is faster, swapping it out soon.
- Tiberium 3y agoOpenRouter is generally a good option (already mentioned), the best part is that you have a unified API for all LLMs, and the pricing is the same as with the providers themselves. Although for OpenAI/Anthropic models they were forced (by the respective companies) to enable filtering for inputs/outputs.
- firejake308 3y agoI personally like Anyscale Endpoints
- mark_l_watson 3y agoI have used both Mistral’s commercial APIs and also AnyScale’s commercial APIs for mixtral-8-7b- both providers are easy to use. I also run a 3 bit quantization of mixtral-8-7b on my M2 Pro 32G memory system and it is fairly quick. It is great having multiple options.
- Philpax 3y agoI've had good experiences with Together, and they have very competitive pricing.
- bearjaws 3y agoTogether.ai seems to be the best, incredibly fast.
- fifteen1506 3y agoJust checking: PDF summarization is not yet implemented, right?
- _aaed 3y agoThe Kagi browser extension can do that, if you're a subscriber
- fifteen1506 3y agoAsk a PDF? I thought it was only the $25 a month plan.
- _aaed 3y agoNo, it's just text, like so: https://i.imgur.com/3NMzyDf.png https://i.imgur.com/3NMzyDf.png
- deleted 3y ago[deleted]
- charcircuit 3y agoIt's interesting that they made it so you can ask LLM queries right from the omnibar. I wonder if they eventually will come up with some heuristic to determine if thr query should be sent directly to an LLM or if the query should use the default search provider.
- finikytou 3y agoquick question I have 24GB VRAM and I need to close everything to run MIXTRAL at 4 bit quant with bitsandbyte. there is no way to run it at 3,5 on windows?
- syntaxing 3y agoInteresting, I must have missed the first Leo announcement. I really like how privacy conscious it is. They don’t store any chat record which is what I want.
- Dwedit 3y agoThere is no way to confirm that claim, just like there is no way to confirm that a VPN service is "no log".
- Erratic6576 3y agoYou gotta trust them by their word
- lolinder 3y agoYes, at some point if you're going to interface with other humans you will eventually just have to trust their word. For some people's threat models that isn't good enough, but for the vast majority of people—people who aren't being pursued by state intelligence agencies but who are squeamish about how much data a company like Google collects—a pinky promise from Brave or Mullvad is good enough.
- bcye 3y agoI would like to think GDPR ensures this pinky promise is good enough
- wolverine876 3y ago> For some people's threat models that isn't good enough, but for the vast majority of people—people who aren't being pursued by state intelligence agencies but who are squeamish about how much data a company like Google collects—a pinky promise from Brave or Mullvad is good enough. Who are you to say it's good enough (and ridicule people who disagree)? We don't have too much evidence of it, because they have very few options and of course most people are not informed and lack the expertise to understand the issues (a good situation for regulation). At one point lots of people used lead paint and were fine with it; they would have told us. > Yes, at some point if you're going to interface with other humans you will eventually just have to trust their word. There's technology, such as the authorization tokens used by Brave, that reduces that risk. Of course, no risk can be complete eliminated but that doesn't mean we shouldn't reduce it.
- kristianpaul 3y agoI run Mixtral locally using ollama
- andai 3y agoAsked Mistral 8x7B for an essay on ham. It started telling me about Hamlet.
- Erratic6576 3y agoIt must start from the beginning. Pig > piglet. Ham > Hamlet
- andai 3y agoWould make sense if it was the first token. But it's the last, presumably with a "end of user message" separator! (Or perhaps not? I don't know.)
- deleted 3y ago[deleted]
- m3kw9 3y agoIf you have used gpt4 and then use mistral, it’s like looking at a Retina display and then have to go back to a low res screen. You are always thinking “but GPT4 could do this though”
- mpalmer 3y agoHave you used mixtral?
- wolverine876 3y agoKudos to Brave (for this and other privacy features): Unlinkable subscription: If you sign up for Leo Premium, you’re issued unlinkable tokens that validate your subscription when using Leo. This means that Brave can never connect your purchase details with your usage of the product, an extra step that ensures your activity is private to you and only you. The email you used to create your account is unlinkable to your day-to-day use of Leo, making this a uniquely private credentialing experience.
- quinncom 3y agoThis is very cool, and something I’d like to integrate in my own apps. Does anybody know how this works exactly, not using foreign keys?
- luke-stanley 3y agoI could guess, an "anonymous payment credential service" could do something like this: 1. User completes payment for the paid for service, 2. To track the payment entitlement, a random, unique ID is generated by the service for the user, that is not related to any of their data. 3. This ID is saved in a database as a valid payment key. 4. The database records IDs in shuffled batches, or with semi-random fuzzy / low resolution timestamps to prevent correlation between payment time and ID generation. 5. Each ID has an entitlement limit or usage stopping point, ensuring it's only valid for the subscribed period. Another way might be Zero-Knowledge Proofs (ZKPs), but that might be more complex. They might even use their BAT crypto stuff for this somehow, I suppose. Whatever solution, would need a fundamental solution for how to avoid correlation, I think.
- emmanueloga_ 3y agoDoes anyone know of a good chrome extension for AI page summarization? I tried a bunch of the top Google search hits, they work fine but are really bloated with superfluous features.
- Terretta 3y agoSee Kagi's Universal Summarizer https://kagi.com/summarizer/index.html https://kagi.com/summarizer/index.html https://help.kagi.com/kagi/api/summarizer.html https://help.kagi.com/kagi/api/summarizer.html "Alternatively use Kagi Search browser extension (Chrome/Firefox) and you can use the most advanced Muriel model right from the extension."
- frozenport 3y agoI've been running the version on poe and chat.groq.com for the last week. Much better than llama 70b.