12 ms·
https://docs.mistral.ai/platform/pricing https://docs.mistral.ai/platform/pricing Pricing has been released too. Per 1 million output tokens: Mistral-medium
by rrsp 3y ago
https://docs.mistral.ai/platform/pricing https://docs.mistral.ai/platform/pricing
Pricing has been released too.
Per 1 million output tokens:
Mistral-medium $8
Mistral-small $1.94
gpt-3.5-turbo-1106 $2
gpt-4-1106-preview $30
gpt-4 $60
gpt-4-32k $120
This suggests that they’re reasonably confident that the mistral-medium model is substantially better than gpt3-5
- infecto 3y agoI don’t think it’s safe to assume any of this. It’s still limited release which reads as invite only. Once it hits some kind of GA then we can test and verify.
- raincole 3y agoIt's safe to assume they are confident it's better than 3.5. But people can be confident and wrong.
- infecto 3y agoWe won’t know anything until it becomes a wider release and can test it.
- ilaksh 3y agoMultiple people have tested it. Code and weights are fully released .
- rockinghigh 3y agoMistral-medium has not been released yet.
- dcastm 3y agoIf you take input tokens in consideration is more like 5.25 eur vs. 1.5 eur / million tokens overall. Mistral-small seems to be the most direct competitor to gpt-3.5 and it’s cheaper (1.2 eur / million tokens) Note: I’m assuming equal weight for input and output tokens, and cannot see the prices in USD :/
- stavros 3y agoDoes the 8x7B model really perform at a GPT-3.5 level? That means we might see GPT-3.5 models running locally on our phones in a few years.
- anon373839 3y agoThat might be happening in a few weeks. There is a credible claim that this model might be compressible to as little as a 4GB memory footprint.
- stavros 3y agoYou mean the 7B one? That's exciting if true, but if compression means it can do 0.1 token/sec,it doesn't do much for anyone.
- infecto 3y agoNot true. Not everyone is building chat bot or similar interface that requires output with latency low enough for a user. While your example is of course incredibly slow, there are still many interesting things that could be done if it was a little bit quicker.
- stavros 3y agoWhat kind of use cases run in an environment where latency isn't important (some kind of batch process?) but don't have more than 4GB of RAM?
- wongarsu 3y agoPrice sensitive ones, or cases where you want the new capability but can't get any new infrastructure.
- TeMPOraL 3y agoNot LLMs, but locally running facial and object recognition models on your phone's gallery, to build up a database for face/object search in the gallery app? I'm half-convinced this is how Samsung does it, but I can't really be sure of much, because all the photo AI stuff works weirdly and in unobservable way, probably because of some EU ruling. (That one is a curious case. I once spent some time trying to figure out why no major photo app seems to support manually tagging faces, which is a mind-dumbingly obvious feature to support, and which was something supported by software a decade or so ago. I couldn't find anything definitive; there's this eerie conspiracy of silence on the topic, that made me doubt my own sanity at times. Eventually, I dug up hints that some EU ruling/regs related to facial recognition led everyone to remove or geolock this feature. Still nothing specific, though.)
- code51 3y agogpt-3.5 is heavily subsidized. Mistral may just be aiming for a more sustainable price for the long run.
- raverbashing 3y agoDo they all use the same tokenizer? (I mean, Mistral vs GPT)
- superkuh 3y agoNo. Mistral uses sentencepiece and the GPT use tiktoken.
- YetAnotherNick 3y ago> This suggests that they’re reasonably confident that the mistral-medium model is substantially better than gpt3-5 How did you reach the conclusion? Maybe they are counting on people paying extra just to prevent vendor lockdown.
- antifa 3y agoThe only vendor lock-in to GPT3.5 is the absence (perceived or real) of competitors at the same quality and availability.
- epups 3y agoI understand how Mistral could end up being the most popular open source LLM model for the foreseeable future. What I cannot understand is who they expect to convince to pay for their API. As long as you are shipping your data to a third-party, whether they are running an open or closed source model is inconsequential.
- antifa 3y agoIf I'm happy with my infrastructure being built on top of the potential energy of a loadbearing rugpull, I'd probably stick with OpenAI in the average use case.
- chadash 3y agoI pay for hosted databases all the time. It’s more convenient. But those same databases are popular because they are open source. I also know that because it’s open source, if I ever have a need to, I can host it on my own servers. Currently I don’t have that need, but it’s nice to know that it’s in the cards.
- epups 3y agoOpen source databases are SOTA or very close to it, though. Here the value proposition is to pay 10-50% less for an inferior product. Portability is definitely an advantage, but that's another aspect which I think detracts from their value: if I can run this anywhere, I will either host it myself or pay whoever can make it happen very cheap. Even OpenAI could host an API for Mistral.
- chadash 3y ago> Here the value proposition is to pay 10-50% less for an inferior product. OpenAI just went through an existential crisis where the company almost collapsed. They are also quite unreliable. For some use cases, I'll take a service that does slightly worse on outputs, but much better on reliability. For example, if I'm building a customer service chat bot, it's a pretty big deal if the LLM backend goes down. With an open-source model, I can build it using the cloud provider. If they are a reliable host, i'll probably stick with them as i grow. If not, I always have the option of running the model myself. This alleviates a lot of the risk.
- raphaelj 3y agoDo we have estimates of the energy requirements for these models? I just did some napkin math, looks like inference on a 30B model with a GTX 4090 should get you about 30 tokens/sec [1], or 100k tokens/hour. Considering such systems consume about 1 kW, that's about 10 kWh/1M tokens. Based on the current cost of electricity, I don't think anyone could get below 2 ~ 4 $ per 1M token for a 30B model. [1] https://old.reddit.com/r/LocalLLaMA/comments/13j5cxf/how_many_tokens_per_second_do_you_guys_get_with/ https://old.reddit.com/r/LocalLLaMA/comments/13j5cxf/how_man...
- avereveard 3y agoBatching changes that equation a fair bit. Also these cards will not consume full power since llm are mostly limited by memory bandwidth and the processing part will get some idle time.
- Filligree 3y agoThe 4090 is considerably more power-hungry compared to e.g. an A100, however.
- fpgaminer 3y agoIf comparing apples to apples, the 4090 needs to clock up and consume about 450 W to match the A100 at 350W. Part of that is due to being able to run larger batches on the A100, which gives it an additional performance edge, but yes in general the A100 is more power efficient.
- jillesvangurp 3y agoDepends how and where you source your energy. If you invest in your own solar panels and batteries, all that energy is essentially fixed price (cost of the infrastructure) amortized over the lifetime of the setup (1-2 decades or so). Maybe you have some variable pricing on top for grid connectivity and use the grid as a fallback. But there's also the notion of selling excess energy back to the grid that offsets that. So, 10kwh could be a lot less than what you cite. That's also how grid operators make money. They generate cheaply and sell with a nice margin. Prices are determined by the most expensive energy sources on the grid in some markets (coal, nuclear, etc.). So, that pricing doesn't reflect actual cost for renewables, which is typically a lot lower than that. Anyone consuming large amounts of energy will be looking to cut their cost. For data centers that typically means investing in energy generation, storage, and efficient hardware and cooling.
- up6w6 3y agoI think the medium is trying to compete with Anthropic's Claude than Openai's products https://www-files.anthropic.com/production/images/model_pricing_dec2023.pdf https://www-files.anthropic.com/production/images/model_pric...
- antifa 3y agoAll they have to do to beat Anthropic's Claude is to skip having a permanent waitlist and let the credit cards get charged.