7 ms·
Qwen3-VL
- natrys 1y agoModels: - https://huggingface.co/Qwen/Qwen3-VL-235B-A22B-Thinking https://huggingface.co/Qwen/Qwen3-VL-235B-A22B-Thinking - https://huggingface.co/Qwen/Qwen3-VL-235B-A22B-Instruct https://huggingface.co/Qwen/Qwen3-VL-235B-A22B-Instruct
- causal 1y agoThat has got to be the most benchmarks I've ever seen posted with an announcement. Kudos for not just cherrypicking a favorable set.
- be7a 1y agoThe biggest takeaway is that they claim SOTA for multi-modal stuff even ahead of proprietary models and still released it as open-weights. My first tests suggest this might actually be true, will continue testing. Wow
- ACCount37 1y agoMost multi-modal input implementations suck, and a lot of them suck big time. Doesn't seem to be far ahead of existing proprietary implementations. But it's still good that someone's willing to push that far and release the results. Getting multimodal input to work even this well is not at all easy.
- Computer0 1y agoI feel like most Open Source releases regardless of size claim to be similar in output quality to SOTA closed source stuff.
- drapado 1y agoCool! Pity they are not releasing a smaller A3B MoE model
- daemonologist 1y agoTheir A3B Omni paper mentions that the Omni at that size outperformed the (unreleased I guess) VL. Edit: I see now that there is no Omni-235B-A22B; disregard the following. ~~Which is interesting - I'd have expected the larger model to have more weights to "waste" on additional modalities and thus for the opposite to be true (or for the VL to outperform in both cases, or for both to benefit from knowledge transfer).~~ Relevant comparison is on page 15: https://arxiv.org/abs/2509.17765 https://arxiv.org/abs/2509.17765
- ilc 1y agohttps://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct
- sergiotapia 1y agoThank you Qwen team for your generosity. I'm already using their thinking model to build some cool workflows that help boring tasks within my org. https://openrouter.ai/qwen/qwen3-235b-a22b-thinking-2507 https://openrouter.ai/qwen/qwen3-235b-a22b-thinking-2507 Now with this I will use it to identify and caption meal pictures and user pictures for other workflows. Very cool!
- willahmad 1y agoChina is winning the hearts of developers in this race so far. At least, they won mine already.
- swyx 1y agoso.. why do you think they are trying this hard to win your heart?
- llllm 1y agothey aren’t even trying hard, it’s just that no one else is trying
- willahmad 1y agoThey might have dozens of reasons, but they already did what they did. Some of the reasons could be: - mitigation of US AI supremacy - Commodify AI use to push forward innovation and sell platforms to run them, e.g. if iPhone wins local intelligence, it benefits China, because China is manufacturing those phones - talent war inside China - soften the sentiment against China in the US - they're just awesome people - and many more
- vanviegen 1y ago> - they're just awesome people Thank you for including that option in your list! F#ck cynicism.
- brokencode 1y agoMaybe they just want to see one of the biggest stock bubble pops of all time in the US.
- protocolture 1y agoI know I do
- binary132 1y ago
- deepdarkforest 1y agoThe Chinese are doing what they have been doing to the manufacturing industry as well. Take the core technology and just optimize, optimize, optimize for 10x the cost/efficiency. As simple as that. Super impressive. These models might be bechmaxxed but as another comment said, i see so many that it might as well be the most impressive benchmaxxing today, if not just a genuinely SOTA open source model. They even released a closed source 1 trillion parameter model today as well that is sitting on no3(!) on lm arena. EVen their 80gb model is 17th, gpt-oss 120b is 52nd https://qwen.ai/blog?id=241398b9cd6353de490b0f82806c7848c5d2777d&from=research.latest-advancements-list https://qwen.ai/blog?id=241398b9cd6353de490b0f82806c7848c5d2...
- jychang 1y agoThey still suck at explaining which model they serve is which, though. They also released today Qwen3-VL Plus [1] today alongside Qwen3-VL 235B [2] and they don't tell us which one is better. Note that Qwen3-VL-Plus is a very different model compared to Qwen-VL-Plus. Also, qwen-plus-2025-09-11 [3] vs qwen3-235b-a22b-instruct-2507 [4]. What's the difference? Which one is better? Who knows. You know it's bad when OpenAI has a more clear naming scheme. [1] https://modelstudio.console.alibabacloud.com/?tab=doc#/doc/?type=model&url=2840914_2&modelId=qwen3-vl-plus https://modelstudio.console.alibabacloud.com/?tab=doc#/doc/?... [2] https://modelstudio.console.alibabacloud.com/?tab=doc#/doc/?type=model&url=2840914_2&modelId=qwen3-vl-235b-a22b-instruct https://modelstudio.console.alibabacloud.com/?tab=doc#/doc/?... [3] https://modelstudio.console.alibabacloud.com/?tab=doc#/doc/?type=model&url=2840914_2&modelId=qwen-plus-2025-09-11 https://modelstudio.console.alibabacloud.com/?tab=doc#/doc/?... [4] https://modelstudio.console.alibabacloud.com/?tab=doc#/doc/?type=model&url=2840914_2&modelId=qwen3-235b-a22b-instruct-2507 https://modelstudio.console.alibabacloud.com/?tab=doc#/doc/?...
- deepdarkforest 1y agoEh i mean often innovation is made just by letting a lot of fragmented, small teams of cracked nerds trying out stuff. It's way too early in the game. I mean, qwens release statements have anime etc. IBM, Bell, Google, Dell, many did it similarly, letting small focused teams having many attempts at cracking the same problem. All modern quant firms are doing basically the same as well. Anthropic is actually an exception, more like Apple.
- BUFU 1y agoThe open source models are no longer catching up. They are leading now.
- buyucu 1y agoIt has been like that for a while now. At least since Deepseek R1.
- jadbox 1y agoHow does it compare to Omni?
- helloericsf 1y agoIf you're in SF, you don't want to miss this. The Qwen team is making their first public appearance in the United States, with the VP of Qwen Lab speaking at the meetup below during SF teach week. https://partiful.com/e/P7E418jd6Ti6hA40H6Qm https://partiful.com/e/P7E418jd6Ti6hA40H6Qm Rare opportunity to directly engage with the Qwen team members.
- dazzaji 1y agoRegistration full :-(
- alfiedotwtf 1y agoLet’s hope they’re allowed in the country and get a visa… it’s 50/50 these days
- richardlblair 1y agoAs I mentioned yesterday - I recently needed to process hundreds of low quality images of invoices (for a construction project). I had a script that had used pil/opencv, pytesseract, and open ai as a fallback. It still has a staggering number of failures. Today I tried a handful of the really poor quality invoices and Qwen spat out all the information I needed without an issue. What's crazier is it gave me the bounding boxes to improve tesseract.
- creativebee 1y agoAny tipps on getting bounding boxes? The model doesn’t seem to even understand the original size of the image. And even if I provide the dimensions, the positioning is off. :'(
- VladVladikoff 1y agoInteresting. I have in the past tried to get bounding boxes of property boundaries on satellite maps estimated by VLLM models but had no success. Do you have any tips on how to improve the results?
- mh- 1y agoDo you have some example images and the prompt you tried?
- BOOSTERHIDROGEN 1y agoalso documented stack setup if could.
- richardlblair 1y agoWith Qwen I went as stupid as I could: please provide the bounding box metadata for pytesseract for the above image. And it spat it out.
- VladVladikoff 1y agoIt’s funny that many of us say please. I don’t think it impacts the output, but it also feels wrong without it sometimes.
- mountainriver 1y agoIncredible release! Qwen has been leading the open source vision models for a while now. Releasing a really big model is amazing for a lot of use cases. I would love to see a comparison to the latest GLM model. I would also love to see no one use OS World ever again, it’s a deeply flawed benchmark.
- mythz 1y agoTeam Qwen keeps cooking! qwen2.5VL was already my preferred visual model for querying images, will look at upgrading if they release a smaller model we can run locally.
- am17an 1y agoThis model is literally amazing. Everyone should try to get their hands on a H100 and just call it a day.
- clueless 1y agoThis demo is crazy: "At what time was the goal scored in this match, who scored it, and how was it scored?"
- addandsubtract 1y agoI had the same reaction, given the 100min+ runtime of the video.
- vessenes 1y agoRoughly 1/10 the cost of Opus 4.1, 1/2 the cost of Sonnet 4 on per token inference basis. Impressive. I'd love to see a fast (groq style) version of this served. I wonder if the architecture is amenable.
- petesergeant 1y agoCerebras are hosting other Qwen models via OpenRouter, so probably
- aitchnyu 1y agoIsnt it a 3x rate difference? 0.7$ for Qwen3-VL vs 3$ for Sonnet 4?
- vessenes 1y agoOpenrouter had $8-ish / 1M tokens for Qwen and $15/M for Sonnet 4 when I checked
- vessenes 1y agoI spent a little time with the thinking model today. It's good. It's not better than GPT5 Pro. It might be better than the smallest GPT 5, though. My current go-to test is to ask the LLM to construct a charging solution for my macbook pro with the model on it, but sadly, I and the pro have been sent to 15th century Florence with no money and no charger. I explain I only have two to three hours of inference time, which can be spread out, but in that time I need to construct a working charge solution. So far GPT-5 Pro has been by far the best, not just in its electrical specifications (drawings of a commutator), but it generated instructions for jewelers and blacksmith in what it claims is 15th century florentine italian, and furnished a year-by year set of events with trading / banking predictions, a short rundown of how to get to the right folks in the Medici family, .. it was comprehensive. Generally models suggest building an Alternating current setup and then rectifying to 5V of DC power, and trickle charging over the USB-C pins that allow trickle charging. There's a lot of variation in how they suggest we get to DC power, and often times not a lot of help on key questions, like, say "how do I know I don't have too much voltage using only 15th century tools?" Qwen 3 VL is a mixed bag. It's the only model other than GPT5 I've talked to that suggested building a voltaic pile, estimated voltage generated by number of plates, gave me some tests to check voltage (lick a lemon, touch your tongue. Mild tingling - good. Strong tingling, remove a few plates), and was overall helpful. On the other hand, its money making strategy was laughable; predicting Halley's comet, and in exchange demanding a workshop and 20 copper pennies from the Medicis. Anyway, interesting showing, definitely real, and definitely useful.
- ripped_britches 1y agoThat is a freaking insanely cool answer from gpt5
- nl 1y ago> predicting Halley's comet, and in exchange demanding a workshop and 20 copper pennies from the Medicis I love this! Simple and probably effective (or would get you killed for witchcraft)
- vessenes 1y ago
- Workaccount2 1y agoSadly it still fails the "extra limb" test. I have a few images of animals with an extra limb photoshopped onto them. A dog with an leg coming out of it's stomach, or a cat with two front right legs. Like every other model I have tested, it insists that the animals have their anatomically correct amount of limbs. Even pointing out there is a leg coming from the dogs stomach, it will push back and insist I am confused. Insist it counted again and there are definitely only 4. Qwen took it a step further and even after I told it the image was edited, it told me it wasn't and there were only 4 limbs.
- brookst 1y agoDefinitely not a good model for accurately counting limbs on mutant species, then. Might be good at other things that have greater representation in the training set.
- user34283 1y agoI'm not knowledgeable about ML but it seems disappointing how we went from "models are able to generalize" and "emergent capabilities" to "can't do anything not greatly represented in the training set".
- ComputerGuru 1y agoI wonder if you used their image editing feature if it would insist on “correcting” the number of limbs even if you asked for unrelated changes.
- vunderba 1y agoIt will. I actually made a test when NanoBanana first went GA which featured a photo of a one-legged man and asked the model to change the clothing into pants. It added the pants as requested and then proceeded to "heal" his missing leg in the process. Very difficult for even SOTA to go against data that is as well-represented as bipedal humanoids. https://mordenstar.com/blog/edits-with-nanobanana https://mordenstar.com/blog/edits-with-nanobanana
- 1y ago
- whitehexagon 1y agoImagine the demand for a 128GB/256GB/512GB unified memory stuffed hardware linux box shipping with Qwen models already up and running. Although I´m agAInst steps towards AGI, it feels safer to have these things running locally and disconnected from each other, than some giant GW cloud agentic data centers connected to everyone and everything.
- buyucu 1y agoI bought an GMKtec evo 2 that is a 128 GB unified memory system. Strong recommend.
- te0006 1y agoInteresting - do you need to take any special measures to get OSS genAI models to work on this architecture? Can you use inference engines like Ollama and vLLM off-the-shelf (as Docker containers) there, with just the Radeon 8060S GPU? What token rates do you achieve? (edit: corrected mistake w.r.t. the system's GPU)
- buyucu 1y agoI just use llama.cpp. It worked out of the box.
- Keyframe 1y agoThat's AMD Ryzen AI Max+ 395, right? Lots of those boxes popping up recently, but isn't that dog slow? And I can't believe I'm saying this - but maybe RAM filled-up mac might be a better option?
- buyucu 1y agoI'm not buying a Mac. Period.
- ricardobeat 1y agoYes, but the mac costs 3-4x more. You can get one of these 395 systems with 96GB for ~1k.
- fareesh 1y agoCan't seem to connect to qwen.ai with DNSSEC enabled > resolvectl query qwen.ai > qwen.ai: resolve call failed: DNSSEC validation failed: no-signature And https://dnsviz.net/d/qwen.ai/dnssec/ https://dnsviz.net/d/qwen.ai/dnssec/ shows aliyunga0019.com/DNSKEY: No response was received from the server over UDP (tried 4 times). See RFC 1035, Sec. 4.2. (8.129.152.246, UDP_-_EDNS0_512_D_KN)
- vardump 1y agoSo 235B parameter Qwen3-VL is FP16, so practically it requires at least 512 GB RAM to run? Possibly even more for a reasonable context window? Assuming I don’t want to run it on a CPU, what are my options to run it at home under $10k? Or if my only option is to run the model with CPU (vs GPU or other specialized HW), what would be the best way to use that 10k? vLLM + Multiple networked (10/25/100Gbit) systems?
- bitflourjikg 1y agoA non-CPU setup will very likely require an electrical service upgrade or tactical positioning of different systems on different circuits for you to be able to run models that large. Several kW setups also cost non-trivial sums of money to run usually
- loudmax 1y agoAn Apple Mac Studio with 512GB of unified memory is around the $10k. If your really need that much power on your home computer, and you have that much money to spend, this could be the easiest option. You probably don't need fp16. Most models can be quantized down to q8 with minimal loss of quality. Models can usually be quantized to q4 or even below and run reasonably well, depending on what you expect out of them. Even at q8, you'll need around 235GB of memory. An Nvidia RTX 5090 has 32GB of VRAM and has an official price of about $2000, but usually retails for more. If you can find them at that price, you'd need eight of them to run a 235GB model entirely in VRAM, and that doesn't include a motherboard and CPU that can handle eight GPUs. You could look for old mining rigs built from RTX 3090s or P40s. Otherwise, I don't see much prospect for fitting this much data into VRAM on consumer GPUs for under $10k. Without NVLink, you're going to take a massive performance hit running a model distributed over several computers. It can be done, and there's research into optimizing distributed models, but the throughput is a significant bottleneck. For now, you really want to run on a single machine. You can get pretty good performance out of a CPU. The key is memory bandwidth. Look at server or workstation class CPUs with a lot of DDR5 memory channels that support a high MT/s rate. For example, an AMD Ryzen Threadripper 7965WX has eight DDR5 memory channels at up to 5200 MT/s and retails for about $2500. Depending on your needs, this might give you acceptable performance. Lastly, I'd question whether you really need to run this at home. Obviously, this depends on your situation and what you need it for. Any investment you put into hardware is going to depreciate significantly in just a few years. $10k of credits in the cloud will take you a long way.
- Alifatisk 1y agoWow, the Qwen team doesn't stop and keep coming up with surprises. Not only did they release this but also the new Qwen3-Max model
- isoprophlex 1y agoExtremely impressive, but can one really run these >200B param models on prem in any cost effective way? Even if you get your hands on cards with 80GB ram, you still need to tie them together in a low-latency high-BW manner. It seems to me that small/medium sized players would still need a third party to get inference going on these frontier-quality models, and we're not in a fully self-owned self-hosted place yet. I'd love to be proven wrong though.
- Borealid 1y agoA Framework Desktop exposes 96GB of RAM for inference and costs a few thou USD.
- michaelanckaert 1y agoYou need memory on the GPU, not in the system itself (unless you have unified memory such as the M-architecture). So we're talking about cards like the H200 that have 141GB of memory and cost between 25 to 40k.
- Borealid 1y agoDid you casually glance at how the hardware in the Framework Desktop (Strix Halo) works before commenting?
- michaelanckaert 1y agoI didn't glace at it, I read it :-) The architecture is a 'unified memory bus', so yes the GPU has access to that memory. My comment was a bit unfortunate as it implied I didn't agree with yours, sorry for that. I simply want to clarify that there's a difference between 'GPU memory' and 'system memory'. The Frame.work desktop is a nice deal. I wouldn't buy the Ryzen AI+ myself, from what I read it maxes out at about 60 tokens / sec which is low for my use cases.
- ramon156 1y agoThese don't run 200B models at all, results show it can run 13B at best. 70B is ~3 tk / s according to someone on Reddit.
- michaelanckaert 1y agoQwen has some really great models. I recently used qwen/qwen3-next-80b-a3b-thinking as a drop-in replacement for GPT-4.1-mini in an agent workflow. Cost 4 times less for input tokens and half for output, instant cost savings. As far as I can measure, system output has kept the same quality.
- buyucu 1y agoThe Chinese are great. They are making major contributions to human civilization by open sourcing these models.
- ramon156 1y agoOne downside is it has less knowledge of lesser known tools like orpc, which is easily fixed by something like context7
- ashvardanian 1y agoQwen models have historically been pretty good, but there seems to be no architectural novelty here, if I’m not missing it. Seems like another vision encoder, with a projection, and a large autoregressive model. Have there been any better ideas in the VLM space recently? I’ve been away for a couple of years :(
- youssefarizk 1y agoAnother day another Qwen model