6 ms·
Gemma3 – The current strongest model that fits on a single GPU
- sigmoid10 2y agoThese bar charts are getting more disingenuous every day. This one makes it seem like Gemma3 ranks as nr. 2 on the arena just behind the full DeepSeek R1. But they just cut out everything that ranks higher. In reality, R1 currently ranks as nr. 6 in terms of Elo. It's still impressive for such a small model to compete with much bigger models, but at this point you can't trust any publication by anyone who has any skin in model development.
- pzo 2y agoopen llm leaderboard [0] is probably good to compare open weights model on many different benchmarks - wish they put also some closed source one just to see what's relative ranking of best open weights to closed source one. They haven't updated yet for gemma 3 though [0] https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard https://huggingface.co/spaces/open-llm-leaderboard/open_llm_...
- sigmoid10 2y agoBeware that they use very narrow metrics. Which is also why you only see fine-tunes over there gaming narrow aspects. If your edge case fits into one of those - great. If not and you just want a good general purpose model you'll have to look elsewhere.
- swores 2y agoThe chart isn't claiming to be an overview of the best ranking models - it's an evaluation of this particular model, which wouldn't be helped at all by having loads more unrelated models in the chart, even if that would have helped you avoid misunderstanding the point of the chart.
- sigmoid10 2y agoHow are better ranking models unrelated? They are explicitly comparing open and closed, small and large foundation models. Leaving the best ones out is just plain disingenuous. There's no way to sugarcoat this.
- antirez 2y agoThe most disturbing thing is that in the chart it ranks higher than V3. Test a few prompts against DeepSeek V3 and Gemma 3. They are like at two totally different levels, one is a SOTA model, one is a small LLM that can be useful for certain vertical tasks perhaps.
- mythz 2y agoNot sure if anyone else experiences this, but ollama downloads starts off strong but the last few MBs take forever. Finally just finished downloading (gemma3:27b). Requires the latest version of Ollama to use, but now working, getting about 21 tok/s on my local 2x A4000. From my few test prompts looks like a quality model, going to run more tests to compare against mistral-small:24b to see if it's going to become my new local model.
- dizhn 2y agoIt might not be downloading but converting the model. Or if it's already downloading a properly formatted model file, deduping on disk which I hear it does. This also makes its model files on disk useless for other frontends.
- Patrick_Devine 2y agoThere are some fixes coming to uniformly speed up pulls. We've been testing that out but there are a lot of moving pieces with the new engine so it's not here quite yet.
- squeakywhite 2y agoI experienced this just now. The download slowed down to approx 500kB/s for the last 1% or so. When this happens, you can Ctrl+C to cancel and then start the download again It will continue from where it left off, but at regular (fast) download speed.
- swores 2y agoSee the other HN submission (for the Gemma3 technical report doc) for a more active discussion thread - 50 comments at time of writing this. https://news.ycombinator.com/item?id=43340491 https://news.ycombinator.com/item?id=43340491
- archerx 2y agoI have tried a lot of local models. I have 656GB of them on my computer so I have experience with a diverse array of LLMs. Gemma has been nothing to write home about and has been disappointing every single time I have used it. Models that are worth writing home about are; EXAONE-3.5-7.8B-Instruct - It was excellent at taking podcast transcriptions and generating show notes and summaries. Rocinante-12B-v2i - Fun for stories and D&D Qwen2.5-Coder-14B-Instruct - Good for simple coding tasks OpenThinker-7B - Good and fast reasoning The Deepseek destills - Able to handle more complex task while still being fast DeepHermes-3-Llama-3-8B - A really good vLLM Medical-Llama3-v2 - Very interesting but be careful Plus more but not Gemma.
- mythz 2y agoConcur with Gemma2 being underwhelming, I dismissed it pretty quickly but gemma3:27b is looking pretty good atm. BTW mistral-small:24b is also worth mentioning (IMO best local model) and phi4:14b is also pretty strong for its size. mistral-small was my previous local goto model, testing now to see if gemma3 can replace it.
- InsideOutSanta 2y agoOne more vote for Mistral for local models. The 7B model is extremely fast and still good enough for many prompts.
- pduggishetti 2y agoRecently phi4 has been very good too!
- m00dy 2y agosshht, don't make it a public debate :P)
- anon373839 2y agoFrom the limited testing I've done, Gemma 3 27B appears to be an incredibly strong model. But I'm not seeing the same performance in Ollama as I'm seeing on aistudio.google.com. So, I'd recommend trying it from the source before you draw any conclusions. One of the downsides of open models is that there are a gazillion little parameters at inference time (sampling strategy, prompt template, etc.) that can easily impair a model's performance. It takes some time for the community to iron out the wrinkles.
- wtcactus 2y agoThe claim of “strongest” (what does that even mean?) seems moot. I don’t think a multimodal model is the way to go to use on single, home, GPUs. I would much rather have specific tailored models to use in different scenarios, that could be loaded into the GPU when needed. It’s a waste of parameters to have half of the VRAM loaded with parts of the model targeting image generation when all I want to do is write code.
- leumon 2y agoIn my opinion qwq is the strongest model that fits on a single gpu (Rtx 3090 for example, in Q4_K_M quantization which is the standard in Ollama)
- moffkalast 2y agoGemma 2 27B at 4 bits would be a drooling idiot anyway, even going down to 8 bits seems to significantly lobotomize it. Qwens are surprisingly resistant to quantization compared to most so it'll pull ahead just in that already in terms of coherence for the same VRAM amount. We'll see if the quantization aware versions are any better this time around, but I doubt any inference framework will even support them. Gemma.cpp never got a a standard compatible server API so people could actually use it, and as a result got absolutely zero adoption.
- hnfong 2y agoQuants at 4 bits are generally considered good, and 8 bits are generally considered overkill unless somehow need to squeeze the last bits of performance (in terms of generation quality) from it. There are papers to that effect though admittedly perhaps specific models might have divergent behavior ( https://arxiv.org/abs/2212.09720 https://arxiv.org/abs/2212.09720 ) All the above is subjective so maybe that’s true for you, but claiming there’s a lack of inference framework for gemma 2 is really off the mark. Obviously ollama supports it. Also llama.cpp. Also mlx. I’ve listed 3 frameworks that support quantized versions of gemma 2 llama.cpp support for gemma-3 is out, the PR merged a couple hours after googles announcement. Obviously ollama supports it as well as you can see in TFA here. I’m really curious how you’d get to the conclusions you’ve made. Are we living in different alternate universes?
- moffkalast 2y ago> Quants at 4 bits are generally considered good, and 8 bits are generally considered overkill Two year old info, only really applies to heavily undertrained models with short tokenizers. Perplexity scores are a really terrible metric for measuring quantization impact, and quantized models tend to also score higher than they should in benchmarks ran as topk=1 where the added randomness seems to help. In my experience it really seems to affect reliability most, which isn't often tested consistently. An fp16 model might get a question right every time, Q8 every other time, Q6 every third time, etc. In a long form conversation this means wasting a lot of time regenerating responses when the model throws itself off and loses coherence. It also destroys knowledge that isn't very strongly ingrained, so low learning rate fine tune data gets obliterated at a much higher rate. Gemma-2 specifically also loses a lot of its multilingual ability with quantization. I used to be in the Q6 camp for a long time, these days I run as much as I can in FP16 or at least Q8, because it's worth the tradeoff in most cases. Now granted it's different for cases like R1 when training is native FP8 or with QAT, how different I'm not sure since we haven't had more than a few examples yet. > there’s a lack of inference framework for gemma 2 is really of the mark I mean mainly for the QAT format for Gemma 3, which surprisingly seems to be as a standard gguf this time. Last time around Google decided llama.cpp is not good enough for them and half-assedly implemented their own ripoff as gemma.cpp with basically zero usable features. > llama.cpp support for gemma-3 is out Yeah testing it right now, I'm surprised it runs coherently at all given the new global attention tbh. Every architectural change is usually followed with up to a month of buggy inference and back and forth patching, model reuploads and similar nonsense.
- antirez 2y agoAfter reading the technical report do the effort of downloading the model and run it against a few prompts. In 5 minutes you understand how broken LLM benchmarking is.
- toinewx 2y agocan you expand a bit?
- antirez 2y agoThe model performs very poorly in practice, while in the benchmark it is shown to be DeepSeek V3 level. It's not terrible but it's at another level compared to the models it is very close to (a bit better / a bit worse) in the benchmarks.
- tarruda 2y agoIn my experience, Gemma models were always bad at coding (but good at other tasks).
- alekandreev 2y agoHey, Gemma engineer here. Can you please share reports on the type of prompts and the implementation you used?
- pcdoodle 2y ago[flagged]
- sgt101 2y agovibe testing, vibe model engineering...
- tarruda 2y agoCan you share the all the recommended settings to run this LLM? It is clear that the performance is very good when running on AI studio. If possible, I'd like to use the all the same settings (temp, top-k, top-p, etc) on Ollama. AI studio only shows Temperature, top-p and output length.
- tarruda 2y agoIs "OpenAI" the only AI company that hasn't released any model weights?
- world2vec 2y agoAnthropic hasn't released anything either AFAIK
- dev0p 2y agoThey need to open source Sonnet 3.7. I know they won't, but a man can dream.
- ddalex 2y agoI'd wish people stop using "open sourcing" when speaking about models. Open sourcing is about being able to change and replicate builds, they make the models "freely available" but the recipe on how they are made is kept secret. It's akin to being able to download Windows shareware executables and calling that "open source" when nothing related to how the executables are build is available.
- andai 2y agoWell, yeah. They need to do that, too!
- whiplash451 2y agoEven if they did release the code, that would not help you much unless you have the $ for traning and the talented individuals for pipelining and distributed training.
- toinewx 2y agowould you be able to run Sonnet 3.7 on a consumer computer though?
- 2y ago
- elif 2y agoGood job Google. It is kinda hilarious that 'open'AI seems to be the big player least likely to release any of their models.
- amelius 2y agolyingAI
- kjrfghslkdjfl 2y ago[dead]
- tekichan 2y agoI found deepseek better for trivial tasks
- aravindputrevu 2y agoI'm curious. Is there any value to do these OSS models? Suddenly after reasoning models, it looks like OSS models have lost their charm
- archerx 2y agoThee are a lot of open source reasoning models. The true value to local models is privacy and the ability to have the models be uncensored.
- lelag 2y agoOSS model do not have to be local models, and it's not just about privacy, imo. DeepSeek R1 hosting is out of reach for most, but it being open is a game changer if you are a building a business that needs the SoTA capabilities of such a large model, not because you will necessarily host it yourself, but because you can't be locked out of using it. If you build your business on top of OpenAI, and they decide they don't like you, they can shut you down. If you use an open model like R1, you always have the option to self host even if it can be costly, and not be at the mercy of a third party being able to just kill your business by shutting down your access to their service.
- whiplash451 2y agoYou can absolutely be locked out effectively if they stop releasing upgrades while the other providers move forward.
- pzo 2y agoAnother benefit is they can be fine tuned. Also it's not only about if Openai will shut you down but decide to deprecate model (like they will do for gpt4.0) or swap the name for different model (like sonnet 3.5 did) or censure it or limit capability.
- whiplash451 2y agoUncensored at inference time does not imply uncensored at training time (not a specific comment about Gemma)
- deleted 2y ago[deleted]
- iamgopal 2y agoSmall Models should be train on specific problem in specific language, and should be built one upon another, the way container works. I see a future where a factory or home have local AI server which have many highly specific models, continuously being trained by super large LLM on the web, and are connected via network to all instruments and computer to basically control whole factory. I also see a future where all machinery comes with AI-Readable language for their own functioning. A http like AI protocol for two way communication between machine and an AI. Lots of possibility.
- wewewedxfgdf 2y agoDiscrete GPUs are finished for AI. They've had years to provide the needed memory but can't/won't. The future of local LLMs is APUs such as Apple M series and AMD Strix Halo. Within 12 months everyone will have relegated discrete GPUs to the AI dustbin and be running 128GB to 512GB of delicious local RAM with vastly more RAM than any discrete GPU could dream of.
- throwaway314155 2y agoThat seems a tad dramatic. GPU's were widespread because of gaming, not AI. That the overlapping market would somehow just all magically have >3,000$ _and_ decide to switch to a non-standard, non-CUDA hardware solution in just 12 months is absurd.
- lvl155 2y agoFWIW GPUs still do not saturate PCIe lanes.
- smcleod 2y agoNo mention of how well it's claimed to perform with tool calling? The Gemma series of models has historically been pretty poor when it comes to coding and tool calling - two things that are very important to agentic systems, so it will be interesting to see how 3 does in this regard.
- PKop 2y agoI wasn't able to get function calls to work for Gemma3 in ollama, nor were others[0]. What is another way to run these models locally? [0] https://github.com/ollama/ollama/issues/9680 https://github.com/ollama/ollama/issues/9680 [1] https://github.com/ollama/ollama/issues/9680#issuecomment-2719702423 https://github.com/ollama/ollama/issues/9680#issuecomment-27...
- chaosprint 2y agoHow does this compare with qwq 32B?
- axiosgunnar 2y agoPSA: DO NOT USE OLLAMA FOR TESTING. Ollama silently (!!!) drops messages if the context window is exceeded (instead of, you know, just erroring? who in the world made this decision). The workaround until now was to (not use ollama or) make sure to only send a single message. But now they seem to silently truncate single messages as well, instead of erroring! (this explains the sibling comment where a user could not reproduce the results locally). Use LM Studio, llama.cpp, openrouter or anything else, but stay away from ollama!
- Technetium 2y agoI looked around to get confirmation, and I did find some related issues. Seems like it works properly when context is defined explicitly. There also appears to be a warning logged about "truncating input prompt", so it isn't an entirely silent failure. https://github.com/ollama/ollama/issues/2653 https://github.com/ollama/ollama/issues/2653 + https://github.com/ollama/ollama/issues/4967 https://github.com/ollama/ollama/issues/4967 + https://github.com/ollama/ollama/issues/7043 https://github.com/ollama/ollama/issues/7043 + https://github.com/ollama/ollama/issues/8144 https://github.com/ollama/ollama/issues/8144
- danielhanchen 2y agoI wrote a mini guide on running Gemma 3 at https://docs.unsloth.ai/basics/tutorial-how-to-run-gemma-3-effectively https://docs.unsloth.ai/basics/tutorial-how-to-run-gemma-3-e...! The recommended settings according to the Gemma team are: temperature = 0.95 top_p = 0.95 top_k = 64 Also beware of double BOS tokens! You can run my uploaded GGUFs with the recommended chat template and settings via ollama run hf.co/unsloth/gemma-3-27b-it-GGUF:Q4_K_M
- tarruda 2y agoThanks for this, but I'm still unable to reproduce the results from Google AI studio. I tried your version and when I ask it to create a tetris game in python, the resulting file has syntax errors. I see strange things like a space in the middle of a variable name/reference or weird spacing in the code output.
- vessenes 2y agoDaniel, as always, thanks for these. I had good results with your Q4_K_M quant on mac / llama.cpp. However, on Linux/A100/ollama, there is something very wrong with your Q8_0 quant. python code has indentation errors, missing close parens, quite a lot that's bad. I ran both with your suggested command lines, but of course could have been some mistake I made. I'm testing the bf16 on the A100 now to make sure it's not a hardware issue, but my gut is there's a model or ollama sampling problem here. EDIT: 27b size
- svachalek 2y ago
- deleted 2y ago[deleted]
- tarruda 2y agoMy usual non-scientific benchmark is asking it to implement the game Tetris in python, and then iterating with the LLM to fix/tweak it. My prompt to Gemma 27b (q4) on open webui + ollama: "Can you create the game tetris in python?" It immediately starts writing code. After the code is finished, I noticed something very strange, it starts a paragraph like this: " Key improvements and explanations: Clearer Code Structure: The code is now organized into a Tetris class, making it much more maintainable and readable. This is essential for any non-trivial game. " Followed by a bunch of fixes/improvements, as if this was not the first iteration of the script. I also notice a very obvious error: In the `if __name__ == '__main__':` block, it tries to instantiate a `Tetris` class, when the name of the class it created was "TetrisGame". Nevertheless, I try to run it and paste the `NameError: name 'Tetris' is not defined` error along with stack trace specifying the line. Gemma then gives me this response: "The error message "NameError: name 'Tetris' is not defined" means that the Python interpreter cannot find a class or function named Tetris. This usually happens when:" Then continues with a generic explanation with how to fix this error in arbitrary programs. It seems like it completely ignored the code it just wrote.
- whiplash451 2y agoWhy did this get downvoted? Asking genuinely
- whbrown 2y agoThose sound like the sort of issues which could be caused by your server silently truncating the middle of your prompts. By default, Ollama uses a context window size of 2048 tokens.
- tarruda 2y agoI checked this, the whole conversation was about 1000 tokens. I suspect the Ollama version might have wrong default settings, such as conversation delimiters. The experience of Gemma 3 in AI studio is completely different.
- tarruda 2y agoI ran the same prompt on google AI studio it had the same behavior of talking about improvements as if the code it wrote was not the first version. Other than that, the experience was completely different: - The game worked on first try - I iterated with the model making enhancements. The first version worked but didn't show scores, levels or next piece, so I asked it to implement those features. It then produced a new version which almost worked: The only problem was that levels were increasing whenever a piece fell, and I didn't notice any increase in falling speed. - So I reported the problems with level tracking and falling speed and it produced a new version which crashed immediately. I pasted the error and it was able to fix it in the next version - I kept iterating with the model, fixing issues until it finally produced a perfectly working tetris game which I played and eventually lost due to high falling speed. - As a final request, I asked it to port the latest working version of the game to JS/HTML with the implementation self contained in a file. It produced a broken implementation, but I was able to fix it after tweaking it a little bit. Gemma 3 27b on Google AI studio is easily one of the best LLMs I've used for coding. Unfortuantely I can't seem to reproduce the same results in ollama/open webui, even when running the full fp16 version.
- casey2 2y agocoalma3
- singularity2001 2y agoHow does it compare to OlympicCoder 7B [0] which allegedly beats Claude Sonnet 3.7 in the International Olympiad in Informatics [1] ? [0] https://huggingface.co/open-r1/OlympicCoder-7B?local-app=vllm https://huggingface.co/open-r1/OlympicCoder-7B?local-app=vll... [1] https://pbs.twimg.com/media/GlyjSTtXYAAR188?format=jpg&name=4096x4096 https://pbs.twimg.com/media/GlyjSTtXYAAR188?format=jpg&name=...
- deleted 2y ago[deleted]
- eogrok 2y ago[flagged]