8 ms·
They need to release benchmarks against Luna/Terra. Luna is much cheaper which feels like it undercuts the need for Flash. I've always considered the Flash ser
by wxw 1mo ago
They need to release benchmarks against Luna/Terra. Luna is much cheaper which feels like it undercuts the need for Flash.
I've always considered the Flash series of models to be for low-cost, high-volume, mostly text-based use cases (e.g. summarization, parsing, formatting), emphasis on low-cost.
[edit: ah, benchmarks here: https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash/ https://blog.google/innovation-and-ai/models-and-research/ge...
more of a Terra than Luna competitor which is an interesting positioning. I feel like differentiation at the mid-tier of models is pretty difficult.]
- timdorr 1mo agoThey compared against 5.6-terra on the model card: https://deepmind.google/models/model-cards/gemini-3-7-flash/ https://deepmind.google/models/model-cards/gemini-3-7-flash/
- deleted 1mo ago[deleted]
- peab 1mo agogemini flash is probably the best model for visual tasks right now. they also make it really easy to ingest videos
- pants2 1mo agoYes, was going to say I use it exclusively for video and audio. The ability to give it a YouTube link through the API and ask questions about it is awesome
- wxw 1mo agoAh, multimodal is a great point. I'll need to try that some time.
- icelancer 1mo agoCrazy it's still the only video understanding endpoint. It's what I use it for and no other model even offers a competitor.
- mike_hearn 1mo agoYou probably can't build such a model without unlimited access to YouTube and Google has been tightening the screws on that over the years pretty systematically.
- peab 1mo agoto be fair, all it's doing is sampling the frames and maybe doing transcription, if I'm not mistaken. So you can do it with the other models too, you just need to sample the frames yourself and do the transcript yourself
- icelancer 1mo agoit does this at a variable rate of frames which you can set - not sure if it is transcribing or natively understanding audio, but I think it's the latter since it is much faster than most transcription models I am aware of Regardless you are right - I can roll my own.... but why
- ipsod 1mo agoAlso best at OpenSCAD, seemingly for the same reason, at least in terms of "iterate on a design, comparing visual output to target".
- sourweasel 1mo agoAre you manually rendering previews of its OpenSCAD output to create images for it to review, or do you have a workflow that automates that?
- gunalx 1mo agoPersonally both, i use openscad with opencode, and paste images but it is pretty often the modell decides by itself it wants to see a render and uses the render shell commands to get a image to look at.
- ipsod 1mo agoI've built some tools using AI. A custom GUI that has "copy context" and "copy image" buttons, to make prompting easy, and the camera position is persisted to disk on each change so that the agent can run a command to get a screenshot of it. I'm thinking about inlining an AI chat window directly - I guess I might fire off a prompt to do that right now.
- gunalx 1mo ago+1 gemini models where really the only ones fullt grasping spatial reasoning even compared to opus (at least when i last cared to check it)
- nl 1mo agoI've found Sol excellent at OpenSCAD. But I don't do "compare visual output to target" as much as 3D reasoning type tasks (eg: "Build a G1 curve where the -X face meets the +Z face" etc)
- ipsod 1mo ago"G1 curve"? Nice, cool to have new vocabulary for the spell book. Thanks.
- pimeys 1mo agoIt is also very good and cheap for computer use.
- andai 1mo agoMatched roughly with Sol on DeepSwe cost per task. Luna way cheaper. DeepSeek used to be, but I think it's somewhere on Sol's curve after the price hike.
- scotty79 1mo agoOn DeepSwe it's strictly beaten by Luna on max, cost and result. Damn, Luna on max is as good on DeepSWE as Kimi k3, I think I dismissed this model unjustly.
- HDBaseT 1mo agoThis is why benchmarks are scary, since Artificial Analysis puts it a fair bit behind Kimi K3. Kimi K3 is a beast though, just costly.
- dannyw 1mo agoStill cheaper than API rates for Opus!
- barrenko 1mo agoSo in a way Luna is the new Gemini Flash? I've been out of the game for a while.
- scotty79 1mo agoLuna is the smallest variant of gpt-5.6 from openai and it seems to beat Gemini Flash 3.7 on DeepSWE benchmark (which is one of the new coding benchmarks that people find more relevant to their daily work than old, saturated and gambled benchmarks). It beats it massively on cost and by a bit on quality.
- anthonypasq 1mo agoflash-lite is more of their luna tier competitor but even still not quite there yet, but gemini's dominance on multimodal and image understanding i think really gets downplayed on this site when most people think the only think you can do with LLMs is write code
- vohk 1mo agoThat's been my association as well. I see Flash get brought up a lot in relation to things like OCR and PDF processing frequently, and a lot of other routine multimodal workloads.
- deleted 1mo ago[deleted]
- MrBuddyCasino 1mo agoYes this is my impression as well. To be fair I didn't compare to Luna yet, but Gemini 3.5 Lite is a very good and cheap multi-modal data extraction model.
- dogomatic 1mo agoIf anyone knows of a cheaper vision llm with the same accuracy I would love to switch
- WarmWash 1mo agoUltimately it would track that in the real world, people will want to point cameras at things and get answers. I pay for ChatGPT and Gemini, and while Sol is a total beast with anything text, it still poisoned my cucumber bed. Which I will be bitter about for at least a few years while the bed recovers. Gemini (even flash) is exceptionally talented at viewing photos and telling you what to do/what it is (and telling me I just misidentified the problem with my cucumbers and spraying off the "bugs" actually just spread the bacteria everywhere.)
- rolisz 1mo agoTell me more about your cucumber bed. We started some raised veggie beds this year and my wife is relying very heavily in ChatGPT and Claude for advice on how to deal with issues.
- jeffbee 1mo agoWhy is Gemini represented by points on this cost-quality plane, while competitor's models are represented by curves?
- scotty79 1mo agoCompetitors release multpile models and their curves reflect reasoning effort of each single model. Gemini doesn't have adjustable reasoning effort (at least on the graph) so each of its curves is just one point.
- jeffbee 1mo agoGemini has levels though
- denalii 1mo agoFor what it's worth the source of the data[0] does have 3.7 flash with all 3 reasoning levels. 3.5/3.6 are in fact just the single points though (high reasoning). The datapoint in the announcement screenshot is either med or high, but they're pretty much exactly the same so can't say for certain. [0]: https://deepswe.datacurve.ai/ https://deepswe.datacurve.ai/
- jeffbee 1mo agoIt is extremely odd that this model has that crowbar, spending more for worse results at "high".
- marcuskaz 1mo agoThat graph has to be made because an Exec didn't like that graph went down to the right instead of up and to the right. How do you make a graph with 0 on the far right and counts up by going left of 0? What number line is that?