

Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- Yikes. I think the concept of a 'flash' model is changing, no? Google used to market this as its lower-intelligence, faster, cheaper option. I appreciate that they are delivering on both of those, but personally I would appreciate if they could create an incremental knowledge improvement while holding price steady. Fortune 500 companies have to make their money I guess.by s3p
- That would be Flash Lite now, and I'm also interested in the cheaper end of things so kinda disappointed they didn't release 3.5 Flash Lite at the same time...by toraway
- My guess is Gemini Pro coming later will be 2x more, bringing it comparable to Opus’s pricing.by likium
- Real smart. I’ve come to associate ”Flash” with ”useless make-shit-up”, and always look for Thinking/Pro when I see it set. Now, suddenly, there is only Flash?by kilpikaarna
- I think flash just means "fast" nowby 2001zhaozhao
- On my Agentic SQL benchmark it scores 19/25. That's... mediocre.
It means performs worse than 3.1 Flash Lite Preview (22/25), is slower (367s vs 142s) and is more expensive (75c vs 2c).
It is outperformed by Gemma4 26B-A4B in every way(!)
https://sql-benchmark.nicklothian.com/?highlight=google_gemi...
(Switch to the cost vs performance chart to see how far this is off the Pareto frontier)
by nl - I'm seeing this too.
I have a SQL agent and my tests with 3.5 are resulting in hitting query budget limits that have never been hit before. On average, to answer the same question, 3.5 is spending 10x more on SQL queries vs gemini-3-flash-preview.
The query patterns can be extremely degenerate too. E.g. the agent will hit the semantic layer tool to pull the schema, then run `SELECT * FROM table LIMIT 1`, which hits the query budget limit and fails.
I've only really been looking this morning, so I need to do a full eval, but the initial results match what your benchmark shows.
---
Side note: your benchmark has an issue. On Q1 medium the model returned gross margin of 0.127 instead of 12.7 (%), and the benchmark failed it. The failures on Q9 and Q21 are the same (I didn't check other questions). Nowhere in the prompt did you specify you wanted the values converted to percentage points and rounded.
If you asked me to write that SQL with that prompt, unless you were throwing it directly into a visualization I would format it the same way gemini-flash did. If I were pulling into a spreadsheet or vis tool this format is preferable because it's easier to format in a client application.
The other failures like Q21 incorrectly averaging the list price are correct failures.
by data-ottawa - Knowledge cutoff: January 2025
Latest update: May 2026
I have a very bad feeling about this lag.
by reconnecting - I thought that was a choice that Google made?
- you really shouldn't have them pulling facts from their weights, they need grounding from real data sourcesby verdverm
- Can you explain what you mean?by hosel
- At least in some cases, there seems to be a move toward training on more synthetic data and strictly curated data, especially for smaller models where knowledge can't be extremely broad, because there just isn't enough room to store the world in tens or hundreds of gigabytes of model weights. So, to achieve higher quality reasoning, the training has to be focused and the data has to be very high quality and high density.
With strong tool use, it maybe doesn't even matter that the models are using older data. They can search for updated information. Though most models currently don't, without a little nudge in that direction.
Also, I believe the Qwen 3 series are all based on the same base model, with just fine-tuning/post-training to improve them on various metrics. Maybe everything in the Gemini 3 series is the same, and maybe they're concurrently training the Gemini 4 base model with updated knowledge as we speak.
by SwellJoe - I have google ai pro plan and tried antigravity with 3.5 flash but it used up all my quota in two prompts. If that is not a bug then it is seriously unusable.by hmate9
- The web version went from 100 Pro Prompts per day to...12 per 5 hours lol. I just did 3 back and forth not even technical planning for an infra project and I am ~25% thorough. Insane.by abeindoria
- I'm seeing this too.
API price for gemini-3.5-flash is 3x gemini-3-flash-preview so they might be throttling it 3x sooner. They should either drop API prices or not advertise AI Pro as supporting Antigravity.
https://ai.google.dev/gemini-api/docs/pricing#gemini-3.5-fla...
by babl-yc - The way they're charging for failed generations is brutal.
Checked my 5 hour quota, it was 0%, got this for multiple attempts:
I'm getting more image requests than usual, so I can't create that for you right now. Please try again later.
or
Can you ask me again later? I'm being asked to create more images than usual, so I can't do that for you right now.
Went back and found they took 34% of my quota for the privilege of repeating that same error.
I think the "Usage Limits" screen is new so who knows how long they've been counting errors against our quota. I guess I should be grateful it's now visible.
by cube00 - Yesterday, or the day before, Google lowered the AI Pro quota from 33x standard usage to 4x.
From the talk on the Gemini subreddit it's severely lower than before. I'm likely canceling my AI Pro.
The update also broke the app for me. Editing a message crashes the app every time. I'm on a Pixel lol
by quirino - 3.5 Flash was more expensive than 3.1 Pro to run the Artifical Analysis test suite. $1551 for 3.5 Flash [0] vs $892 for 3.1 Pro [1]. That's 74% more cost while ranking lower. It's 2.5x as fast but I don't think the bang for the buck is there anymore like it was with 3.0 Flash. I'm a bit bummed out to be honest.
I did not expect such a huge (3x) price increase from 3.0 Flash and I bet many people will not just blindly upgrade as the value proposition is widely different.
One interesting point to note is that Google marked the model as Stable in contrast to nearly everything else being perpetually set as Preview.
[0] https://artificialanalysis.ai/models/gemini-3-5-flash [1] https://artificialanalysis.ai/models/gemini-3-1-pro-preview
by eis - That's what I came here to check. Last model release they only put it into preview[0] at first.
Does that mean this model is production ready?
by mijoharas - >3.5 Flash was more expensive than 3.1 Pro to run the Artifical Analysis test suite
That's everything I needed to know.
by ls_stats - How do they calculate that?
3.1 has 57M output tokens from Intelligence Index, 3.5 Flash has 73M, so not a lot more, and 3.5 is a bit cheaper, I don't get how 3.5 can be 74% more expensive.
by pingou - Seems like the only good thing about 3.5 Flash is its speed. Not cost-competitive or benchmark-leading by any means.by ekojs
- Ouch. That's going in completely the wrong direction.
How many people complain that we have too much low quality AI output for humans to read, let alone evaluate vs. how many people are complaining that they want higher quality, more trustworthy output?
by hedora - Gemini 3.5 Flash's 2000 token clocks aren't bad. https://clocks.brianmoore.com/by lanewinfield
- Fascinating, kimi k2 has good clock too from my limited time being on the site.by acters
- From looking at all of them, it actually seems to be the best one, followed by Deepseek 3.1. And something went wrong with GPT-5's.by Valakas_
- Click on "Listen to article", make sure the voice is "Umbriel" and skip to 4:15 - there's a hallucinated part at the end in Russian (I think). On a blog post about the latest and greatest AI model. Oh the irony.by lmazgon
- I ran it through speech-to-text and it starts with something among the lines of "dear colleagues, just like a doctor tells a patient 'health can wait'...", after that it's nonsense.
I don't know if what the doctor said is some kind of idiomatic expression, but appears to be the opposite of sound medical advice. :)
by Tade0 - Thank you for this gem.by luk4
- Looks like they removed the option to "listen to article". I wonder why.by marknutter