Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • 3.5 Flash was more expensive than 3.1 Pro to run the Artifical Analysis test suite. $1551 for 3.5 Flash [0] vs $892 for 3.1 Pro [1]. That's 74% more cost while ranking lower. It's 2.5x as fast but I don't think the bang for the buck is there anymore like it was with 3.0 Flash. I'm a bit bummed out to be honest.

    I did not expect such a huge (3x) price increase from 3.0 Flash and I bet many people will not just blindly upgrade as the value proposition is widely different.

    One interesting point to note is that Google marked the model as Stable in contrast to nearly everything else being perpetually set as Preview.

    [0] https://artificialanalysis.ai/models/gemini-3-5-flash [1] https://artificialanalysis.ai/models/gemini-3-1-pro-preview

    by eis
  • Gemini 3.5 Flash's 2000 token clocks aren't bad. https://clocks.brianmoore.com/
  • Click on "Listen to article", make sure the voice is "Umbriel" and skip to 4:15 - there's a hallucinated part at the end in Russian (I think). On a blog post about the latest and greatest AI model. Oh the irony.
  • Am I really so old that when someone says "Flash" my immediate response is... "consider HTML5 instead" ??
  •   > Create animated SVG of a frog on a boat rowing through jungle river. Single page self contained HTML page with SVG
    
    3.5 Flash: Thinking Medium - 7516 tokens

    https://gistpreview.github.io/?5c9858fd2057e678b55d563d9bff0...

    3.5 Flash: Thinking High - 7280 tokens

    https://gistpreview.github.io/?1cab3d70064349d08cf5952cdc165...

    3.1 Pro - 28,258 tokens

    https://gistpreview.github.io/?6bf3da2f80487608b9525bce53018...

    Though 3.1 took 3 minutes of thinking to generate, but it only one that got animated movement.

    by SXX
  • Per million input/output tokens:

    Gemini 2.5 flash: $0.30/$2.50

    Gemini 3.0 flash preview: $0.50/$3.00

    Gemini 3.5 flash: $1.50/$9.00

    Interesting pricing direction. I don't think we have ever seen a 3x price increase for in the immediate next same-sized model (and lol @ 3 only ever getting a preview).

    3.5 flash costs similar to Gemini 2.5 pro which was $1.25/$10

  • The pelican is a lot: https://github.com/simonw/llm-gemini/issues/133#issuecomment...

    Not a great bicycle though, it forgot the bar between the pedals and the back wheel and weirdly tangled the other bars.

    Expensive too - that pelican cost 13 cents: https://www.llm-prices.com/#it=11&ot=14403&sel=gemini-3.5-fl...

  • For those who would like to know the total and active parameter count of this model: even though Google doesn't disclose the model technicals, we can infer them within relatively tight margins based on what we do know.

    We know they serve the model on TPU 8i, which we have plenty of hard specs for (so we know the key constraints: total memory and bandwidth and compute flops). We can also set a ceiling on the compute complexity and memory demand of the model based on knowing they will be at least as efficient as what is disclosed in the Deepseek V4 Technical Report.

    We can also assume that the model was explicitly built to run efficiently in a RadixAttention style batched serving scenario on a single TPU 8i (so no tensor parallelism, etc. to avoid unnecessary overheads... Google explicitly designed the 8th-generation inference architecture to eliminate the need for tensor sharding on mid-sized models).

    We know Google intends to serve this model at a floor speed of around 280 tok/s too.

    Putting all these pieces together, we can confidently say this model is ~250-300B total, and 10-16B active parameters. Likely mostly FP4 with FP8 where it matters most.

    Visual:

      ┌────────────────────────────────────────────────────────┐
      │                   TPU 8i VRAM (288 GB)                 │
      ├───────────────────────────┬────────────────────────────┤
      │   Static Model Weights    │  Dynamic Allocations &     │
      │   (250B - 300B @ Mixed    │  Compressed KV Caches      │
      │   FP4/FP8)                │  (RadixAttention / SRAM)   │
      │   ~110 GB - 150 GB        │  ~138 GB - 178 GB          │
      └───────────────────────────┴────────────────────────────┘
    
    I do model serving optimization work. This is napkin math.

    Edit: There's one factor I under-rated in my initial estimate... TurboQuant. This is a compute to KV memory use tradeoff. It's plausible with TurboQuant at a quality-neutral setting they've gotten the model up to 400B with similar economics. This is a variable effecting concurrency and the the way they decided total model size was likely based on what they see for the average user's average KV cache depth in real-world usage.

Explore Birbla archives