Top
Best
New

Posted by logickkk1 9 hours ago

Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber(blog.google)
https://console.cloud.google.com/agent-platform/publishers/g...
565 points | 451 commentspage 3
youssefarizk 8 hours ago|
3.5-lite is the real showpiece here; agentic models of this size are a huge value-add for 90% of knowledge work agent tasks
brap 3 hours ago||
From my experience, this thing is crazy fast.

Spawn 10 on the same problem and have them debate to reach a consensus, you’ll get Fable-like results but 100x faster.

u1hcw9nx 7 hours ago||
Google has not changed. Following two facts are like tautologies by now.

1. Their AI efforts are very fundamental research oriented. They are really good at it.

2. Their productization sucks. The end products gets little attention compared to competition. It can be canceled at any time. You should never build anything around Google only APIs, AI or not.

Alifatisk 5 hours ago||
In other good news "the model has been trained to minimize refusals for beneficial uses.".

Otherwise, this news feels like a tiny incremental improvement on Gemini Flash series to make it more efficient with token usage, subagent and cost. Nothing big.

Regarding their benchmark scores on CyberGym, I wonder why they didn't compare their 3.5 Flash Cyber model with Fable 5. I mean they included Mythos and GPT-Cyber, so why not Fable 5 too?

They also mentioned Gemini 3.5 Pro is in testing and its about to become available very soon. Another thing maybe worth discussing is the announcement of pre-training Gemini 4. Sadly, not much technical details to discuss on. Many comments in here seem to mostly be about how Google is behind the others, but honestly, is it really worth the investment to be #1 in Artifical Analysis every week?

ConfusedDog 8 hours ago||
Why would 3.6 flash perform a little worse than 3.5 flash on Artificial Analysis Coding Index...

https://artificialanalysis.ai/models/gemini-3-6-flash?intell...

sosodev 8 hours ago||
Because AA Coding "Index" consists only of two benchmarks (Terminal-Bench v2.1, SciCode) and generally fails to be meaningfully representative of agentic coding capabilities.
firethunder7 3 hours ago|||
AA coding index has been updated to use DeepSWE, Terminal-Bench v2, and SWE-Atlas-QnA.
sosodev 1 hour ago||
When? It literally says on the page for Gemini 3.6 Flash "Artificial Analysis Coding Index represents the weighted average of coding benchmarks in the Artificial Analysis Intelligence Index (Terminal-Bench v2.1, SciCode)"
Alifatisk 8 hours ago|||
Whats a better option for AA Coding Index?
WASDx 7 hours ago||
DeepSWE and FrontierCode are more realistic if you read up on what they actually measure. But the most realistic is to try it yourself. Benchmarks can only vaguely represent typical usage, and how you judge the result. Giving the same real task you have to a few models will make you understand them better than chasing benchmarks.
kimjune01 34 minutes ago||
it would be nice if these benchmark reports actually specified which tasks they passed and which ones they didn't.
wmedrano 8 hours ago||
Could be a good tradeoff for the flash model though. 3.5 -> 3.6 is a tiny bit cheaper and maybe faster?

artificialanalysis.ai has it going from 165 tps -> 304 tps. openrouter.ai needs more data but it has it going from ~100 tps -> ~150 tps, though at peak 3.5 has reached 156tps.

waldrews 4 hours ago||
3.5 Flash-Lite seems available in US region, as was 3.5 Flash; but 3.6 Flash looks Global only so far when pinging. If Google employees are watching, will this issue go away?
jdthedisciple 3 hours ago||
Bottom line it looks about on equal footing with GLM 5.2 in terms of both overall intelligence and cost per task, while being significantly faster (in fact it is the fastest model on artificial analysis as of rn [0])

[0] https://artificialanalysis.ai/models/gemini-3-6-flash

ianberdin 6 hours ago||
Pelican svg and a near-perfect 3D MacBook at max effort for $0.16, about a fifth of Fable's price.

Fable 5 still wins on detail with no visible errors, but it's close. And this isn't a memorized pelican;

https://playcode.io/blog/macbook-svg-benchmark#gemini-3-6-fl...

WarmWash 6 hours ago|
I don't know if it's a rendering error since it looks like your site renders the SVG instead of hosting a static image of it, but the 3.6 macbook looks like an abstract art piece lol, both ff and chrome desktop
revolvingthrow 8 hours ago||
Tons of guardrails, lazy model, super confusing plans, expensive 3.5/3.6 flash and lite and 3.5 pro MiA?

Rough patch for google ai

zwaps 3 hours ago|
Here's the issue:

GLM 5.2 is better, also cheaper, and almost as fast.

So essentially, a big L for Google. Combine this with them not being able to produce a frontier model this generation... hmm implications

WarmWash 2 hours ago|
3.6 is roughly 50% faster, which isn't totally insignificant for being marginally more expensive.[1]

[1]artificialanalysis.ai

zwaps 2 hours ago||
Sure, but there's no sota alternative from Google. That's it, and its beaten by GLM 5.2 on every measure except somewhat speed.

I find that quite staggering. GLM is open weights

Narkov 1 hour ago||
Speed is definitely a marketable quality. All these things are a trade-off and solely measuring against SOTA I don't feel is always helpful.
More comments...