Top
Best
New

Posted by logickkk1 18 hours ago

Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber(blog.google)
https://console.cloud.google.com/agent-platform/publishers/g...
686 points | 522 commentspage 11
tiahura 18 hours ago|
3.5 Pro must really suck.
holistio 18 hours ago||
They are comparing against their own previous models instead of competitors. Not a great sign.
geooff_ 18 hours ago||
At this point just put the Pareto in the bag bruh
accountrequired 17 hours ago||
whatever, dude. give gemma5
raffael_de 17 hours ago||
is it just me or is this one-upping each other every few days getting ridiculous secreting a whiff of desperation?
JacobAsmuth 17 hours ago|
Just you. This is typical market competition in a fast moving field.
ChrisArchitect 18 hours ago||
Some more discussion:

Gemini 3.6 Flash https://news.ycombinator.com/item?id=48993130

lostmsu 8 hours ago||
3.6 Flash has the same performance on artificial analysis benchmarks as 3.5 Flash. So... what... is... the... point?..
lostmsu 8 hours ago|
Nevermind, I just realized it is point six!
llmslave 18 hours ago||
I keep saying this and people dont believe me, but I have b2b saas systems with actual agents running around the clock, and the performance/stability of the flash model is higher than most other models.

Meaning, its predictable with tool calls, wont spin off a million tools/do weird behavior, its reasonable. Even sonnet in a real world decision making scenario is not reliable, or will reason so long its incredibly expensive.

The benchmarks arent catching all the value, and most people have never actually ran an ai agent in a real context that matters

sureMan6 18 hours ago|
Who's most people? What are you talking about? Most people here use agents every day and I wouldn't trust flash or pro to touch any important project of mine because they're both terrible compared to the competition, waste of time every time I give them a chance
llmslave 17 hours ago||
I mean like an ai agent doing some sort of HR work, not a coding agent. Very few businesses are trusting an autonomous agent.
dismalaf 17 hours ago||
With all the naysayers on Gemini models I'm curious how many people actually use Gemini regularly?

For me, Gemini models are the most usable. Claude Opus and Mistral always try to turn queries into one-shot enormous commits, which just burns tokens, time and annoys me for something which is still wrong more often than not.

Gemini seems far better at listening to instructions and giving me what I actually want, on top of using far fewer tokens and wasting my time. Fable is the only model that's come close to Gemini Pro for me.

And as this is about Flash, it's exciting, I find Flash can usually get the right answer pretty quickly and without too much nonsense.

dudeinhawaii 9 hours ago|
I use all of the major providers daily and I tend to go to Gemini for "fast lookups" where a good enough answer is probably OK. I use ChatGPT and Claude for anything where it matters and generally when I invoke all three -- Gemini is the most surface level with responses, and also sycophantic.

It gets worse from there. Gemini is terrible at agentic coding, primarily because Agy is terrible. I noticed Google updated Agy with this release, so perhaps that's finally going in a good direction. I'll have to test it. Thus far, my experience in countless experiments has been Gemini models being 2x faster yet with less depth in their solutions and a lot more going off track.

I very rarely have to stop Codex or Claude Code sessions because they're doing something random and unexpected (or not asked for). I genuinely think Gemini models are brilliant but virtually useless in agentic scenarios in my experience of the last few years (2.5, 3, 3.1, 3.5).

I should also note that Gemini web UI annoying resets to its lowest intelligence which feels scummy and Google is not transparent about what "extended thinking" really is. Past posts have pointed to "extended" being medium. Every other provider gives you the raw value (medium/high/etc).

So honestly, I feel Google would rather I don't use their models. They just want to get a little bit of mindshare and stay in the conversation. I had the Ultra plan and cancelled it once it was apparent they were not improving the agentic experience nor trying to compete.

tobias2014 8 hours ago||
I agree, I see how Gemini itself with a usable harness can be excellent. But in agy with forced eager compaction (~125k with 3.1-pro, ~200k with flash) a kind of laziness and forgetting shows through that leads to an endless sequence of stopgap instructions, even with rigorous GEMINI.md and isolated task delegation and a good task tracking system. Agy is basically useless for more complex problems as far as I am concerned, at least when used somewhat autonomously as one could expect from claude. For strictly mechanical one-shot tasks it might be fine. I've spent way too much time working around these limitations instead of just continuing to use claude. Hoping that things would have improved with flash 3.6 I feel that it's actually worse in following instructions, and always acts even when just asked a question. If just agy offered a better experience and got rid of the terrible forced automatic eager compaction.

PS: That opus-4.6 via agy works so much better points in another direction though!

fur-tea-laser 18 hours ago|
not a google fanboy by any stretch... though i've been thrilled with the flash line of models... i exclusively use it on high, and have found it to be a great fit for increasing productivity 10-fold while maintaining quality... sure it can't just go off and one-shot a bunch of work, but at the complexity level i tend to work at, neither can the frontier in a robust way that i can be confident in... sure i have to be in the loop more, but that helps keep me grounded and course-correct earlier before wasting tokens... and when you sufficiently spec out a coding/software problem, and i mean really document all of the critical nuance, it will successfully satisfy the constraints... the quality is rarely acceptable on first-pass, but it forces me to stay connected to the architecture more than i would be if using a frontier model... i've found this to be a happy middle-ground of productivity and awareness...
More comments...