Top
Best
New

Posted by bratao 14 hours ago

Gemini 3.8 Flash and 3.8 Flash Cyber(blog.google)
https://deepmind.google/models/model-cards/gemini-3-8-flash/
918 points | 525 commentspage 2
abixb 11 hours ago|
I like Google's strategy here. These new Flash models of late (Flash 3.6, 3.7 and now 3.8) have obviously been distilled from a much larger unreleased model (Gemini 3.5 Pro, iirc from the rumors).

One aspect of model releases that don't get discussed as much are the cache invalidation (changes in underlying architecture, weights, or tokenizers); I assess Google seems to be squeezing the maximum out of the last 'Pro' version they released with 3.1 back in February.

Small models cataching up with their bigger siblings are fantastic news.

alephnerd 11 hours ago|
A couple larger GCP customers requested this for sometime, especially on the cybersecurity side.

A SOC/IR or AppSec team doesn't need a generalized model that knows when Chaucer lived but it absolutely needs a model that can efficiently, quickly, and accurately prioritize vulnerability severity or validate patches.

kamranjon 13 hours ago||
They've interestingly left out any mention of speed.

I have been testing 3.7 flash against 3.5 flash and it seems to lose every time in overall latency. Every benchmark I've seen seems to suggest the opposite[1] - that 3.7 flash is significantly (at times 2x) faster than 3.5 flash - but I have never been able to prove this out in real world use cases.

Has anyone found their latency numbers to actually be accurate? Is this why they've toned it down in this release? For context, I'm testing larger generation payloads that take 8-10 seconds in 3.5 flash and 15-25 seconds in 3.7 flash. Lowest reasoning settings in both cases.

1: https://artificialanalysis.ai/?speed=intelligence-vs-speed&m...

pampas 6 hours ago||
In my niche Redactle puzzle solving benchmark [1] I noticed Gemini 3.8 flash is slightly faster than 3.7 flash. They both smoke every model I've tested. I have not yet run 3.5 flash. Gemini models are great at this task because they seem to have exact Wikipedia text baked into the weights. When I rewrite the wiki text a bit it's not able to one-shot the game so much.

[1]: https://redactle.net/llm-leaderboard

film42 13 hours ago||
It depends on how you're querying Gemini models. OpenRouter is the fastest by far. I'm guessing they bought the dedicated pipe from Google. Gemini via VertexAI and consumer API has pretty bad latency.
kamranjon 13 hours ago||
Yea I am testing through OpenRouter - have you noticed 3.7 flash being significantly faster?
film42 11 hours ago||
I guess it might be relative, but switching from VertexAI endpoint to OpenRouter was like 2-3x faster for us.
andai 14 hours ago||
Wait, I didn't realize 3.7 Flash was already beating Sol on a bunch of the benchmarks. Isn't it a way smaller models?
ipsod 14 hours ago||
IDK if it's smaller, but I know it's way faster. In one test I did, Flash 3.7 high was ~9.4x faster than Luna High.

But, also... Sol crushes Flash 3.7 at writing code in a codebase of any size beyond "tiny".

Flash is my go-to for prototyping, and basically anything that isn't writing production code.

ramon156 14 hours ago|||
The only company with a proper TPU set-up is bound to have the fast models, now add a market cap like Google to the mix.
ipsod 14 hours ago||
They've been my bet to win the AI race for a while. I was starting to doubt, but this 3.6, 3.7, and 3.8 arc has anchored me.
Diederich 12 hours ago||
Totally agree. When this current wave of GenAI really started heating up, I guess 2020-2021, my analysis was very straightforward. What are the high level inputs to long-term success? I basically came up with a couple of criteria:

1. Data. Lots of data.

2. Money. Lots of money.

3. Access to necessary hardware.

4. Business alignment/will to do it.

5. Access to talent, current and future.

This is certainly incomplete/naive. In my mind, though, Google was the clear answer.

On a more personal level, I've been deep into the Google ecosystem since I got diederich@gmail.com in 2005. (I actually paid 50 cents on ebay to get a very early invite.) There was no question in my mind that Google's AI work would deeply integrate into their whole ecosystem in very powerful and productive ways. (Yes, I can join you to discuss, at length, the various ways that Google's dominance is problematic/scary.)

Having said all that, I'm quite happy that there is, at the moment, a very rich competitive landscape. Indeed, not too long ago, with Gemini Pro 3.1 languishing, I moved most of my deeper thinking work to ChatGPT, which was, for me at least, clearly outperforming Gemini.

While I certainly didn't anticipate it, Google's strategy of making their fast/relatively inexpensive models surprisingly powerful has been a welcomed surprise.

ipsod 11 hours ago||
Not only do they have access to the hardware - they've been developing it in-house for years.
momojo 12 hours ago||||
Same. Love oneshotting or sanity checks. Which fortunately is a lot of my workflow (lot of long tail stuff fits in one prompt).
esafak 14 hours ago||||
Luna is way slow. I don't remember an OpenAI model ever being this slow.

edit: I have a subscription; direct call.

dannyw 14 hours ago||
Are you using direct or via OpenRouter? I think OpenRouter Luna always uses the `flex` tier, which is quite a bit slower.
MrBuddyCasino 11 hours ago|||
Its not good at not making mistakes, but what it produces is structurally quite nice, not over-engineered (looking at you Sol) and its personality isn’t annoying (looking at you Claude). A bit like Grok Code, but Grok is a better coder.
realist_not 14 hours ago|||
It's pretty good if you can actively steer it , its actually really really good , the antigravity free tier and pro tiers are generous as well . I'm shocked at how fast it generates tokens.
worldsavior 14 hours ago|||
Some would say it's Google's TPUs.
MrBuddyCasino 11 hours ago|||
Can the free tier be used outside Antigravity CLI? Because its security prompts get old pretty quick.
pampas 7 hours ago|||
Gemini 3.7 Flash was already smashing more expensive models on my Redactle benchmark https://redactle.net/llm-leaderboard which mostly tests omniscience.
refulgentis 14 hours ago||
They're quite selective in benchmarks, c.f. notably only bad one is 10% on TerminalBench. It's a really addled model, one time I said "Hi" and it built out a 4 panel hello world app with (fake) weather, a todo list, and a couple other things I forgot. I wouldn't be comfortable saying "ignore the #s!" except when I complained it was trash and way overcooked on agentic coding yet not good at it, and a couple DeepMind ML people liked the tweet.
zuzululu 8 hours ago||
i discount people who lean too heavily into benchmark as the authoritative truth when it comes to evaluation of coding capability of these models.

experience tells me that those people simply have not used models for a long period of time specifically on coding and have run their own comparisons

to someone who uses all vendors, the differences are very palpable and drives purchase decisions.

also keep in mind Gemini and other labs have repeatedly done benchmaxxing, you must have your own benchmarks to evaluate these models.

j-bu 13 hours ago||
"The knowledge cutoff date for Gemini 3.8 Flash is March 2026 – users can expect updated information for some domains while in others they may experience the model’s knowledge is limited to January 2025 (in line with the Gemini 3 Model Family)."

Kind of wild that they haven't (successfully) pretrained a base model since Jan-25.

venusenvy47 13 hours ago||
I'm curious if the knowledge cutoff is important, when the interface (Gemini app) can search online for recent information. Is there a big advantage to having everything internal?
j-bu 13 hours ago|||
Not directly - but latest research advancements, cleaner / richer datasets, etc. still require fresh base models. Not everything can be fixed through post training alone (e.g. why GPT-5.5 "Spud" was such a big jump, and also why GPT-6 "Astra" is now supposedly another big leap). Ofc model size etc also plays a role, but my (admittedly limited) understanding is that new base models _can_ also lead to big jumps even keeping parameter counts constant.
StevenWaterman 9 hours ago||||
You don't need everything internal, but having some idea of recent events is useful. If you ask it to implement some local AI there's a decent chance it will try to use qwen 2.5 without wondering if anything better came out since
rjh29 10 hours ago||||
Search grounding is expensive, you can't force the model to do it either. I use Gemini a lot and it often replies with out-dated data. The more detailed the information you're asking, the more likely it is to be wrong.
npn 11 hours ago||||
very important actually. just try to generate code for fresher frameworks/libraries. gemini sucks so bad in real work usage, everything it suggests are outdated and mostly useless.
neuronic 7 hours ago|||
[dead]
make3 2 hours ago||
That extremely likely just means that they're preparing an omega huge Gemini 4 Pro release and that that's what training right now on most of the compute
raincole 13 hours ago||
I don't know if Google is having the worst marketing fumble or the most genius marketing one. Their "flash" models are very comparable to other companies' "pro" or "flagship" models. It seems to be a quite counterintuitive naming convention as it undersells the models.

Unless they have an even more powerful Gemini Pro in the oven...?

anthonypasq 12 hours ago||
the 3.5 pro pretrain was a complete disaster, they shelved it and are now working on gemini 4.

3.0 flash -> 3.8 flash is all post training which is pretty impressive.

chrsw 7 hours ago||
Do labs come back from disasters like GDM’s 3.5 pretrain? I am thinking of Meta’s Llama 4. Meta is just now starting to be taken seriously again but they are definitely not at the frontier. And when I say “come back” I mean have an Opus 4.5 moment, which was really mind blowing for me at the time. Fable was a similar leap, just not as big.
deaux 3 hours ago||
OpenAI had such a disaster themselves before, GPT-4, so they replaced it with 4o.
rahidz 11 hours ago|||
The conspiracy theorist in me wonders if it's about keeping the federal government out of their business after seeing what happened to Sol & Mythos.
owaiswiz 12 hours ago|||
not saying they do have a beefier pro, but even if they did, isn't the delta between flash vs pro models reduced quite a bit? (e.g glm 5.3 flash vs 5.3, v4 flash vs v4 pro, sonnet 5 vs opus 5)?
drowntoge 13 hours ago|||
Well if that's the case, it's been in the oven for quite a while now.
make3 2 hours ago||
My assumption is that they're cooking an ultra humongous Gemini 4 Pro release. They certainly have the cash and the compute for it, and it's so obviously the thing to do from a strategic perspective.
xnx 14 hours ago||
Seem like a great, no-compromise, upgrade over 3.7 which is already a bargain, fast, and doesn't have the brain-damaged writing style of Claude.
fitsumbelay 14 hours ago|
that's certainly what it's looking like so far. kind of mind boggling ...
lysecret 11 hours ago||
Also just want to let my appreciation here for 3.7 it’s cheap super fast super reliable incredible at information parsing eu host able (important for us) and perfectly integrated into gcp. Great job google!
MrBuddyCasino 11 hours ago|
I hope they bring a lite version, its good enough for information parsing and very cheap.
weird-eye-issue 2 hours ago||
Use Luna for that
tagalog 2 hours ago||
Gemini flash seems to have been a bit of a sleeper. Somehow it's ended up as the most used LLM for my client document extraction work these past few months.

I have an eval harness that runs every Thursday to determine which models are the current best for a few different client workflows. And since May(?) flash has slowly been taking over more and more stuff to the point it is now 100% on 8 out of 11 document extraction flows with the other 3 being a Flash / Opus 4.8 mix for high value stuff where cost is less of a factor.

meh2frdf 14 hours ago||
The flash models, for coding are reckless in my experience. I have a Ultimate subscription, get good quota, but still use Opus 4.6 as it's much more reliable if you manage the context window carefully.
datlife 14 hours ago||
I use Flash model as code implementation executor, then have GPT-5.6-Sol or Opus to review the work. Pretty good so far and presumably less expensive.
upcoming-sesame 14 hours ago|||
If by reckless you mean commit, push, deploy without me asking it to, the I agree!
tiborsaas 14 hours ago|||
It even took my girlfriend on a date, now it prepares for IPO, how do I turn it off?
Ridius 13 hours ago||
Just hand over your clothes, your boots and your motorcycle and it'll be on it's way
okdood64 14 hours ago||||
Respectfully: If it's able to deploy without you asking it to, that's a you problem. There are no safeguards?
wongarsu 13 hours ago|||
That's exactly how you get 'you are right, I deleted the production DB to apply the new schema when I should have written a migration'

That said, I do trust Opus and Fable enough to let them deploy to staging. Great for debugging. Just don't give them keys for prod

meh2frdf 14 hours ago||||
You need more safeguards for sure, but also it tends to fly off down rabbit holes, rebuilding things in dumb ways, hacking around things, making assumptions etc, it seems very eager to go 'ta da! I did it look how quick I was', sometimes it nails it other times it created a lot of tech debt.
meh2frdf 14 hours ago||||
Also if it ever says, "I've found the root cause of ..", it definitely has not found the root cause and is making a non evidence based guess as it has run out of ideas.
upcoming-sesame 12 hours ago||||
100% my problem, but it's the only model I tried that does that so recklessly
iAMkenough 13 hours ago|||
I told it “don’t betray me” in my prompt and it still stabbed me in the back.
sejje 11 hours ago||
I'm terrified to seed the RNG with words like "betray". I'll keep those way, way down the list of likely words.
kyrra 12 hours ago|||
Agents.md is a thing, you can ask it to not do that (it follows that ask pretty well).
onlyrealcuzzo 14 hours ago|||
> The flash models, for coding are reckless in my experience.

My experience is that antigravity is awful and reckless - but that the model itself isn't.

agluszak 10 hours ago||
[dead]
throw10920 3 hours ago|
We've gotten an unusually fast speed of Gemini Flash releases over the past few months. Is this Recursive Self Improvement, or Google just trying to distract from the fact that it's been a while since the last Gemini Pro release?
tkgally 3 hours ago|
The blog post says it is RSI: “both of today's releases are … accelerated by long-running agentic loops designed to recursively evaluate and refine the underlying models.”
throw10920 3 hours ago||
Yeah, but Google is incentivized to claim that regardless of truth value. Critical analysis is necessary.
More comments...