Top
Best
New

Posted by bratao 15 hours ago

Gemini 3.8 Flash and 3.8 Flash Cyber(blog.google)
https://deepmind.google/models/model-cards/gemini-3-8-flash/
944 points | 533 commentspage 3
throw10920 4 hours ago|
We've gotten an unusually fast speed of Gemini Flash releases over the past few months. Is this Recursive Self Improvement, or Google just trying to distract from the fact that it's been a while since the last Gemini Pro release?
tkgally 4 hours ago|
The blog post says it is RSI: “both of today's releases are … accelerated by long-running agentic loops designed to recursively evaluate and refine the underlying models.”
throw10920 4 hours ago||
Yeah, but Google is incentivized to claim that regardless of truth value. Critical analysis is necessary.
throwa356262 14 hours ago||

    "available to trusted defenders through our new Fairwind Program"

Then why even bother announcing this? Ordinary people can use K3 and GLM 5.3 or whatever drops next and avoid all this hassle.
JacobAsmuth 12 hours ago||
You're telling me for only 5x the cost and 1/10th the speed I can use a Chinese model which performs worse than Gemini 3.8 Cyber? And I get to do all the hosting and setup work myself instead of just using a model and framework which is already integrated with GCP? Dang!
129867 12 hours ago||
I'm sorry, is this a bot that is optimized for sealioning? The point is that you don't have access to Cyber.
uif124 12 hours ago||
Agreed. The only legitimate use case is restricted to a secret guild. Imagine:

"Valgrind is only available to trusted defenders in our new UnfairAdvantage program"

pampas 10 hours ago||
Gemini 3.8 Flash is top of the Redactle LLM benchmark but so was Gemini 3.7 Flash. Both one shot all puzzles in the evals though 3.8 is just a bit faster. It also does the evals cheaper and faster than almost all the other models I've tried.

https://redactle.net/llm-leaderboard

leopoldj 14 hours ago||
Blog: https://blog.google/innovation-and-ai/models-and-research/ge...
adbachman 14 hours ago||
Still zero on the felony bench.

Is this weakness in their training regimen the impact of operating under regulatory frameworks for too long?

hmate9 14 hours ago||
It is more expensive per task than 5.6-sol high: https://artificialanalysis.ai/models/gemini-3-8-flash#price-...
HJain13 14 hours ago||
Cheaper at medium level while still being same score as Sol medium
radicalriddler 14 hours ago|||
Huh, according to some of those charts, it's both dumber, and more expensive to run against their benchmarking tasks than Fable??? Seems crazy to me.
sejje 12 hours ago||
Perhaps the model is able to evaluate that it's not done, and to keep pressing on in the face of mounting failures, until it eventually arrives at a solution. Where Fable can skip that.
jdthedisciple 13 hours ago||
Sol is still underrated imo, especially for the current discounted price
gere 13 hours ago||
I have mixed feelings about Gemini 3.7 Flash. I used it for a personal project in Java and it was ok: it was crazy fast and it reached the correct result, but the code quality was barely passable.

I also used it for a an app for my Garmin watch, and it wasn't good. The code was compiling, but functionality was totally broken and even with a lot of steering it wasn't able to make it work. GLM 5.3-flash instead was up for it and the code wasn't bad at all. I am curious to see if 3.8 is an improvement in this use case.

AM1010101 14 hours ago||
Seems to do reasonably well in opencode according to artificial analysis. https://artificialanalysis.ai/agents/coding-agents

If I had to pay per token I would probably consider using this (they seem to be on the pareto of performance) but not being able to use opencode with a subscription is not really something I'm realistically going to do when claude and codex are around. Also never gotten along well with gemini-cli / antigravity-cli.

buntp 14 hours ago||
It seems like this is one of the most powerful models for the price, really didn't see that coming from Google
f311a 15 hours ago|
Is the google infra stable enough right now? At the start of the year, the flash model was unusable for a whole month via gemini CLI. They could not fix it for a whole month and I was a paid customer.
ipsod 15 hours ago||
I haven't had any issues lately.
elias_t 14 hours ago||
I use it quite a lot and after a week of use I’m being hard rate limited
More comments...