Top
Best
New

Posted by bratao 21 hours ago

Gemini 3.8 Flash and 3.8 Flash Cyber(blog.google)
https://deepmind.google/models/model-cards/gemini-3-8-flash/
1054 points | 592 commentspage 9
ldm0 17 hours ago|
It’s strange that its score on Terminal‑Bench 4.0 is so low. They aren’t fast enough to benchmaxx that section.
algoth1 12 hours ago||
I asked gemini 3.8 high to review the site I'm working on for points of high cpu/ram consumption - it failed spectacularly and also halucinated the server i/o limits
prometheus1992 20 hours ago||
Google keeps flashing everyone where everyone is expecting to get PRO'bed.
kzrdude 20 hours ago|
We also had GLM-5.3 flash and Qwen 3.8 Flash Next, everyone's getting flashed and I think it's a good trend.

Almost suspect that the rate of improvement to post-training is so fast that small models have an advantage - it takes much more compute to train a bigger model, so the flash models are just running in circles (well, not exactly of course) around the larger models right now.

atemerev 20 hours ago||
Everyone is censoring models now with anything remotely resembling cyber or bio. I already have problems with my research in mathematical epidemiology because of that - both Sol and Fable simply refuse. They keep pushing people towards Chinese models that can be decensored.
vehemenz 20 hours ago|
Supposedly Fable 5.1 is better, but I haven't tried it yet. I've run into the same thing with mundane work that is barely bio/cyber adjacent.

Re: Chinese models, even if the model itself isn't censored, some of the big model providers have guardrails now that you can't exceed, which somewhat defeats the purpose.

atemerev 17 hours ago||
"Uncensored" means "weights modified to remove refusals". Abliterated. Providers do not serve such models, at least not frontier-grade. You have to run the weights yourself. For Kimi K3, this is about $60/hour for hardware rental. But you can have about 100 sessions simultaneously.

And yes, Fable 5.1 has the same refusal rate, and significantly nerfed reasoning.

im_soul 18 hours ago||
disclaimer : Introductory price expires on December 31, 2026. Starting January 1, 2027, $1.50/1M input tokens and $7.50/1M output tokens will apply.
lgl 20 hours ago||
Am I the only only one thinking that Google might still "win" the AI race, despite the apparent gap?

They're apparently evolving slower than most SOTA models but "slow and steady wins the race" is probably still a thing.

And since Google doesn't depend exclusively on AI models, they can probably afford to "wait and see" where all this craze is heading.

sejje 17 hours ago||
Staying a little ways behind the leaders is not "slow and steady." Every company is moving very fast right now.

I don't think slow and steady will win this race, but I think anyone can still win--especially Google.

evilhackerdude 18 hours ago||
i always thought alphabet’s own youtube videos must be a comparatively good source of new training data. if slop and other garbage is reliably filtered out it should leave plenty of higher quality content.
fitsumbelay 21 hours ago||
shows up in /models though and encourages you to use it over 3.7 Flash I prefer this over reading specs: the "just show me" way
realist_not 21 hours ago||
Anyone has a cached page / mirror ? 404
Namahanna 21 hours ago|
Page - https://web.archive.org/web/20260902151410/https://deepmind.... PDF Card - https://web.archive.org/web/20260902150007/https://storage.g...
johnnyApplePRNG 11 hours ago|
Why is it still such a bad coding agent? Does anybody have any insight?

I am continually impressed with Gemini's chat responses, which encourages me to test their agentic capabilities and... no... no... and no... every single time.

It's terrifying watching it, really.

More comments...