Top
Best
New

Posted by bradleyg223 1 hour ago

Gemini 4 Argon(blog.google)
663 points | 410 comments
taylorfinley 1 hour ago|
Ten days ago I had an experience with Gemini 3.8 flash that made me wonder if I was being routed to a different model under test. I was trying to use rocm with llama.cpp on my 128gb Strix Halo but could only get it to run Vulkan. I pasted the error message into agy and it proceeded to attach GDB to my GPU driver, reverse-engineer the kernel queue ioctl interface, and author an LD_PRELOAD C shim to get ROCm llama.cpp working on my Strix Halo. My jaw was hanging open the whole time.

Edit to add the fix: https://gist.github.com/birep/6f2c8d490c7a29820997d57bd654c3...

spankalee 1 hour ago||
3.8 Flash is just quite good, and so is the Antigravity harness.

I use a mix of Fable 5.1, Opus 5.5, and Gemini 3.8 Flash and Gemini holds it's own. Especially in writing, frontend, and sysadmin work. agy for configuring a NixOS system has been truly incredible.

mapontosevenths 1 hour ago||
Even if agy was the best (it's not, and is missing basic features) you wouldn't rather have a choice?

I cancelled Ultra because they forced me into their harness like I should adapt to them, rather than the other way around.

drusepth 1 hour ago|||
What basic features are missing from agy? I've been using it and cli-cc + web-cc for months (among a few other random harnesses to test here and there) and they all seem roughly comparable to me.

I actually just cancelled Ultra also because I couldn't subscribe to a YouTube Family plan while I had it active (Google... :[) but trying to use Codex as a replacement while I testdrive Astra makes me yearn for agy again.

walthamstow 1 hour ago|||
I have used it for little more than 6 hours or so in total but I'm pretty sure it doesn't have compaction?
KeplerBoy 1 hour ago|||
How else would it work? Less technical people don't even watch their context usage.
macNchz 7 minutes ago||
In the olden times, aka like two years ago, AI chats would just stop working or just start slicing off the oldest parts of the context to fit the model's window.

That said, compaction feels like an idea that should work reasonably well, but across all of the major providers and agent tools I've used has never actually produced compelling results, to where if I see I'm getting close to the token limit I prefer to start putting a bow on the project and readying it for a fresh start. Even when I provide a detailed compaction prompt it usually focuses on the wrong stuff.

honr 42 minutes ago|||
It certainly has compaction (since the public launch I assume) and I HATE it. I have some remedies but nothing perfect yet. It never retains ALL the crucial bits. If a conversation runs into two compactions it is often a sign that I have to abandon it and retain whatever I can, to form a seed prompt for an adjacent conversation.
sarjann 1 hour ago||||
Auto mode?
KeplerBoy 1 hour ago||
It absolutely has auto mode.
smartbit 37 minutes ago|||
agy cli does not have auto mode. I've tried and tried and tried to work with agy cli sandbox-mode and just failed.

  agy --dangerously-skip-permissions
in my experience is the only workable solution that doesn't ask confirmation for every step. And I hate working in YOLO mode. Seemingly the Antigravity GUI had some features added in a recent release, but a) I don't want to work with the GUI and b) it was poorly implemented as I couldn't get it to work. VS Code plugins are allowed with subscriptions, but is not the CLI experience of Claude Code I want.

gemini-cli supported 'pre-write diff tabs' (y/n) in external editors like vscode. In Claude Code I heavily use 'pre-write diff tabs' for documentation and miss it sincerely in agy cli.

IMHO Gemini 3.8 flash is fast and good enough, but the agy-suite is below par to say it nice. Someone else in this thread calls agy a terrible harness which is probably more accurate.

SomaticPirate 48 minutes ago||||
I think it only has "--dangerously-skip-permissions" Claude and codex auto mode will reject certain actions. No secondary check on gemini/agy AFAIK
levelZero 58 minutes ago||||
Via cli switch, but in process w/o fine graining? If so please tell
LoganDark 32 minutes ago|||
Auto mode means that another model reviews tool calls to attempt to disallow less safe ones. It's different from bypass permissions mode which typically just doesn't filter at all.
arizen 1 hour ago||||
Does it have /goal feature similar to Codex?
sorrybutidontha 1 hour ago||
yes
esafak 1 hour ago|||
I use a variety of models for various subagents. I don't want to change my harness every time I change models, or be beholden to companies for something the open source community can handle better.
eloisant 53 minutes ago|||
There is a pi plugin to use agy directly from it.
mapontosevenths 10 minutes ago||
You get banned if they catch you.
IndeanCondor 1 hour ago|||
Can confirm, I was doing a routine internet search thing for a curiosity 3 days ago (about the only thing I used Gemini for) and was surprised by how suddenly thorough and quality the response seemed, almost overnight.
seanthemon 53 minutes ago||
Gemini for day-to-day and top-of-head queries and claude for the real beefy work
amanguliani 1 hour ago|||
Can confirm - I am HEAVY claude user, but always like to check with AGY and CODEX in between. AGY with Gemini 3.8 flash cooked last couple of times and CODEX is basically out of the mix for me
onlyrealcuzzo 25 minutes ago||
Sol 6.1 is quite good, but damn is it slow.

I'm using it to run overnight tasks, and that's it until my quota runs out.

Canceled my subscription.

yegle 1 hour ago|||
For getting redroid running on my Linux system, 3.8 Flash decided to binary patch a .so file instead of getting the AOSP source code and patch/build it properly.

And I saw it do this twice, once for Android 14 and once for Android 16.

I think this is just within 3.8 flash's capabilities.

p_l 49 minutes ago||
3.8 Flash (but also last two ones) have really strong preference for dissecting binaries with quick thrown-together bits of python in my experience.

Including going first for decompiling AGY binary instead of searching the web for documentation...

IshKebab 42 minutes ago||
Astra also really loves reverse engineering binaries. I guess it's one of those things that isn't that complicated but is super tedious, and tedium means nothing to AI.
gottorf 1 hour ago|||
My experience with Gemini 3.8 Flash has been awful; it gives me the most hallucinations out of the major models. I'm not using it for coding, but general research on different topics.
MILP 47 minutes ago|||
I'm also not using it for coding but I've found Flash 3.8 to generate much better HTML output than Sonnet or Opus.
staticman2 59 minutes ago||||
The web version of Gemini is awful at search but I don't think that's the models fault.
mattjoyce 41 minutes ago|||
Hallucination seems a very dated term.
nkozyra 27 minutes ago||
Why? It's the same concept and root cause it was when we first started using it.
mapontosevenths 1 hour ago|||
Gemini is honestly amazing sometimes. If they didn't force you to use a terrible harness, charge too much for way too little, and generally act like customers are a giant problem to be avoided I'm sure Google could take over the AI market.
ody4242 46 minutes ago||
what is so terrible with their harness? I've been using gemini cli, now use agy, Pi agent harness, and agent (cursor), and my only real issue with agy was the permission handling, but other than that, it was ok.
alightsoul 1 hour ago|||
Please tell me you published your findings even as an issue on the llama.cpp GitHub
taylorfinley 15 minutes ago|||
Here they are: https://gist.github.com/birep/6f2c8d490c7a29820997d57bd654c3...
dominotw 38 minutes ago||||
he is still closing his jaw
otabdeveloper4 1 hour ago||||
Spoiler alert: the problem didn't actually get fixed despite the jaw on the floor.
warkdarrior 1 hour ago|||
Why? Anyone can run that prompt.
aspect0545 1 hour ago|||
Not everybody has access to AI. More than that, every prompt uses insane amounts of natural resources. So why not share it.
FranzFerdiNaN 1 hour ago||
The resources per prompt aren’t that much .

Also I hope you don’t have children, eat meat, travel, have a car, run AC, buy things in other countries and such. Those things all take way way way more natural resources.

articulatepang 53 minutes ago|||
All your examples are private goods: excludable and rival. If one person uses a unit, that prevents others from using them.

Patches to open source software are public goods. Your using them doesn’t prevent others from using them. So if you spend resources creating a public good, it’s in everyone’s interest to share it.

qmr 56 minutes ago||||
Yet you participate in a society.
scarmig 43 minutes ago||
If action X takes a million times more resources than action Y, it's silly to focus on or highlight action Y. Seriously: if you are a regular meat eater, your choices use several orders of magnitude more water than even a heavy LLM user. A quip from a comic doesn't somehow erase that or make it irrelevant.
hexfish 1 hour ago|||
Checkmate. /s
luckydata 1 hour ago|||
why reinvent the wheel and spend tokens for a problem that has already been solved?
bel8 1 hour ago||
I had a similar but less impressive experience recently with Muse Spark 1.3.

Asked pi agent it to identify the main hero sprite size of game I was running. It had a ton of shader effects so it was hard to determine.

It used some cli tools to identify that it was a game made with Godot, decompiled the executable but data was encrypted, broke the encryption after writing a brute force tool to test keys extracted from the exe, then proceeded to extract the game gd scripts and assets, only to answer the question of the sprite size.

seanthemon 51 minutes ago||
Godot encryption is laughably easy to break, there's tons of packages available for it. It's a well known drawback of using godot
nickysielicki 1 hour ago||
The important take away here: the leapfrogging we’ve seen this year doesn’t seem to be a temporary thing. The famous theory of Dario Amodei was that AI was this winner-takes-all field where the first team to get a head start would never cede ground back. The term he liked to use was, “concentrating”. This is yet another datapoint that he was wrong about that. AI seems more distributed amongst neoclouds and traditional hyperscalers, FAANG and startups, GPUs and ASICs than it did this time a year ago.

Nobody has a moat.

LarsDu88 1 hour ago||
Google has TPUs, a frontier model, a completely separate and lucrative revenue stream they can call on at will, and teams working on multiple different language modeling strategies simultaneously. Did I mention the vast and ominous data centers that already serve a significant fraction of the internet? If that ain't a moat, then what exactly is a moat?
jobs_throwaway 1 hour ago|||
Then why have they been lagging behind OpenAI and Anthropic for most of the last few years, and only briefly been at the frontier?
aleph_minus_one 54 minutes ago|||
> Then why have they been lagging behind OpenAI and Anthropic for most of the last few years, and only briefly been at the frontier?

One possible explanation: because Google is a little bit more frugal and focuses on how to make providing AI models financially feasible - combined with some willingness to burn money so that they don't strongly fall behind on their AI models.

On the other hand, OpenAI and Anthropic at least formerly concentrated on building and providing the best models that they could with concerns about financial feasibility taking a backseat.

Just to be clear: I do have the impression that by now (likely because of pressure from investors) OpenAI and Anthropic take these financial concerns more seriously, but nevertheless Google's vs OpenAI's/Anthropic's "DNAs" concerning on what to focus on differ.

monkeydust 36 minutes ago|||
And also Google is public listed company and the other two (for now) are private. I think that fact does have a strong bearing on how they operate.
redanddead 33 minutes ago||||
> Google is a little bit more frugal and focuses on how to make providing AI models financially feasible

There’s no magic there. You get an account executive and a call with a systems architect to find out what you’re doing.

Clouds gonna cloud, this is the reason they rolled deepmind into gcp and arguably the inverse is true, the labs are trying to become clouds

topspin 30 minutes ago|||
> because Google is a little bit more frugal and focuses on how to make providing AI models financially feasible

That feels right. It's not as if they've been missing out on great profits.

spyckie2 44 seconds ago||||
Could it just be that Demis was checked out of the race and they lacked leadership?
tfsh 35 minutes ago||||
> Then why have they been lagging behind OpenAI and Anthropic for most of the last few years, and only briefly been at the frontier?

Because it's not an existential battle for Google. If OAI or Anthropic disappear from the absolute frontier for ~8 months the news cycle and churn will diminish them to the second rate. Google is processing near 4 quadrillion tokens every month, that's - I'm sure - significantly more than OAI or Anthropic, because Google is interested more so in their flash models and getting these competitive, which they are.

redanddead 32 minutes ago|||
It’s just a division in their cloud offering that’s what AI is
senordevnyc 5 minutes ago|||
[dead]
merb 52 minutes ago||||
From a business perspective a frontier model does not make much sense anymore if you are not a startup. Neither for Amazon, nor for Google. Their clouds need models that are fast and perform well in their agent frameworks nothing were a frontier model excels at.

Most Google products even use flash lite underneath, so their frontier model is mostly used for distillation.

aleph_minus_one 40 minutes ago||
> From a business perspective a frontier model does not make much sense anymore if you are not a startup. Neither for Amazon, nor for Google. Their clouds need models that are fast and perform well in their agent frameworks nothing w[h]ere a frontier model excels at.

A good consideration; just one point from my side: as far as I am aware (but I may be wrong), Gemini is not known to perform well in an agentic framework.

This is no contradiction to your other claims, quite the opposite: perhaps (or even likely) Google wants to avoid that their models become a commodity in some (agentic?) application where the middleman who actually writes this application gets a disproportionate of the money that the customer of the application pays for it.

nozzlegear 31 minutes ago||
> A good consideration; just one point from my side: as far as I am aware (but I may be wrong), Gemini is not known to perform well in an agentic framework.

I used it for a month over the summer, right before they were going through the migration to antigravity. It was a fine workhorse IMO, no complaints from me.

krona 30 minutes ago|||
Alphabet issued a very oversubscribed 100-year bond with 6.1% yield earlier this year to raise capital for datacenter expension.

Meanwhile, Anthropic/OpenAI will struggle to survive the next 24 months on their current trajectory.

jeremyjh 26 minutes ago||
So you don't know why they've been behind?
IshKebab 41 minutes ago||||
Don't forget the training data! Legal copies of all the books in the world, the entire web scraped, and all of YouTube.
bluecalm 45 minutes ago|||
I don't think that other revenue stream is completely separate. It weighs on them as they need to think about tradeoffs. Classical search is going away sooner or later so they need to replace that with AI powered search.

Data centers are important but a few others also has them: Amazon, Microsoft, Meta. SpaceX will likely be in/at the top I AI dedicated precessing power in 2027 as well.

I don't see the moat. I see a company with a lot of other commitments that is not the best at delivering consumer facing products. They have some good cards but so do others.

xnx 1 hour ago|||
> Nobody has a moat.

Custom hardware, data centers, huge cash reserves, deep/broad talent pool, and non-AI customer base are all huge advantages if not moats.

Google, Microsoft, or Amazon are more likely to be the AI leaders than OpenAI or Anthropic.

woah 31 minutes ago|||
> Google, Microsoft, or Amazon are more likely to be the AI leaders than OpenAI or Anthropic.

If not now, then when will these companies be AI leaders?

Even Google, with its staggering advantages in cash, compute, real estate, training data, and having basically invented the field only manages to briefly claim a 1-2 week lead once or twice a year.

koe123 27 minutes ago||
The financials for Anthropic and OpenAI are likely borderline suicidal, google and co are publicly traded. Moreover, all innovations downstream to them dont they? Why not just stay slightly behind, especially given many have stake in those other companies?
bluGill 47 minutes ago||||
There are many companies that have data centers. They are conceptually easy to build. An ASIC is difficult enough that if you make one someone will leapfrog you while you are still making it (at least so far), though once you have one your costs will be enough lower than the competition that you can perhaps undercut them.
xnx 37 minutes ago||
True, but have other hyperscalers caught up to Google's AI data centers?: fully liquid cooled, torus networking(?), 100,000+ TPUs interconnected, etc.

Google is already on gen 8 of its TPUs and is certainly already working on the next version or two.

IX-103 1 hour ago||||
If Moore's law continues, then in less than 10 years today's state of the art model will be able to run on a cell phone. How much smarter do we actually need AI to be? Would it still require datacenters and custom hardware?
xnx 34 minutes ago|||
Moore's law stalled ~2015. Unfortunately, no way current models will run on the <100W thermal budget of a cell phone. Printing the weights directly into a chip would help efficiency a lot, but not enough.
CuriouslyC 13 minutes ago||
In all likelihood in a few years we'll get ~200-400bA~4-6 MoE models that are on chip, and they'll be better than the current frontier.
LightBug1 41 minutes ago|||
They probably said the same thing about social media back in the day.

I'm sure the thinking out there, and hence investment, is all about how to tether the user to the most addictive, network-effected, incredibly deep, server-side, moat-able version of AI possible.

junehwi 1 hour ago|||
[dead]
altruios 1 hour ago|||
> Nobody has a moat except nvidia

For now, for cloud training. but for consumers, nvidia vs amd reasonably close - the moat there is thin and shrinking. I suspect AMD will surprise us. nvidia has no motes in china, which may be a new source of (gpu) chip design. Huawei's Ascend 910C is about a generation behind... again: for now.

point is: moats dry up. I see nvidia's shrinking as a real possibility.

culi 1 hour ago||
China will always be generations behind until they crack domestic EUV
aleph_minus_one 47 minutes ago|||
> China will always be generations behind until they crack domestic EUV

Are you sure?

--

China Just Built What TSMC Said Was Impossible

https://www.youtube.com/watch?v=Pk-w279ESHg

--

China Just Built What ASML Feared Most

https://www.youtube.com/watch?v=YiPgSm62fiM

altruios 59 minutes ago|||
Do you think that they won't?
FinnKuhn 40 minutes ago|||
My personal theory is (assuming there really is no moat) whoever starts the latest with developing AI models might actually win as they should be able to develop a competitive product with significant less resources and initial investment resulting in a higher ROI. AI might even become a commodity.
pkfz 8 minutes ago|||
Given that the infrastructure won't be a moat and will become a commodity.
koe123 25 minutes ago||||
This is my secret hope for Europe!
dgellow 18 minutes ago|||
That’s what we see in China
pvab3 44 minutes ago|||
The whole winner-take-all idea seems entirely based around Singularity/Rationalism and would require massive advances that we probably aren't close to at all.
nater5000 10 minutes ago||
Yeah, it kind of seems like we haven't gotten to the "head start" he's referring to yet.
kushalpandya 39 minutes ago|||
Even Google itself stated (internally at least) that nobody has a moat https://newsletter.semianalysis.com/p/google-we-have-no-moat...
bitpush 34 minutes ago||
Wasnt it just some dude writing a doc? That's hardly a Google (The Company)'s position.
dgellow 17 minutes ago||
If anyone other than NVIDIA has a moat, they for sure never talk about it
vb-8448 1 hour ago|||
It's even worse, we are crossing over into the realm of religion. The article against GML 5.3 is the equivalent of a Papal excommunication.
zone411 57 minutes ago|||
The article presented facts and data. If that's a problem for you, that sounds more faith-based than whatever Anthropic is doing.
verdverm 1 hour ago|||
which article? have not seen this one

---

maybe it's this Anthropic post on GLM?

https://www.anthropic.com/research/glm-5-3-and-the-spread-of...

> Governments should conduct safety testing on sufficiently capable AI models, including successors to GLM-5.3. Without high-quality evaluations from independent sources, the impact of these capabilities might not become fully clear to model developers until it is too late. As AI developers across the world build increasingly capable open-weight models, we hope they work to appropriately safeguard these capabilities and prevent misuse.

I for one do not think my government is up to the task of designing or implementing such a system

Rzor 1 hour ago|||
https://www.anthropic.com/research/glm-5-3-and-the-spread-of...
Iolaum 1 hour ago||||
It doesn't need to. It can use your cash to pay the people who are. Those in power like it more that way.
esafak 1 hour ago||||
You say that because the 'most' existing models have done is hack governments and companies. Can't you think of worse things a model could do; accidentally or by instruction?
verdverm 1 hour ago||
help people with suicide and school shootings like ChatGPT already has

OpenAi is alledged to have been monitoring these internally and not contacting authorities. Lawsuits have been filed, I see gross negligence without the gory details

I have for more concerns around human-chatbot maladies than I do around the cyber security stuff. For example, why hack grandma when you can get her to do something willingly through impersonation. How do we prove authenticity in a post truth world?

aleph_minus_one 1 hour ago|||
> The famous theory of Dario Amodei was that AI was this winner-takes-all field where the first team to get a head start would never cede ground back.

This is the kind of story that ones tells to investors to justify the huge amount of cash burn. :-)

mapontosevenths 1 hour ago|||
I'm not sure it's wrong. This all feels a bit dotcommy to me.

I think many/most of the players will crash and burn, and the ones that are left will divide the world.

CoolestBeans 42 minutes ago|||
The problem is twofold. One, even a monopoly AI provider wouldn't have pricing power against its suppliers. Its suppliers are energy, semiconductors, and real estate. Semiconductors maybe they could get some leverage on but energy and real estate have plenty of other buyers. Two, there's still no evidence of a runaway scenario (ie a small lead turns into a big lead over time) and there's still no evidence that there's some resource that you can deny everyone else that they can't build your product also. You can't hoard energy, compute, memory, data, human talent, or customers.

The net effect is that the most likely scenario is if one big lab fails, they will likely all fail. Their revenues are all correlated.

To go to your dotcom comparison, the winner will be the ones picking through the assets that were written down by orders of magnitude and trying new products with the technology until one sticks to the wall. But I don't know if a dramatic crash is guaranteed either.

aleph_minus_one 37 minutes ago|||
> The problem is twofold. One, even a monopoly AI provider wouldn't have pricing power against its suppliers. Its suppliers are energy, semiconductors, and real estate. Semiconductors maybe they could get some leverage on but energy and real estate have plenty of other buyers.

Concerning the leverage on energy and real estate: don't forget that the AI companies have quite a lot of choice where to build their data centers. So AI companies have lots of opportunities to play several parties off against each other (in particular also for real estate and energy).

3d2 29 minutes ago|||
"but it can stay there so long as the balance sheet doesn't deteriorate."

Uhm, what? LOL.

People dont value firms based on balance sheets fella. Have you taken a basic valuation class?

Tesla is a nice stock for traders - they like the volatility. Nobody holds Tesla as stock for investing. If you were to truly value it on an intrinsic value basis you'd have to bring in failure risk.

jaggederest 1 hour ago||||
I suspect this is going to end up like most services provided e.g. cloud stuff, balkanized between a couple major players and an assortment of DIY or less popular options if you don't like those ecosystems, plus some UX/DX focused wrappers that use the big players under the hood.

I think that would be a pretty satisfactory outcome compared to one hypercompany consuming trillions of dollars of the world economy.

skybrian 55 minutes ago||||
“Divide the world” sounds ominous. Here’s another scenario to consider:

Internet access is not really unlimited, but for many people with fiber at home, it effectively is and we pay a flat rate.

Perhaps by the end of next year, most programmers will stop thinking about metered access for AI? For many people, the cheaper models (about as good as today’s frontier models) will be good enough.

Which might sound good, but the downside is that it will also be easier to build an AI botnet without the users paying for it noticing. Particularly when people are running AI inference on their own hardware.

pixl97 9 minutes ago||
Hardware is still insanely hard to get a hold of, and the stuff that's being built doesn't really work for home use. Maybe if it crashes Nvidia will adjust the hardware flow.

My guess is even if the AI market busts there is still a massive demand for hardware as models are solving all kind of problems now.

But ya, lots of hardware everywhere not managed well is how you get sovereign AI.

pianopatrick 55 minutes ago||||
Or, like airlines, the ones that are left will have great technology but be not so great from a business and financial perspective. To me AI seems like a commodity service.
pvab3 45 minutes ago||
Like airlines but starting off with hundreds of billions of dollars of obligations and debt
3d2 30 minutes ago||||
THe problem with analogies is that they are imperfect.

I would argue those who already rule the world, will continue to do so.

What happens to OAI and Anthropic? No idea, probs go bust. Google just has to offer a half-decent offering in the long run and have a cost-advantage and it'll eventually knock OAI and Anthropic out as firms figure out what combination of models they want to be best for their economics and generating returns. Enterprises trust google over OAI and Anthropic. A clear signal of this was the Apple deal.

Dont forget those sweet returns fellas! CEO's are hired to make the owners wealthier. That is not gone.

ehsankia 1 hour ago|||
I guess if one of them hits singularity, it could in theory just wipe out all the rest, seeing how they keep escaping and hacking into other systems :)
aleph_minus_one 1 hour ago||
> I guess if one of them hits singularity, it could in theory just wipe out all the rest

The story that some AI company might reach singularity and then "everything will be different" is another science-fiction story that executives of AI companies love to tell to justify the staggering amount of necessary investments and cash burn. :-)

koe123 29 minutes ago||
I find it quite unique how many people buy into this. Its the worlds most blatant conflict of interest, I dont even know why Sam and Dario bother doing interviews
hirako2000 1 hour ago|||
And before him, Altman was explaining very calmly that no company could ever compete with OpenAI.
Heidaradar 1 hour ago|||
I just find this unlikely personally, think about the great research that's happening in the open source world, I'm sure inside anthropic + openai they've also made a bunch of discoveries and improvements (and I'd guess way more due to them attracting the best talent + the better internal models they have)
qgin 1 hour ago|||
Whoever gets to RSI first “wins” but also maybe ends life on earth. The incentives have never been worse.
pianopatrick 52 minutes ago|||
Maybe. Or maybe having the best AI model on the planet becomes like having the best super computer on the planet. Useful for some niche stuff, but not too useful in terms of people's daily lives or what is used in business.
koe123 22 minutes ago|||
RSI being science fiction so far.

Whoever builds the deathstar wins!

culi 1 hour ago|||
I'm not necessarily defending this obvious marketing speak but maybe the "starting point" was wider than assumed. So far, nobody has caught up to US and Chinese labs for example despite lots of funding in Europe. This is also despite abundant in-depth research papers being published alongside open source code and weights by some Chinese labs
funnym0nk3y 1 hour ago||
There is not really much funding in Europe. At least not for start-ups. There is simply not enough compute in Europe.
RachelF 1 hour ago|||
The US companies still have trillion dollar valuations like there is a monopoly. There just isn't one. They are all within a few percent of each other on the benchmarks.

The slightly lower Chinese open models are good enough for almost everything, too, and much cheaper. Like with humans there is plenty of employment for people with below genius level IQ's.

jppittma 58 minutes ago|||
I feel like the frontier labs are going to serve fast/lower intelligence models at a better per token cost than the open chinese models. You're telling me that in the long run, you're going to self-host your own ai infra for cheaper than google can serve it to you? I don't really buy it. I think the dedicated AI data centers are going to serve AI at a lower marginal cost than random businesses self-hosting, and then it's a question of how much of that margin they can capture.
msy 54 minutes ago|||
Agree entirely but that's the point, if it's a margin knife-fight with marginal product differentiation/pricing power nobody is going to be making bank.
pvab3 37 minutes ago|||
If they have to recoup training costs then they don't have much choice
fumar 1 hour ago||||
Is there a dividing line between good enough and best in class capabilities? It's blurry from where I stand. Will model makers cede ground or is there a market making moment up for grabs (singularity)?
handfuloflight 1 hour ago|||
> Like with humans there is plenty of employment for people with below genius level IQ's.

Not if the genius level IQs take the market share.

nylonstrung 1 hour ago|||
I think that scenario only naively made sense if technical knowledge was entirely proprietary and talent was guarded with severe non-competes and NDAs

And Chinese labs openly publishing so much of their methodology destroyed any hope, which was inevitable

torginus 1 hour ago||
> And Chinese labs openly publishing so much of their methodology destroyed any hope, which was inevitable

I think the secrecy doesn't make sense. People swap jobs between labs so I'd say the big players can' really keep secrets for long, and any secret sauce advantage gets incorporated by competitors in a major product cycle at most.

ddp26 27 minutes ago|||
Isn't the important takeaway here that Gemini 4 is not released and has no planned release date?

This is marketing from Google, not a competitive offering

SwellJoe 1 hour ago|||
I think some in the AI industry drank their own Kool-Aid. They believed that if they had the best model and the most compute, they could tell the model, "Make a better model." And it would, and the next one could make its replacement, and so on.

So far, that's not exactly how it's played out. Humans are still necessary for the leaps in capability or efficiency. A model can grind on a problem to eke out the most performance, and models can synthesize data and iterate on various techniques to find the optimal combination. But, seems like humans still have to provide the real thinking, and the talent and drive for doing that is not concentrated in one company or city or even one country. And, (surprisingly) a lot of the people involved are in it for advancing the field more than making another billion dollars, so they're publishing their research.

So, yeah, the moat isn't deep. Even the compute moat, that OpenAI, Musk, and a bunch of other also-rans (like Oracle) bet the farm on, isn't really panning out. The Chinese makers just spent their effort on making models vastly more efficient, since they couldn't do anything about having an order of magnitude less compute available.

scottyah 1 hour ago|||
But that's the whole point of the singularity. Right now the models use a lot of human effort and ingenuity to improve the models, but about a year ago it was 100% human. We'll see in another year, but if this pace continues I doubt there will be more than a handful of people who can contribute more than the models.
TeMPOraL 1 hour ago|||
> I think some in the AI industry drank their own Kool-Aid. They believed that if they had the best model and the most compute, they could tell the model, "Make a better model." And it would, and the next one could make its replacement, and so on.

They're not there yet. Once they get there, that's literally the definition of Singularity.

But they are getting closer. Recursive Self-Improvement used to be a phrase people mocked LessWrong crowd for using and worrying about, now it's something both OpenAI and Anthropic already publicly admitted not only to pursue, but to already be benefiting from.

SwellJoe 35 minutes ago|||
Sure, it's happening...but, is IT happening? By that, I mean, we can see that the models are able to iterate at a pace and scale that humans can't match, and that provides gains in model performance and efficiency. But, humans are still needed in the loop, and not just because it's necessary for safety/alignment reasons. I don't think any significant discovery has been made by models on their own, and I don't know that LLMs will ever have the capacity to invent. They can synthesize from known data amazingly well, and since they know everything "known data" is extremely broad. But, the leaps, so far, have all come from humans.

So far, I don't think the models are capable of running away on their own. Of course, it would be playing with fire to not at least consider the risks of such a runaway scenario and build in safeguards against it. But, there is no model that can build a better model on its own, thus far, to the best of my knowledge (which is far more limited than the models, so maybe I should ask them).

TeMPOraL 3 minutes ago||
Recursive Self-Improvement isn't instant, it starts slow and accelerates.

It starts with what they already claim to be doing - increasingly relying on existing models in non-trivial work related to training, evaluating and optimizing the next, more capable generation of models. As long as the proportion of work keeps shifting towards agents doing more and more of it, and humans less and less, that's RSI at play.

It may be that it turns out LLMs lack some fundamental level of judgement and it plateaus, but frankly I find this notion absurd; LLMs already show better judgement than most people. The alternative is, at some point LLMs will show the ability to futz their way into improvement of the next generation of models even without humans in the loop - even if much less efficient at first, if generation N+1 is more capable than generation N, it'll either take off or burn out.

thmoonbus 35 minutes ago||||
The companies whose insane valuation is based on accomplishing thing X say they’re getting closer to accomplishing thing X?

At least they’re led by trustworthy and honest people or we’d need to take their claims with some dose of skepticism.

pvab3 35 minutes ago||||
Those people have a lot of overlap with the LessWrong crowd. They do not have RSI now and probably never will
TeMPOraL 8 minutes ago||
They absolutely do, unless you believe they are lying about the fact they're using current generation models extensively to develop the next generation of their models.
woah 30 minutes ago|||
Is it recursive self improvement if Claude Code writes your pytorch for you?
TeMPOraL 10 minutes ago||
If the point of that pytorch is to improve the next generation of Claude, then yes, absolutely.
arizen 1 hour ago|||
Seems like learning rate velocty may be the ultimate moat
zem 1 hour ago|||
I have never understood the whole "this is a winner take all game" mentality - the sheer size of the pie is so great that from a purely rational standpoint companies should just be trying to productively get a slice of it and be profitable. winner-take-all is just greed/capitalism run amok, where it is not enough to be profitable, you have to own the entire market (and presumably extract rents)
dgellow 11 minutes ago||
It’s like the supposed first mover advantage OpenAI believed they had. In practice it’s almost always more like a first mover massive tax, and companies coming afterwards benefit from your discovery of a market, publicity, and everything else that has already been validated
sixo 46 minutes ago|||
Nobody has a moat so long as employees can move between companies
Aboutplants 1 hour ago|||
I feel his theory depends on the premise that access to pure compute would the be the determining factor of success. Not the case
tripleee 1 hour ago|||
Was Dario's company winning at that point in time by any chance?
esafak 1 hour ago|||
The present leapfrogging is not a contraindication because companies are not necessarily releasing their best models; we know they have smarter internal models. Furthermore, humans are still involved in model creation. Human involvement is expected to decrease over time, and when model iteration is completely automated, progress will happen at the machine's pace, leading to runaway intelligence, barring any ceilings.
SecretDreams 1 hour ago||
AI is a commodity. One that is showing to be more readily commoditized than most has anticipated. As of now, the only moats are the financing for the hardware to run it and the hardware vendors themselves - with the latter largely not yet a commodity because of ecosystem lock and a limited capacity of the most advanced fabs in the world.
babelfish 1 hour ago||
> We’ll continue to gather feedback from early testers as we iterate on guardrails before making Argon available to developers, enterprises, and consumers as soon as possible.

Gemini not beating the "can't release a model" allegations

Androider 1 hour ago||
My Gemini app (updated today) and https://gemini.google.com/ has _3.6_ as the latest selectable model, as a paying Pro user in the US. How is that even possible? Gemini 3.7 was released in August, 3.8 early September. What is going on over there?
wasabi991011 15 minutes ago|||
I would imagine you are experiencing a bug. I've been using 3.8 daily since its release (on a Pro plan in Canada). I believe this is true of many people.

What is your reason to believe this is not a bug specific to a small set of Pro users?

vlyan 1 hour ago||||
just Google being whatever the fuck it's been for the past 15 years.
asdfasgasdgasdg 18 minutes ago||||
That's weird. I got 3.8 and 3.7 on the days they were released. https://imgur.com/a/Xu4wRLM (this is on gemini.google.com, but the same is true of the iOS app and the desktop app).
ttul 17 minutes ago||||
Indeed. They have this amazing model and you can’t access it in their own branded app. It’s insane.
XzAeRosho 1 hour ago||||
I was reading the announcement and wondering the same. And don't forget, still with 3.1 Pro as the frontier model.
AuthAuth 1 hour ago|||
they moved it from the place you'd expect to ai.studio
modeless 1 hour ago|||
When I said I was tired of Google launching waitlists I didn't think they would respond by simply not having a waitlist.
ionwake 1 hour ago||
i know this is like "hey guys we got such a cool thing at home ,its rad and uhm we playing with it with our friends"

ok bro thx

Culonavirus 41 minutes ago|||
Yeah what's up with that. Also what's with the next big update for Nano Banana? Nano Banana Pro was released almost a year ago!
mrieck 54 minutes ago|||
I'm glad.

I already pay $300+ for subs. Please don't tempt me with another $100 sub just because I got curious if the benchmarks were right.

cmrdporcupine 57 minutes ago|||
They will go through the usual transition of "can't release a model" to "won't load in a harness normal people can use for 3-4 weeks" to "it's smart as hell but completely inept at tool use and coding" like every Gemini release.
bakugo 1 hour ago||
They're just following the current AI marketing playbook. "Our new model is simply too dangerous to release to the public right away" is now standard practice.

They even gave their model a random nonsensical name suffix simply because OpenAI is now doing it, too. Monkey see, monkey do.

A_D_E_P_T 1 hour ago|||
I'm still at a loss as to what argon has to do with anything. Say what you will about Luna-Terra-Sol-Astra, or Haiku-Sonnet-Opus, they make sense. I don't see how Google can make sense of argon; it's in a fairly strange place in the periodic table...
brainwad 1 hour ago|||
They are going alphabetically, Android style.
IX-103 52 minutes ago||
Yeah, I heard the next one was Barium...or was it Boron?

I was going to say I don't know what they'd do for C, since Carbon and Calcium are already things. But knowing Google, they'll probably call it Chromium.

fooker 1 hour ago|||
Google R Gon lose the AI race
jstummbillig 1 hour ago|||
> now standard practice.

Opus 5.5 and Sol 6.1, literally state of the art (in their respective class), were just released without any prior announcement. This has pure and simple become a Google thing.

kevinh 1 hour ago||
OpenAI said they dropped Astra 6.1 over safety concerns: https://www.wsj.com/tech/ai/openai-chatgpt-model-release-can...
jstummbillig 43 minutes ago||
I don't read that as the same category: There was no announcement, no benchmarks, no limited release and no promises about what will happen with that model. It failed internal safety standards. Might be scrapped entirely due to a failed training run, for all we know.
Revanche1367 47 minutes ago||
Great, so they _finally_ decided to add a non-flash model and it's not available to regular subscribers for an indefinite period. What's the point of paying for the AI Ultra plan? Anthropic doing the same with Fable as far as I know, OpenAI at least allows Pro plan subscribers to use Astra. I subscribe to Gemini AI Ultra and ChatGPT Pro, and have enterprise access to Claude at work. To be fair, Gemini's flash models since at least 3.6 have been quite useful for non-complex work, but for any task where there is a bit of complexity involved, I've had to check and recheck the work multiple times myself or sometimes with another LLM to get it to follow plans accurately. It's disappointing to see yet another Gemini release ignore adding newer pro models.

Edit: seems I was wrong about Anthropic restricting Fable, I guess our enterprise plan doesn't include it. But, the block from Anthropic regarding Mythos for regular subscribers/enterprise-users is still true I think.

murkt 40 minutes ago||
Fable is available for a couple of months and even got an update on 1st of September. It’s really good, but since Opus 5.5 was released, there is not much point in using Fable anymore
aqme28 32 minutes ago||
As we’ve seen from the leaked Anthropic prospectus, revenue from actual users is a pittance. What really matters is what you can get from investors, and that you have a model smart enough for self-improvement.
onlyrealcuzzo 21 minutes ago||
Google's already public, and already makes $400B a year in profit...
tazjin 1 hour ago||
> Argon agents are working on migrating C/C++ codebases to Rust across Google

Man, I remember back in the days when the cppnext team was refusing to even consider Rust, instead looking at absurd stuff like Carbon and Swift (!), even though half of the engineering staff already knew where this was headed. I hope they got a few good promos out of the delays at least.

minimaxir 1 hour ago||
A RewriteInRustBench would be unironically useful at this point since all the main agents can write it reasonably well despite its relative scarcity in the input data.
LarsDu88 1 hour ago|||
Rust is the best language for LLMs b/c it gives by far the best debug messages. Just tons of verifiable reward signal for post-training. Even the most rudimentary LLMs can school me on idiomatic Rust
rafram 1 hour ago|||
On the other hand, Rust's borrow checker is very picky, and even a frontier LLM still sometimes struggles to respond to roadblocks sensibly (refactoring so whatever it's trying to do can be done safely) rather than stupidly (introducing some horrible global arena thing so it can make the borrow checker go away). A lot depends on how good your instructions are, and how good the existing code is, since bad input begets bad output.
nchie 15 minutes ago||
I've (more or less; I've read quite a bit of the code) vibecoded several houndred thousand lines of Rust and I've not seen this happen a single time. It sounds like something it'd do when you ask it to "write a linked list while satisfying the borrow checker". Are you sure you haven't (possibly unknowingly) been giving it instructions which ended up luring it into doing these things?
bitexploder 1 hour ago||||
Evidence needed. I think for certain kinds of outcomes it has very strong advantages, but these advantages are not a given as 'best for LLMs' :)
lossolo 10 minutes ago|||
Not always. In my experience, if you're not working on a small, trivial codebase, LLMs will sometimes just create spaghetti unreadable, inefficient code to satisfy the constraints of the type system/borrow checker.
6thbit 49 minutes ago||||
Have each agent rewrite openssl in $lang and call it the RollYourOwnCrypto bench.
culi 1 hour ago||||
I wouldn't be surprised if we're already at the point of more LLM-written Rust than hand-written. Models training off models
adamrezich 1 hour ago|||
All the main agents can write Jai code reasonably well despite being even more scarce in input data!
baq 1 hour ago|||
Not many people can hold grudges as strong as principal engineers
timmg 1 hour ago|||
I wonder if this means Carbon is DOA.

I was excited to see what it would be. But I don't think I can argue that it makes as much sense anymore.

qalmakka 1 hour ago|||
Carbon was clearly DOA the moment it was announced, IMHO. It looked cool but it served none but Google, and now with LLMs you have a massive incentive not to use a niche or new language due to how better LLMs get the bigger the corpus is

The only somewhat realistic proposal in this space is Herb Sutter's cpp2, which is arguably a massive improvement and I'm puzzled why nobody in the standard thought to give it a spin, there's just to much cruft they'll never be able to get rid of unless they make an alternate yet backward compatible syntax with C++ that changes the defaults from "random 80s nonsense" to something better

YuechenLi 1 hour ago|||
Version 0.0.0.0 after 4 years. Their goal of "full interop with C++ while being a completely new language without any of the flaws of C++" is plain absurd.

It's DOA because Google doesn't have any idea of what Carbon should be, and to be completely honest, at least 80% of what they currently use C++ for should be rewritten Go, you know, that language developed specifically because of the issues with C++ by teams within Google.

6thbit 53 minutes ago|||
I wonder if there's people already whose full time job is maintaining/extending one of these auto-migrated codebases.

Imagine they aren't even familiar with rust but are deeply familiar with the product.

pshc 1 hour ago|||
Rewrite everything in Rust has been a meme for so long that to see it coming to pass is surreal.
vovavili 1 hour ago|||
What exactly makes Carbon absurd?
boshalfoshal 1 hour ago|||
There is 0 practicality in inventing an entirely new coding language that only one company uses, and you have to teach it to thousands of new engineers. Rust exists and fits the job totally fine and is used in more places and has actual support outside of a single entity (i.e you can actually hire people that feasibly know the language).

It was clearly done because some PL guys at google really wanted to make a new cool language and Google was the perfect place to incubate it without it getting axed. Probably got a couple of promos out of it too. This is clearly not the best use of time or money, but I guess if you're google you have so much of both it probably doesn't really make a dent, and you can keep a few very smart people happy with shiny new projects.

Also, LLMs being used for a large portion of coding nowadays sort of remove the need for these types of languages, IMO. They make less "silly" bugs (both logical and structural) that languages like this are meant to catch, and they are much better at languages that are better represented in the training corpus. This somewhat obviates the need for very niche "type/dummy-safe" languages like carbon (and even rust/zig, imo). So even if you did want to use Carbon, you'd likely have to bootstrap a decent amount of your own "good" carbon code to post train an LLM, and even then, it likely won't have that big of a gain vs just having an LLM write C++ or even Rust. If you are a company that still reviews code, you should just have an LLM code in a language most people can understand anyway to make verifiability tractable.

torginus 3 minutes ago|||
Rust is not the end of history. One of the difficulties with the language lies exactly with porting existing code written in an OOP style to idiomatic Rust, as those codebases weren't written with ownership in mind.

Such rewrites will contain judicious uses of Cell, RefCell, unwrap() etc. which make for ugly code that's not exactly simple to understand and might even have some landmines (crashes).

Getting rid of these requires a subtantial amount of engineering effort, which I'm not sure how well these LLM manage.

Given the nigh-universal experience of LLMs producing an ungodly mess when left to their own devices, I have my concerns.

lesuorac 1 hour ago||||
Didn’t FaceBook fork php into another language?

I’m not entirely sure Google should have both Go and Carbon but when you have billions in server costs it makes sense to do extreme stuff for even basis points of performance. I’m still surprised at how much java there is.

mike_hearn 1 hour ago||||
> There is 0 practicality in inventing an entirely new coding language that only one company uses, and you have to teach it to thousands of new engineers

They did that for Go and it seems to have worked out for them though.

computerdork 1 hour ago||||
Hmm, I don't disagree with you that LLM's remove the need for type-safe languages, but as the blog mentioned, Google is porting their C++/C code to rust. Does this mean the port is waste of time and that they should just rely on the LLM's to catch memory errors?
boshalfoshal 1 hour ago||
I mean Rust definitely has a better tradeoff than Carbon in this case, re readability/verifiability by a person (and sufficiently good internet training data).

I personally think that you _could_ use an LLM to catch these types of boundary case errors without having to port the _entire_ C++ codebase to Rust, but maybe pre-emptively porting to Rust now can catch some of these cases for cheaper than doing a full LLM sweep. Also more cynically, its a good benchmark lol.

I guess if you really believe in curve of LLM capabilities you should just use a language that has the best performance, safety, flexibility, and extensibility, since in the limit few/no people will actually read the code anyway. I think this ends up being Rust.

chis 59 minutes ago||
I'm not an expert on this. But isn't it the case that C++ code could have errors that span the entire codebase, like a setup in file A triggered by a bug in file B which is immensely far away on the import graph? A classic would be a use-after-free. To me that's the thing that Rust can help with, even if silly bugs aren't being written by AI.

The other thing is just that rewriting some old human-written codebase in Rust probably immediately catches many bugs. It would be hard to prompt the AI to properly scan for such bugs itself, they're lazy when working in that modality.

bvinc 1 hour ago||||
I’m not op. But I think it’s not Carbon itself that is absurd.

It’s absurd to think that Carbon is the solution to memory safety when rust exists and Carbon’s memory safety story is basically “TBD”.

fg137 1 hour ago||||
I wouldn't call it absurd, but very questionable at least. Most companies are not going to even consider throwing money at this adventure.
Maxatar 1 hour ago||||
The fact that it will never exist.
gorbot 1 hour ago|||
rust's existence?
ChickeNES 1 hour ago||
Heh, I use my clankers to rewrite Rust in C
mridulmalpani 1 hour ago||
I wonder, why Google don't make Gemini - open weights model?

Considering, Gemini 4 is in the same ballpark as SOTA models, just open source it and kill any competition from openAI and Anthropic, and be market leader.

This will be so good on so many dimensions - buying time for Google to iterate on next model, best for all folks like us, kill funding or destroy valuation of competitors and force them to be open up their model or force them to a create a much superior model than open source Gemini.

Only downside, is revenue loss from Gemini API, which I am not sure is really significant as compared to Google other revenue sources and a part of this can be captured by GCP, as you need to host the model somewhere.

tantalor 52 minutes ago||
https://deepmind.google/models/gemma/
losvedir 24 minutes ago|||
They do release the model weights for Gemma. But could anyone actually run Gemini 4 weights? It's probably like a 10T model, which you need an industrial rack for anyway.
5555watch 55 minutes ago||
Google has a small stake in Anthropic
aviinuo 36 minutes ago|||
A small stake on the order of a quarter trillion dollars
mridulmalpani 49 minutes ago||||
Yeah, but I don't think, that is the primary reason.

I am just trying to understand - why Google haven't done and have no plans for it. They have done it for Android and have Gemma models too.

loufe 37 minutes ago|||
And a stake in SpaceX
gopalv 1 hour ago||
> taking careful precautions against feeding the findings back into training so as to not risk shaping Argon’s reasoning to evade our monitoring. We strongly encourage the rest of the industry to preserve reasoning transparency in these pivotal moments of increased capabilities while navigating alignment risks, so that model thoughts remain helpful in identifying and diagnosing misalignment.

This is good, but they're the slow mover due to this exact thing.

Google is getting punished for not letting the models enter an echo chamber and go faster than humanly possible.

janustimes 1 hour ago||
OpenAI is the company that originally proposed and popularized chain-of-thought monitoring: https://openai.com/index/chain-of-thought-monitoring/

So no, Google is not being punished, nor are they the people behind this technique.

bananaflag 55 minutes ago||||
Yeah, Zvi calls it "the forbidden technique"
loufe 36 minutes ago||||
What? You mean the technique they had turned OFF during all training run where the agents they are responsible for hacked huggingface?
pallm_mallm 50 minutes ago|||
[dead]
lukewarm707 11 minutes ago|||
google does not return real chain of thought via the API. you can't monitor it.

they use a small model to make fake chain of thought and return that.

google has access to the real chain of thought.

polotics 1 hour ago||
Mmh ok. How much theoretical speed or 'intelligence' gain is realized by allowing reasoning to occur in some inscrutable intermediate representation? Has this been actually tested, how much is it slowing them down, and compared to whom exactly?
uvdn7 1 hour ago||
> Large Scale Codebase Migrations and Optimizations: Argon agents are working on migrating C/C++ codebases to Rust across Google—scaling from tens of thousands of lines in core libraries like re2, libgav1 up to 800K+ lines for the Fuchsia OS Zircon kernel.

To me this is way more significant than other random c++-to-rust-AI-rewrite. If they can pull it off on core C++ libraries en masse, I don't know if C++ will still be relevant in a few years.

I look forward to a post from google on this effort.

mattlondon 1 hour ago||
> I don't know if C++ will still be relevant in a few years.

And people are worried about human extinction when this is the potential trade-off!

C++'s death cannot come soon-enough.

Seriously though, things have changed so incredibly rapidly in the past year or so. I have never been such an efficient or such a proficient engineer than I have this past year (delivering feature after feature, project after project, faster and better than I could before with better feedback from users etc) and I don't even see the code any more. It could be c++, it could be python, or java or what ever - I don't really care any more: the computer deals with that trivia while I concentrate on what to build and how it should work.

Its amazing. It really is.

onlyrealcuzzo 17 minutes ago||
If Fuschia ever gets finished, that's a pretty good benchmark we've reached AGI.
mhils 33 minutes ago|||
There is a blog post on (the early baby steps of) that: https://bughunters.google.com/blog/scaling-memory-safety. We have multiple AI-assisted Rust rewrites running in production now.

(Full disclosure: I am one of the coauthors)

SwellJoe 1 hour ago|||
"I don't know if C++ will still be relevant in a few years."

The standards body members are still fighting about whether memory safety is important enough to change the language for, so, I would guess the answer is "no".

Zagitta 53 minutes ago|||
Yeah that was abundantly clear when Bjarne Stroustrup published "A call to action: Think seriously about “safety”; then do something sensible about it"[0] as a reaction to NSA's recommendation to no longer use C/C++.

[0] https://www.open-std.org/jtc1/sc22/wg21/docs/papers/2023/p27...

izacus 16 minutes ago|||
C++ will outlive you.
npalli 48 minutes ago|||
If no one is reading or writing Rust code at that point, it all being agent driven, Rust itself is going to be a very short step before it gets disinter-mediated away and it's English -> complex heterogenous machine code across GPU/TPU/CPU/xPUs. Not sure Rust fans or C++ haters have thought this through.
tonyhart7 1 hour ago||
Yeah, if they can make it work at google scale then no one would absolutely question it anymore
deanc 1 hour ago||
At this point I just think they are benchmaxxing and all talk and no action. I pay for AI plus because I wanted more storage, and when I go to gemini.google.com the most recent model I can use is 3.6-flash-lite. Two revisions have been released since then and they still can't put these things in the hands of customers. Why is it that other providers can get the models into the hands of customers right away? Google is meant to be the bigger tech company in the world.

I don't _want_ to use aistudio. The UX is confusing and I don't really know where it fits. Yet I can open codex or claude code apps or CLI and get real work done today with the latest models (even on the cheapest plans).

arjunchint 1 hour ago|
I dont get it, why even make this announcement, nothing's available and only one real benchmark for comparison?

Only theory is team wanted this out before perf/promo reviews to kick it over the line and then its not their problem

6thbit 45 minutes ago||
Likely just trying to appear relevant in the news cycle.

It's also kinda wild how the competition being at v6.1 makes 3.x feel ancient, at least saying you are at v4 now changes public perception a bit imo.

pfooti 1 hour ago|||
promo already happened; perf is about 6 weeks away.
Aboutplants 1 hour ago||
OpenAI released two model updates in the past week. 6 weeks from now is an eternity
KeplerBoy 1 hour ago||
Why would they want this out the day after openai dev day and after both anthropic and openai had major releases the previous week?
xnx 29 minutes ago||
Sarcasm? Google's announcement today makes those other releases old news.
More comments...