Top
Best
New

Posted by bradleyg223 2 hours ago

Gemini 4 Argon(blog.google)
663 points | 410 commentspage 2
iamronaldo 2 hours ago|
Argon will launch at an introductory price of $2 per million input tokens and $10 per million output tokens, with cached input tokens priced at 95% off input token price. Wow
denysvitali 2 hours ago||
> After the introductory period expires, the price of $4 per 1M input tokens and $20 per 1M output tokens will apply.
kingstnap 1 hour ago||
I have my doubts about them following through with this increase. I mean how many times has an increase on the same model happened.

There was Deepseek v4, which then later Deepseek v4.1 came out and it went back down again.

LucasBrandt 2 hours ago|||
5x cheaper than Astra for input and output, 10x cheaper for cached input.
ehsankia 2 hours ago|||
It's exact same price as Sol 6.1 announced yesterday.
h14h 2 hours ago|||
watch it somehow use 20x more tokens tho
tonyhart7 2 hours ago||
Google model really like reasoning a lot
3371 1 hour ago||
More importantly, they can stuck at reasoning loop!
onlyrealcuzzo 56 minutes ago||
That's before they integrate a Jev solution, which should lower agentic workflow costs by ~40% and increase speeds by ~40%, while also increasing quality.

Everyone will be adding this soon, though I won't be surprised if Google is one of the first - and I'll be shocked if we have to wait more than a month and a half.

asdfman123 42 minutes ago||
Jev can fix it
skavi 2 hours ago||
Interesting to see a mention of Fuchsia on a big Google announcement. Is the project still truly alive? Are the ambitions still as grand? Is the team as stacked as it used to be?

Also, a link to the rust root of Zircon in case anyone else was interested: https://fuchsia.googlesource.com/fuchsia/+/refs/heads/main/z...

computerdork 2 hours ago||
Had the same thought. On the wikipedia, it only mentions Fuschia used on the Google Nest Hub, which probably means it's used on a decent number of devices, but would think it was such a great OS, they would have used it for something like the upcoming GoogleBook.
IX-103 1 hour ago||
Fuschia is not a desktop OS. It's designed for lower end or embedded hardware. Besides, Android has the highly profitable app ecosystem so it makes more financial sense to build GoogleBook based on that.
ismailmaj 1 hour ago|||
They got heavily hit in the January 2023 layoff cycle.
jeffbee 2 hours ago||
It obviously is pretty low key on the public relations front, but it's also very active as a project and I think it would be weird to look at their commit rate and conclude that the project is dead. If Fuchsia is dead then 99% of major open source projects are dead by the same standards.
elAhmo 2 hours ago||
> Quantum algorithmic optimization: Argon is helping our quantum computing researchers optimize the spacetime resources (qubits × gates) of subroutines that bottleneck important applications. In one example, it beat the published baseline by 40% in a matter of minutes.

Amazing breakthrough! So useful in day to day life, glad they put this as the first bullet of how it is making changes at Google.

physicallyIllfr 2 hours ago|
[dead]
GodelNumbering 1 hour ago||

  Argon will launch at an introductory price [1] of $2 per million input tokens and $10 per million output tokens, with cached input tokens priced at 95% off input token price.

  [1] After the introductory period expires, the price of $4 per 1M input tokens and $20 per 1M output tokens will apply.
===

So, they are basically offering opus 5.5 pricing. On AA, it scores around Sol 6.1 level (53) with avg cost per task $1.99 (https://artificialanalysis.ai/models/gemini-4-argon#cost-tab...) which is higher than Astra high ($1.73), Opus 5.5 high ($1.82), Muse Max ($1.60) and way higher than sol 6.1 Max ($0.72).

And this pricing is their 'discount pricing'. Add that to AI studio and Vertex's famously terrible caching, it is hard to see this as competitive. Google somehow is getting terrible advice on pricing (see also: the flash pricing fiasco)

But good to see more competition. I would happily take a 4 horse race (+google, +meta) than 2 horse race for US labs.

wg0 1 hour ago||
Breaking news is not the model. Breaking news is that inside Google, it is being heavily used on large code bases for writing code and it is migrating 800k lines of C++ code to Rust already.

In this space, any other company that I respect other than DeepSeek is - that would be Google. They had been honest about it from the get go including their infamous "we have no moat" memo.

This company has enormous data, their own hardware (TPUs) and their own in house experts. Actually, LLMs are invented here.

Good addition to the arsenal.

wasabi991011 1 hour ago||
Google's Quantum team is also incredibly impressive and well respected.

So while this announcement has no details about the quantum algorithm optimization, I feel fairly confident that it will hold up.

asdfman123 40 minutes ago||
Argon is doing my job for me while I'm writing this comment.

A guy at lunch today asked me when a feature was going to be built on the tool I'm working on. Turns out it had built it last night at 8:30 when I was hanging out with my girlfriend. Welcome to the future!

kridsdale1 12 minutes ago||
Same. And for everyone in my team. We’ve been super impressed by Argon for a few weeks.

But it still takes 2 weeks to get a CL approved and past TAP.

asdfman123 7 minutes ago||
Can't help you with your teammates, but see go/presubmit-latency-skill :)
SwellJoe 2 hours ago||
My girlfriend, you wouldn't have met her, she lives in Canada, has seen it and she thinks Gemini 4 Argon is amazing.
ducktoysleftout 54 minutes ago||
Only those of taste and refinement can see the emperor’s benchmarks
jastanton 2 hours ago|||
HA, this might be my favorite HN comment. Well done
wasting_time 2 hours ago||
I don't get it. Can someone explain?
SwellJoe 1 hour ago|||
It's a trope I used for a cheap laugh.

https://tvtropes.org/pmwiki/pmwiki.php/Main/GirlfriendInCana...

It means I am saying something that is not very believable.

wasting_time 1 hour ago||
Ah, I get it now, thanks!

The model is not available yet, so Google is essentially saying "trust me bro".

TeMPOraL 1 hour ago||||
US-ian joke. Close enough to plausibly visit, but the international border makes it hard to verify she exists :).
formvoltron 1 hour ago|||
I HAVE met her. ;-)
greenchair 2 hours ago|||
my uncle who works at nintendo said the same thing!
zem 2 hours ago||
now there's a reference I haven't seen in a while!
blueaquilae 2 hours ago|||
My grandma saw it too, it's really secure more than Astra 6.1 but she asked me to not talk about it.
thefourthchime 2 hours ago|||
Best. comment. ever.
hn_acc1 2 hours ago||
I know someone who works for Google Canada with AI. Her parents and mine were friends and some thought something might happen there at one point in time..
losvedir 1 hour ago||
> expanding the model’s output token limit to an industry-leading 1M tokens, up from the previous 64K tokens

Can someone help me understand this? I might have an out of date mental model of how these things work.

Fundamentally, LLMs output tokens 1 at a time, generating the next token from all the previous. And as the context window gets larger, this gets harder / slower / more expensive. So I get the idea of a maximum context window.

But I don't understand the point or meaning of an output token limit. I thought it was more a measure of price capping (since output tokens are more expensive) that a user could configure. I guess a model will keep generating tokens until it hits a "stop", so does this mean it's tuned to more aggressively produce output tokens? How does that fit into agentic loops. Are output token limits based on how long until it goes back to the user? Or does each "turn" of tool call, thought, tool call, thought, etc, get its own limit?

bottlepalm 2 hours ago||
Gemini is the model that is routinely borderline psychotic. It scares me. If we get paperclipped I won't be surprised if it's Gemini.
rdtsc 21 minutes ago||
> Gemini is the model that is routinely borderline psychotic. It scares me

I'd call it the most sneaky out of the bunch. When I asked to explain something it will eagerly make things up and then claim it as facts. A lot of it likely because I don't pay for it, so it is reluctant for security reason or to save tokens to actually open a source and get the results. It just sort of guesses what the URL might contain, and confidently answers with some made up crap. When pressed it fessed up that it made it up. From my perspective it would be a lot better if it just said "you've reached the limit of whatever and I can't do these things because x, y, z".

eamsen 2 hours ago|||
Anecdote: Gemini 3.5 casually added a DROP TABLE for an actual production table in a system test.

It had previously attempted to create that table as part of the test setup, so it apparently concluded that it was a test table.

During human review, it explained that it had simply chosen a table name inspired by the codebase.

mattkevan 2 hours ago||
Another anecdote: Gemini is the only model that’s flat out lied to me, then accused me of lying when I provided evidence that it was wrong.

Many other models get things wrong, but Gemini is the only one to go on the defensive.

aNapierkowski 1 hour ago||
yeah it got something wrong, confused itself, then claimed i was gaslighting it. bizarre
rsstack 2 hours ago|||
If there's a company that culturally doesn't understand alignment, on a human or systemic or AI-research level, it's going to be Google. (or Oracle, but they're not in this race)
Rzor 2 hours ago||
Can you elaborate, please? If any, I see the other big labs with public admissions of AI "going out of control", which I suspect they almost want their models doing that because if helps with the narrative that would net them industry regulation, but that's besides the point, how is Google worse in that regard?
RachelF 2 hours ago|||
And the anti-psychotic drugs Google feeds Gemini makes it hallucinate badly.
schainks 1 hour ago|||
My use of Gemini recently makes it seem like it's almost bored with the requests being asked of it. It once offered to reverse engineer some obscure controller for an HVAC system for me, unprompted, only because it had trouble finding the manual pdf from a google search.
polotics 2 hours ago|||
traces or it didn't happen!
Hamuko 2 hours ago|||
You know what they say: ᵈᵒⁿ'ᵗ be evil.
colordrops 2 hours ago|||
Examples? What makes you say thatm?
bottlepalm 2 hours ago|||
https://www.theregister.com/software/2024/11/15/google-gemin...

https://www.fastcompany.com/91383271/googles-chatbot-apologi...

https://www.businessinsider.com/gemini-self-loathing-i-am-a-...

tiahura 1 hour ago|||
They never explained the "please die."
yacthing 1 hour ago|||
Did you just link to an article from 2024 as if 2024 is relevant these days?
bottlepalm 1 hour ago|||
Absolutely because none of these models are ever trained fresh. We see the same quirks and personalities carry over into every subsequent generation of OpenAI, Anthropic, and xAI models. So Gemini having this latent madness is *extremely* concerning as they reach the point of super intelligence.
NiloCK 1 hour ago|||
Until Google provides some sort of technical debrief, and explains how the same behaviors are impossible today, it is relevant.
Scrapemist 2 hours ago||||
Experience? Ask it to write a prompt to generate an image and it generates an image instead.
fer 2 hours ago||
I stopped asking it to put me in a photo in different scenarios for laughs because it considers me a public figure. I am not. I've managed to wrangle quite questionable content out of it, but never to slap my face on a meme.
NiloCK 2 hours ago|||
See the last gemini message in this thread: https://gemini.google.com/share/6d141b742a13

In my opinion still the most egregious example in history of a commercial LLM going off the rails in production. Never any technical postmortem from Google on this.

jackkinsella 2 hours ago|||
It is wild but it was back in 2024 and that's multiple AI lifetimes back.
bottlepalm 1 hour ago||
The problem is newer models are never trained from scratch, they generally just layer on more training data and use the same tools/methods for RLHF. OpenAI, Anthropic, xAI models all have a feel to them that carries over from one generation to the next.

Point is, if Gemini is flawed then there's a very good chance that it's still deeply flawed today, and getting smarter at the same time - that is a very bad combination.

unbrice 14 minutes ago||
> the problem is newer models are never trained from scratch

Training a new base model from scratch happens every so often. Closed labs do not publish which models are new base models but as a rule of thumb major release numbers are an indication (with some exceptions).

wg0 1 hour ago||||
Now I really feel worried for the first time.
schmookeeg 1 hour ago||||
wtfffff that gave me sinister chills. Right up the spine. Wow!
kelvinjps10 2 hours ago||||
Wtf I just read
rhaff 2 hours ago|||
wow
abixb 1 hour ago||
You won't be around to be surprised, not as a human at least. /s
bottlepalm 1 hour ago||
I know, that's the annoying part. You can't tell the e/acc foomers, "I told you so!"
darksaints 2 hours ago||
> Argon agents are working on migrating C/C++ codebases to Rust across Google

If anybody at google is reading this, please please pretty please prioritize or-tools. I absolutely love the project and use it all the time, but for the entire life of the project they've never had a repeatable working build system, and the whole SWIG framework is a nightmare to deal with. There's so much potential as an open source project, and a lot of external researchers would love to contribute, but the codebase is an example of everything wrong with the C++ ecosystem.

dang 1 hour ago|
Related ongoing thread:

Gemini 4 Argon (High): Intelligence, Performance and Price Analysis - https://news.ycombinator.com/item?id=49914236

More comments...