Top
Best
New

Posted by bradleyg223 3 hours ago

Gemini 4 Argon(blog.google)
See also: Gemini 4 Argon (High): Intelligence, Performance and Price Analysis - https://news.ycombinator.com/item?id=49914236
770 points | 509 commentspage 3
treefry 1 hour ago|
The benchmark scores are not exciting. I understand Google is playing a catch up game right now, but still wish to see some overwhelming improvements.
woko 57 minutes ago|
Gemini 4 Argon (High) looks comparable to Claude Opus 5.5 (High). https://artificialanalysis.ai/models/comparisons/gemini-4-ar... From that perspective, the benchmark is not too disappointing, given that there are only 8 days between the blog posts (September 30 vs. September 22).
iamben 1 hour ago||
Wonder if this one will be smart enough to run the automations in my Google Home that all broke now they've forced Gemini to replace the Google Assistant.
darksaints 3 hours ago||
> Argon agents are working on migrating C/C++ codebases to Rust across Google

If anybody at google is reading this, please please pretty please prioritize or-tools. I absolutely love the project and use it all the time, but for the entire life of the project they've never had a repeatable working build system, and the whole SWIG framework is a nightmare to deal with. There's so much potential as an open source project, and a lot of external researchers would love to contribute, but the codebase is an example of everything wrong with the C++ ecosystem.

throwitaway222 1 hour ago||
While this is exciting - I have no clue when/if I will be able to use this. Unlike OpenAI or Anthropic where a model release announcement == GA.
helsinkiandrew 3 hours ago||
> Google Grapples With Employee Skepticism About New Gemini Model

https://www.bloomberg.com/news/articles/2026-09-30/google-gr...

bitexploder 2 hours ago||
Opinions my own but I have been using this model for a bit. I would say it is a good model and the skill with which people use AI varies widely.
shawabawa3 2 hours ago||
Hardly gushing praise for an internal only bleeding edge model

I would guess it's like opus 5ish level from this

bitexploder 1 hour ago||
It is definitely smarter than that. It is mostly mannerisms and how it likes to work. I would take it seriously as an Astra or Fable or Opus 5.5 level model. It just needs polish, but where and how you harness and use it matters a lot. But it has amazing long horizon attention and gets things done.
readams 2 hours ago||
Skepticism is gone now.
dom96 3 hours ago||
Why announce this if it’s not available yet? Why not at least announce when it will be released to the public?

None of the other AI labs do this. Really frustrating.

gengelbro 3 hours ago|
Mythos?
aqsnow 28 minutes ago||
Yes but google always does this crap.
moostii 1 hour ago||
Excited to see Google competitive at the frontier level again. Hopefully they sort out their infrastructure and model versioning so that we can feel confident building production applications on top of their APIs. The capacity limitations I've experienced with them in the past have been deeply problematic.
losvedir 1 hour ago||
> expanding the model’s output token limit to an industry-leading 1M tokens, up from the previous 64K tokens

Can someone help me understand this? I might have an out of date mental model of how these things work.

Fundamentally, LLMs output tokens 1 at a time, generating the next token from all the previous. And as the context window gets larger, this gets harder / slower / more expensive. So I get the idea of a maximum context window.

But I don't understand the point or meaning of an output token limit. I thought it was more a measure of price capping (since output tokens are more expensive) that a user could configure. I guess a model will keep generating tokens until it hits a "stop", so does this mean it's tuned to more aggressively produce output tokens? How does that fit into agentic loops. Are output token limits based on how long until it goes back to the user? Or does each "turn" of tool call, thought, tool call, thought, etc, get its own limit?

waldrews 2 hours ago||
Dear Google, please don't turn off your old generally available Pro-class model before your new Pro-class model is generally available (previous discussion https://news.ycombinator.com/item?id=49668196 )
rcr-anti 2 hours ago|
According to Artificial Analysis, one metric is standing out significantly: hallucination rate. Beats frontier models by a good margin at 15%, while latest OpenAI are in the 40s-50s and Anthropic in 60s-70s (mostly). Other near frontiers are closer, Grok 4.7, GLM5.3, and Muse Spark 1.3 are all around 30%. Only other model I recall getting close was Minimax M3 at 18%.
More comments...