Top
Best
New

Posted by bradleyg223 4 hours ago

Gemini 4 Argon(blog.google)
See also: Gemini 4 Argon (High): Intelligence, Performance and Price Analysis - https://news.ycombinator.com/item?id=49914236
849 points | 561 commentspage 4
losvedir 2 hours ago|
> expanding the model’s output token limit to an industry-leading 1M tokens, up from the previous 64K tokens

Can someone help me understand this? I might have an out of date mental model of how these things work.

Fundamentally, LLMs output tokens 1 at a time, generating the next token from all the previous. And as the context window gets larger, this gets harder / slower / more expensive. So I get the idea of a maximum context window.

But I don't understand the point or meaning of an output token limit. I thought it was more a measure of price capping (since output tokens are more expensive) that a user could configure. I guess a model will keep generating tokens until it hits a "stop", so does this mean it's tuned to more aggressively produce output tokens? How does that fit into agentic loops. Are output token limits based on how long until it goes back to the user? Or does each "turn" of tool call, thought, tool call, thought, etc, get its own limit?

rcr-anti 2 hours ago||
According to Artificial Analysis, one metric is standing out significantly: hallucination rate. Beats frontier models by a good margin at 15%, while latest OpenAI are in the 40s-50s and Anthropic in 60s-70s (mostly). Other near frontiers are closer, Grok 4.7, GLM5.3, and Muse Spark 1.3 are all around 30%. Only other model I recall getting close was Minimax M3 at 18%.
huydotnet 1 hour ago||
I recently subscribed to Claude, and was very unhappy about the usage limit of the $20 pro plan. Then I found a trick, since I got free Google AI Pro via my phone carrier, I use Opus 5.5 High for planning, and then dispatching agy to do works.

Most of the time 3.8 works fine, but it's a bit slow if compare to 3.7 Flash. If there's already a detailed plan, 3.7 can complete the task much faster. And the best thing about agy is the usage limit was very generous.

pietz 3 hours ago||
I know companies benchmaxx, but after what Google pulled with Gemini 3.8 Flash, I give zero f*cks about any numbers they report. No other model on Artificial Analysis dropped harder after they adjusted their weighting. Just look at their DeepSWE scores and then try to do any serious coding with the model.

Google is desperate. They haven't been performing in half a year. It's clear their researchers have been forced to integrate existing benchmarks into their training.

These numbers are meaningless. Shame on them.

bobkb 3 hours ago||
IMHO Google first needs to make it easy for humans to find where to find the models and its documentation. With aistudio/model garden / Gemini enterprise etc it takes minutes to find the model.
dang 2 hours ago||
Related ongoing thread:

Gemini 4 Argon (High): Intelligence, Performance and Price Analysis - https://news.ycombinator.com/item?id=49914236

yzydserd 3 hours ago||
"argon" is derived from the Ancient Greek word ἀργόν meaning lazy or inactive.
thefourthchime 3 hours ago||
I was just thinking, I bet if I refresh hacker news, a new model will come up.
scirob 4 hours ago||
"Rolling out soon" don't let them hype without any release
itzikkatz 3 hours ago|
They waited a whole year—until the "free year for students" promotion ended—to release their flagship model. I can't believe I've been stuck with a crappy model like the 3.1 Pro until now.
More comments...