Top
Best
New

Posted by bradleyg223 13 hours ago

Gemini 4 Argon(blog.google)
See also: Gemini 4 Argon (High): Intelligence, Performance and Price Analysis - https://news.ycombinator.com/item?id=49914236
1345 points | 862 commentspage 15
landdate 8 hours ago|
I don't like AI, but it's very enjoyable for me to see my predictions on the success of gemini come to fruition.

I only use gemini, and while I don't use it for actually writing up code, I use it to help me troibleshoot my logic and help find bugs. Its easily the best model I have tried. And yes I am talking about 3.1 pro.

Also I have found gemini the only model to be the least likely to douse me in flattery, and will follow my pre built instructions to never output anything unless it can be directly sourced, pretty well. Chatgpt i tried for a bit and it was by far the worst thing I have ever used. I can understand why people develop psychosis when prompting chatgpt because it is disgustingly scyophantic to the point I was grossed out and felt like I just got done with some other type of self gratification.

Anyway, death to AI. All those who use, create, facilitate, or even just sit by and do nothing in the face of AI will perish in Hell.

VirusNewbie 13 hours ago||
It's fucking insanely good.
nananana9 12 hours ago|
That's good to hear, recent models haven't great at this particular use case.
tamimio 13 hours ago||
Now AI models will turn into vaporware, a bunch of numbers on a table without even releasing the model, because it’s toooo scary to release!
gravisultra 12 hours ago||
Google has the audacity to "protect us from ourselves" and talk about "safety" and in the very same blog post highlight the Israeli "security" company Wiz, that they acquired for a very exaggerated sum of money.

This is why I will never take any of these leading model houses seriously when they talk about alignment. They are literally complicit in genocide and the worst crimes against humanity imaginable.

pliiight 13 hours ago||
Hate to say i will never be touching this model for anything except for youtube video understanding
small_model 9 hours ago||
5 points off Opus 5.5 on AA, not a good release. Falling behind and not able to catchup. Ant probably has opus 6 in the works. Fumbled so hard on this, they should have owned AI.
wewewedxfgdf 13 hours ago|
Gemini is so far behind that it is effectively useless compared to Claude.

It's a surprise that Google has let themselves lose the game given their infinite cash, massive computing resource, gargantuan information store/training data, and vast number of programmers.

The truckloads of ads revenue mean they don't have the single focus drive needed to win.

jjice 13 hours ago||
We're like 3.5 years into this new era - I'm not counting winners or losers yet.
mattlondon 12 hours ago|||
How is it far behind? The benchmarks published in the blog post show it is superior to Opus 5.5 and Astra 6?

Behind how?

wewewedxfgdf 12 hours ago||
Within one question of their web interface, it has lost context and asks you to clarify what you are talking about.

I am very often giving the same programming task to multiple LLMs for various reasons - the answers from Google are so bad that I gave up.

I have no interest in benchmarks.

mattlondon 12 hours ago|||
So you have no experience of their latest model release then? Just repeating the usual tropes about Google having messed up? Or basing your opinions on their website chatbot?

If you have actual independent benchmarks and evidence about how this new model release is "so far behind" and refutes the stuff from their blog then please do share because I think we'd all love to see that?

wewewedxfgdf 12 hours ago||
No I am commenting on my real world experience of using Gemini daily. I still ask it questions alongside Claude and OpenAI and Gemini is always the worst of the three.
mattlondon 12 hours ago||
So you've not used this new release then? So how can you say that they are "so far behind" if you are not using the most recent model for your comparison. This is their first 4.0 model, that you are not using and instead basing all your opinions on on some ancient months-old model from a previous generation?

With respect, I don't find your arguement about them being "so far behind" especially convincing when you are using previous-gen releases and not actually using their current release.

fwip 12 hours ago||||
Yeah, they've definitely got some recurring tooling/infrastructure problems around the models.
bel8 13 hours ago|||
I wonder if Google bans internal use of Claude/Codex.

And I wonder if Google's main monorepo is already in Anthropic/OpenAI training data because of some stubborn dev.

krat0sprakhar 12 hours ago|||
(I work at Google) Yes, internally we all use Jetski (internal version of Antigravity). Outside of Gemini, Opus models are supported and allowed for internal use. No OpenAI models since they are not on Vertex
heyjamesknight 11 hours ago||||
No way to run OAI on a machine with monorepo access even if you wanted to. Claude runs on Vertex so it's not leaving Google infrastructure.
lunarboy 12 hours ago|||
Claude used to be GDM only, but recently opened up Opus for all googlers
ASalazarMX 12 hours ago|||
Funny how we start to see people supporting LLMs like we support sport teams, political parties, or celebrities.

- Person 1: X is garbage compared to Y!

- Person 2: Why?

- Person 1: Because I like Y.

dhdjcjcjnd 12 hours ago|||
Google's strategy is to let their competitors bankrupt themselves while they continue to offer good-enough models near breakeven.
LoganDark 13 hours ago|||
I've tasted Gemini through an intermediary and it feels far better at attention to detail than other models I've tested (Claude Opus/Sonnet, GPT whatever it's called nowadays). But it's less likely to get one-shots right.
gniv 12 hours ago|||
They are playing a longer-term and more enterprise-oriented game.
georgemcbay 12 hours ago|||
> Gemini is so far behind that it is effectively useless compared to Claude.

I fundamentally don't understand LLM "brand loyalty".

All of the models are constantly leapfrogging each other and always have been.

Google had a long lag between releases (and still hasn't released Argon), but why wouldn't they be able to compete? It isn't like any of this stuff requires secret knowledge, the Bitter Lesson has proved true again and again, and Google can certainly scale computation, it is like the one single thing they've always done well in spite of all their other foibles.

wewewedxfgdf 12 hours ago|||
Its not brand loyalty. I use them all the time and have no loyalty - I'd happily ditch an LLM for better results - that's how I got to Claude from ChatGPT.
786562354238 12 hours ago||||
Gemini has never ever leapfrogged any competitor.
singingtoday 12 hours ago|||
I hope it can. Today it is very far behind.
VirusNewbie 13 hours ago|||
-
handfuloflight 13 hours ago||
You have access to Argon?
osti 12 hours ago|||
Google employees do.
matthewfcarlson 12 hours ago|||
Their profile says: > Currently at Google as a Sr. SWE SRE on the cloud.