Posted by leumon 1 day ago
I've been using Gemini to live chat in Afrikaans and do impromptu Afrikaans grammar lessons during my solo drives around town. It is phenomenal at speaking the language - like, it really shocks my family members when they hear it.
This is probably the most joy I get from any of my usages of LLMs/AIs. It's been really, really nice getting to speak my language regularly again. =)
So, I'm excited about this release and live chat getting better. I also hope the other frontier labs pick up niche languages like this as well so that I have more options.
I once asked it to summarize The Hobbit in Catalan to explain it to my daughter before sleep. I was expecting a lot of mistakes as I see regularly if I ask anything in my native language when using GPT or Claude, but it was surprisingly good. I was going just to kind of skim ahead and retell it my own way, but ended up almost saying it verbatim because it was good already.
She loves Zelda so I asked it to explain the story of Breath of The Wild keeping the original names, and to make it fun, etc.. I was surprised again. I did retell some bits in my own style and taste but it is very convincing.
I haven't tried Catalan on newer models like GTP-6 Astra or Fable tho. We have all these benchmarks based on software development, and AGI, etc.. but it would be cool to have some language benchmarks for different communities.
As I work in english and use them in english, I wonder if using LLMs in a different language to code renders a different result as well. Like, if some of these benchmarks were made in other languages, would the result be similar.
I am not a Chinese speaker but my understanding is that all of the models have substantially different behavior in Chinese, to the degree that it's kind of like a second model. Would be interested in hearing more if anyone has direct experience.
Relatedly, when chatting with frontier models about science education content design, I've found it very helpful to mix in Chinese education terms. In English, for example, NGSS is such a massive attractor, discussing nearby topics often yields NGSS "slop". And "estimation" (educational) in the US means one (dysfunctional:) thing, which similarly distracts. Perhaps if AIs become increasingly multilingual, but remain weak at deep conceptual reasoning, it may be fruitful to have multilingual thesauruses, to use language-associated cultural conceptual differences as a way to convey conceptual nuances with which LLMs otherwise struggle?
The one criticism of NGSS I found from a quick skim of https://en.wikipedia.org/wiki/Next_Generation_Science_Standa... is that it teaches evolution.
What does "estimation" mean in the US?!
"Estimation" as a skill in US education, primary school on, is pervasively about point estimates - single numbers. Bounding, a range of safe/reasonable/likely/possible numbers, is very rare. In Fermi questions/problems also.
Even in such goodness as Sanjoy Mahajan's Art of Insight and Street-Fighting[1]... the only only "bound" mentioned is "Printed and bound in [USA]", and there is no "round".[2]
In contrast, IIUC, China is pervasively about "Da Gu" and "Xiao Gu", "round up" and "round down". Point estimates are secondary. Lower and upper bounds are apparently even a cultural concept: an individual's degree vs capability/fate; policy equity vs excellence; economic targets.
In education, my understanding is point estimates have much poorer conversational/collaborative dynamics than range estimates. Bounding lends itself to incremental collaboration. "Can anyone suggest another low or high bound?". Discussions around point estimates seem usually merely about the collection of individual estimates.
> NGSS
Ah, NGSS isn't bad particularly. I meant that discussion/mention of NGSS appears so very much in English training data, that AI chat on topics related to NGSS seem to draw regurgitations of NGSS. AI "slop" in the sense that "ignore NGSS, just reason from a superset of its underlying concepts"... isn't an LLM strength.
> hard to understand
Sorry - Thanks for asking!
[1] open access: https://direct.mit.edu/books/oa-monograph/5345/The-Art-of-In... https://direct.mit.edu/books/oa-monograph/5339/Street-Fighti... [2] search "bound": https://www.google.com/books/edition/The_Art_of_Insight_in_S... https://www.google.com/books/edition/Street_Fighting_Mathema...
Yes, and that brings in a shit load of slop to the problem space from the creationists.
Like how the word life will have a strong correlation with death, the word animate, while it also means alive will have a much less strong correlation with death.
Or do you mean professionally... Instead of merely searching for "ways to do X", one might translate the query into a couple of languages, and translate the results, to get a broader culture sampling. Hmm... That seems something an AI-based search could integrate smoothly - cross cultural insights as a normal part of search results?
While there's a lot of diversity in policy and culture among US regions and states, contrasting them isn't leveraged much. Perhaps failing a "that's analysis, not news" threshold for being part of public dialog? So maybe it's not "non-English cultures" neglected as much as comparative discussion in general?
One "criminally underutilized" which really struck me, was Japan having a far more sensible and powerful COVID safety one-liner than the US managed - "Three C's" (Closed spaces, Crowded places, and Close-range contact - poorly-ventilated / crowd / conversation-contact - avoid overlapping all three) vs "6-foot rule".
For that particular attractor, even German or Spanish or so might help avoid it? Or perhaps even just using British English?
How did the early Roman Empire interact with Greek city states?
How are LLMs planning on learning new information on the fly without new context or retraining runs?
Why did mom leave?
You know, standard stuff.
Reinforcement learning by human feedback.
Even as a human being who has a working brain and all this thinking power, I am not sure how I would answer that. It's kind of funny Gemini will still come up with something and even suggest follow-ons to the conversation but saying all that stuff as a human to another human would be a wild response to that question.
It tells me there's a Costco 10km from me and I'm like, yeah let's go there, that's the one I meant! And then it sends me to one that's like 40 minutes from me all across town and I'm like WTF?!
Never mind asking Android Auto any question that's not driving related. It just doesn't understand me at all. Or even some driving related ones. "Find Alternative route". "Sorry, I can't help you with that". WTF?
As jy wil, ons kan saam praat op Discord :) maar my Afrikaans is sleg
If you're really interested, shoot me an email.
What I've been building, with the help of a few language teachers, is our best guess at a really good way of learning new languages from the comfort of your home. It's not ever going to be perfect, but it is pretty good (at least so far in testing within our small group), and very cost-effective.
If you're interested in funding it, I'm happy to go live in Japan for a year :)
You can use it as supplement or do it standalone on your commute before a trip and end up with good pronunciation, vocab and a conversational skills.
For language learning, translation, research - it is fantastic.
It works very well for certain dialects as part of a local language heritage as well: From Welsh to Bavarian - it is a joy to really get approval from people used to these dialects as being really accurate.
Happpy times! :)
I find no greater joy than listening to a TTS model say Poes (Google translate is one of the funnier ones to make swear). Reminds me of being 14 lol.
It makes total sense it’s good at regurgitating the edge cases of language.
Copes well with thick accent, voices are pleasant and latency seems low.
Oh and I can actually use it on a workspace account - which for most of the recent releases was an account stuck in limbo. Not personal enough for personal offering, not enterprise enough for enterprise.
Well done G - will definitely be using this
And looks like one can trigger live mode via siri
This difference is probably due to Google's web knowledge and free pass to YouTube. However, I'm happy what I got from it so far.
When I ask the question once in a blue moon, it can generally one-shot the answer, even.
This is frustrating because when I discuss AI with laypeople they think it's still incapable of counting the number of Rs in "strawberry." They believe it to be essentially useless and incapable of basic tasks. Which, to be fair, is the case with the free models.
I totally disagree. I pay for both ChatGPT and Anthropic (haven't tried the chinese models yet) and yet Gemini is my go-to model (and I pay for it too through Google Workspace subscriptions for several domain names tied to Google/GMail) for anything that is not coding.
I find Gemini better/quicker/more polished for basically every single subject out there that is not "write me lines of code".
I use gemini for everything not-coding, from doing research, to have custom personas for more niche topics (and feeding more detailed knowledge in these cases).
For coding and image editing, right now I find ChatGPT superior. And for software architecture designs or planning Claude is the best since a while. I still have to try Grok to be fair.
One of the problems is local zoning laws regarding an expansion of my house. Comparison of annex vs extension, boundaries, precedent, costs, etc.
3.8 Flash didn't check most of the required zoning laws. It relied on parametric knowledge, which is outdated and inaccurate. It checked zero precedents. It did made a very cursory check of the boundary area, but didn't validate it, so it missed a lot of important nuance and exceptions to the boundary. Its cost estimates were wildly inaccurate. Ostensibly because it was inferring an average based on historical pricing data rather than gathering current info.
I could go on, but if I had to judge this attempt I would give it a 3/10. It's very fast, but wildly inaccurate. It's clear that the model is designed for speed over accuracy.
But don't take my word for it. [Most benchmarks show it to be significantly below frontier models like Astra.](https://llm-stats.com/models/compare/gemini-3.8-flash-vs-gpt...)
This has been a useful exercise. It's important to understand the developments taking place. I am disappointed to see that Google has made very little progress in six months relative to the frontier labs.
If you really want to compare apples to apples you need to test Gemini models against other models using the same third party search harness.
Otherwise you are largely measuring how much computation the model provider is allocating to a search harness.
Pro tip, ChatGPT is the normiest of all normie websites right now. You're not part of the cognoscenti just because you learned how to type prompts into one of the most popular websites in the world.
P.S. You're probably not using OpenAI models for complex or non-standard tasks. It shits the best just as often as Qwen when you need precision and detail in a non-obvious problem.
The benchmarks clearly show otherwise. This is your cue to tell me the benchmarks are made by the Illuminati and only your superior and subjective methods of evaluation are correct.
It does also seems Gemini's main problem wasn't that it was stupid, but that it was good at doing slightly different things than what I asked it to, very well. Which might well also have to do with me being better at wrangling DeepSeek's quirks than Gemini. Still, at that price tag, it's not worth it.
By all accounts they are far more expensive than DeepSeek, and vs. Gemini I've found out that myself.
Edit: For those who are not familiar with it, this model is quite a bit faster, and about 100x cheaper, per token, than Fable and Astra.
Take your meds.
I was like you before, Gemini was the worst model to me.
Then 3.8 came out. At first I was sceptical, but this model *is* able to do useful things ! Complex things.
Of course, it is NOT perfect. But for things like small/medium complex tasks subagents, it's perfect.
Now, is it worth the money vs Astra ? I don't think so. But still my point remain relevant.
Claude does it occasionally but it's a more a soft landing earlier context seems to be compacted, not completely lose the plot.
I just cancelled my pro subscription. I really wanted it to be good but not yet.
I couldn't even get it to tell me how to pay for antigravity, it sent me on some fruitless paths and eventually said "you shouldn't pay for this, it's too hard to figure it out."
they'll live in my rc for decades
I can't think of a better general purpose model than 3.8 flash right now. It also writes more naturally than the other big models too.
Astra wins handily on every other benchmark, too. Research, science, 3D, visual, Humanity's Last Exam, etc.
However, if I want it to DO something then Gemini is in absolute last place. I don't trust it for anything more than renaming files that I don't care about very much or extracting data (though it's too expensive for data extraction at scale).
It’s for this exact kind of scenario where a random question pops in to my head.
Plus, it’s the most grounded by real live data of all the chatbots.
Silicon Valley people are majorly sleeping on Google Search AI Mode.
People use text with LLMs but it's great to have a high fidelity "analyze this image"
I think they’ve made a shrewd move in focusing on search integration and everyday users (Gemini app) vs software power users. They have their corner and nobody is really competing with them, plus it feeds directly into their existing revenue stream.
Maybe everybody is wrong about how valuable these companies will be, but atm it looks very dumb to not be competitive in coding.
Occam's Razor is overwhelmingly that they just don't have the organisational capability to capture this market. If they did then they absolutely would have.
If OpenAI or Anthropic go away tomorrow, people find a new model and move on. If Google goes away tomorrow, people’s lives would be severely disrupted, and in the case of Gmail / Drive / auth access, even temporarily collapse.
Google in a sense won (me over) like that. I also expected them to brute force their way into everything and dominate. This is how it played out though. Image generation is great as well, but ChatGPT one is more lenient on copyright and nannying - for example when my kid asks me to "take a photo of him and Sonic". Gemini cops out either because of the kid or Sonic, disappointing us both, but ChatGPT can be.. persuaded.
Google has essentially give up the race for the frontier. The gap is widening very fast. Their strategy appears to be to focus on niche areas like voice, small models for on-device (see their deal with Apple), and specialised models for tasks like search, which are very inference efficient.
Ask yourself, how does Google -- a company that famously does everything -- benefit from SWEs outside of Google having access to powerful coding models? They would just be competition.
Famously, Google just eventually discards almost all businesses that don't have the same fire hose of revenue that ads does. Selling coding plans isn't something they are going to want to do.
Google is clearly motivated to make better search and information finding tools and stuff that will ultimately drive users through their existing search/ads/youtube ecosystem. That's really why they're in Android, that's why they do Chrome. Everything else with them is a sideshow.
Google is also full of beancounters obsessed with data centre quota and resourcing. Even massively profitable ads projects have to justify and fight for it. (Source: used to work there).
I can't think of anything less resource & revenue sensible than providing outside parties access to your TPUs for the purpose of letting them write stuff which could just end up competing with you.
Yes, maybe as part of their cloud business, selling token access could be useful money. But I doubt they'd tune it for coding.
Bragging rights to say they have a SOTA model. I guess that was more like Google of 10 years ago with moonshot projects. Nowadays, yeah, perhaps if it's not helping sell ads, it doesn't make sense.
It would be utter stupidity if they keep competing with hyper agile frontier labs. They sensibly moving to big infra provider and that's good niche.
> .. perhaps if it's not helping sell ads,
Perhaps cloud business is also a thing which is making big revenue
At worst Google just need to keep the lights on until the frontier labs file for bankruptcy under the weight of that >$1tn debt they've built up paying for FAANG's data centers. Then Google walks away from all of this with the the best remaining models, best R&D lab, and several hundred billion dollars in infrastructure on their balance sheet.
Catch is, even Googlers internally do not have access to top-tier models. (or did not until recently, when apparently Claude was made accessible to the SWEs internally).
Funny enough, after trying it, I went back to G3.8
these are essentially monthly releases, the dated releases that Deepseek and Qwen do make more sense
As opposed to what, them not having it and burning money that isn't their instead like openai and anthropic? At least Google is feeding itself instead of having to create a bubble to stay alive
Big rich companies take on debt for reasons that are sometimes inscrutable from the outside. Recently, they have been borrowing for ~5%, about a half point above what the US government gets for 10-year Treasuries.
Apple has been financing operations with debt for a number of years as part of a complex optimization plan.
No, Google is not broke.
Some expert wall street analysts discussing what they found and how they dissect things, have a healthy skepticism of Big Ai
If my quick search is correct, Google is sitting on a quarter trillion dollars in cash and marketable securities. They could keep doing the negative cash flow thing at this scale for another decade.
Most of the time I don't need what the bench tests and I'm not really giving them completely ambiguous tasks without any refinement.
I only find marginal differences between models at this point and it almost feels like personality quirks in each model than anything.
alias agy="agy --dangerously-skip-permissions"Anecdotally I'd rate Gemini behind Claude and OpenAI models at fiction and I can't find any benchmarks showing Gemini is the clear winner at this task.
I doubt they even intended it to be, but it seems like I kept going from resorting to 3.5-3.8 (over time) to realizing that Claude and GPT, while great at Python, will make rudimentary mistakes with R; even when they compose giant complicated R code.
I'm worried in their push to catch up on the SOTA front, it's going to lose that natural sounding touch it currently has.
My ChatGPT env only says "low", "medium", "high".
Is this a "pro" thing? I have totally no idea what I'm talking to, so actually I'm thinking of stopping my plan. Gemini and Claude are much more clear about it.
Anyway, I like the speed at which Gemini responds so indeed for simple things it is preferable.
I’ve been using Work for all my queries, since it seems to just be the same interface as Chat but with more features. I don’t understand why they’re two separate things.
Their AI leadership team has taken some hits recently too, in the form of departures. I believe when they get their bearings they will be competitive again. 3.8 Flash has been a great model for me.
Meanwhile Claude and Astra like to couch all their agreements with caveats and provisos.
Sometimes that's what being smart sounds like.
Someone confident but incorrect, can often sound more convincing than someone with actual expertise. The expert must add caveats/hedge, because those are the facts on the ground, whereas the person reciting google can be entirely confident.
Of course the people judging aren't experts, so they side with confidence and simplicity. Heck, just writing shorter replies on Reddit is rewarded. Nobody reads the articles, let alone a paragraph-long reply.
That all being said though, there are limits. Sometimes LLMs on high-thinking go off on full tangents based on little, and don't have the self-awareness to bring it back.
So I won't be addressing this, for those reasons and others.
To an expert communicating with a layperson is a form of compression. You must turn some very complex idea into one that you suppose the other person can grasp given their limited frame of reference. It's always lossy, and you have to guess how much you can remove without sounding patronizing or being inaccurate. It's tough, and the more you know the tougher it gets.
Ever done that "explain what happens when I visit Google in my web browser" interview question?
A sales guy will answer in a sentence. An engineer might be able to talk about it for several days and still not be sure they didn't miss anything important. That much knowledge can actually be detrimental to communication.
When you're a ChatGPT Projects or Claude Projects user, those caveats and provisos are your worst enemy because they'll change caveats into hard rules (either for the session or committed to memories) and you end up in absolute hell having to make it investigate to figure out why it can no longer produce anything but read-only pre-check code that never actually does anything but keeps performing stupid safety checks.
> This mirrors how Apple has always segmented Pro vs. non-Pro iPhones: base models got LTPS panels while Pro models got LTPO, and only with the mainline iPhone 17/17 Plus did that gap close the standard versions previously lacked the smoother 120Hz ProMotion technology and the always-on display feature, unlike the Pro models — the 17e is the one model line still using the older, cheaper panel.
(emphasis mine)
I mean, I can guess what it is trying to say, but who RL'd this nonsense?
Only worked in a 1:1 in a quiet place. Still, can't complain for free.
Tell Gemini Live which speaker you want to engage with, and it is relatively intelligent about it.
well then its not model problem
It reminds me of a pedantic grad student.
And like, it does this despite it speaking in extremely dense math, which both makes it sound correct and requires a lot more effort to prove when it is wrong... yet, it isn't actually correct more often, and so that time sink just isn't worth the benefit. I then think many people--including people who can speak math (as can I)--just stop bothering to check everything, as if you come across a human who speaks like this it probably does correlate with slow and careful thought that helps prevent errors.
Instead, Claude has the mistake rate of a somewhat accelerated beginner impossibly combined with the language of an expert professor; and we as humans just aren't good at that combination: it becomes very dangerous and makes it take longer to spot its egregious mistakes and trained-in biases. If you have to use Claude, I thereby claim you really need to have a team of not-Claudes to help insulate you from this, and Gemini (while being a bit senile) is a lot more collaborative and approaches problems in ways that makes it harder to get tricked.
(To translate this into more of an engineering analogy: Claude always feels to me like the engineer who put more effort into learning how to program in functional languages than into how to actually develop working code, and then confidently presents you answers in Haskell or Lisp that never quite work. To find their errors is then very costly. In contrast, Gemini feels more like a Java or Go developer who knows they are a cog... that's helpful! <- Which maybe just goes to show that AI has finally turned me into a manager, omg.)
Lately it became load-bearingly-reality-difficult to not only read, but to comprehend the Claude output
On my TODO is try and run all of the analysis pipeline in dense "machine speak" to save on tokens and just let Gemini sort it out at the end.
I've set my documentation sub agent to Gemini and my code agent to Luna
I am not so advanced enough as a human being.
Excited to try this out! Shame on Google for not releasing Gemini 3.8 for Google AI Plus users yet, though.
I'm not very happy with live conversations with Claude, so this seems like it might be a good option.
The default for the Backend is a Responses Delegation that just sends it to another OpenAI hosted model. But you can set the API up to use Client Delegation instead, and have your client/harness receive the delegation request. At that point, your client can delegate it to a Claude instead, and have Claude do the deeper thinking. Unfortunately you're not actually talking to a Claude, even if it identifies as one, but you can at least talk to something informed by Claude.
As for how to build that - I just had Fable build/vibe the whole thing for me. I used Go for the language, and SDL 3 to handle the audio & microphone side, even though it's just a command line app for now. I've previously been working on SDL3 bindings to Golang, and a personal AI harness with tool-calling & local MCP stdio support, so it's possible my Fable re-used some of that code. But Fable can whip something together, and the most basic tools (read,write,edit,datetime) are all really easy to add to a harness as built-in tools that you can let it call.
The main downside is that Anthropic probably wants you to run it as API, so costs can blow out. And depending how you setup the harness, either every backend call is a fresh new Claude window (to which you're sending a LOT of context), of you need a way to maintain a persistent Claude conversation that you queue requests to, but at least then the context is all in one window & you can benefit from caching.
But if you get it running, and play with it a bit, you'll realize this is obviously the future interface. Typing into a terminal window feels archaic. Even if we're not quite there yet, we're tantalizingly close... and anyone who has a 3D-printer & a robot arm connected via MCP/tool calls, well, this gives them Iron Man's Jarvis right now.
Every non tech company I know is using Gemini or copilot, because the same companies already were on Google or Microsoft suite and got those as extensions.
NotebookLM is way more popular in the real world than anthropic work or crap like that.
In business world contracts, data retention and procurements are more important than made up benchmarks only nerds care for.