Having a voice under 1Mo is crazy, even if it sounds robotic.
The original is the best: https://youtu.be/qxWwEPeUuAg
I listened to some of the voices. The male voices are believable while all the female voices sound the same and artificial. For some reason, it also reminds me of the voices in Toy Story movies.
Bias in the training data?
Um, what?
Imagine someone showing you that they've trained their dog to hold a paintbrush and paint. There would be no contradiction between "this is incredible" and "these paintings suck".
They are creating gemini SOTA (not really any more), flash versions, text-to-speech, video (omni), etc.
I can see they want to create an ecosystem, but I see no focus in any one area.
On top of that I would imagine the research side of Google might disk over things from one modality being useful to another - something like an audio optimising or memory optimising for text to speech could maybe also be useful for translation or world models. Etc etc
Basically spread many AI's everywhere, get people develop/user their ecosystem, and lock them in eventually.
But I find their models' intelligence lacking still.
I don't ever use their built-in gemini features (i have paid gmail) because they don't work well.
e.g. I ask gemini to format my Google docs per Google's material design spec with spacing, etc. It does a real bad job. Many times it does it line by line, and when I finaly get it do it for the whole doc, it does it sloppy, and extremely slow (takes 5 minutes for 10 page doc)
At best, good, but not great.
Price per hour:
- 3.8 Flash TTS, standard: $0.81
- 3.8 Flash TTS, batch: $0.41
- 3.8 Flash‑Lite TTS, standard: $0.54
- 3.8 Flash‑Lite TTS, batch: $0.27
Official pricing can be seen here: https://ai.google.dev/gemini-api/docs/pricing#gemini-3.8-fla...
$0.50 per hour pricing could last a long time with back and forth conversation use.