Top
Best
New

Posted by swolpers 7 hours ago

Gemini 3.8 text-to-speech(blog.google)
213 points | 108 commentspage 2
maelito 5 hours ago|
Related, for embedding small models, this lib is incredible.

Having a voice under 1Mo is crazy, even if it sounds robotic.

https://tts.ampixa.com/sanoTTS/

112233 6 hours ago||
"Super tinny monotone robotic voice" does not sound neither tinny nor monotone. Compared to what TTS from 90s sounded like. Or even how actors impersonated robots in movies. Has the model been eating too much hype DJs?
burkaman 5 hours ago|
None of these examples are really what the prompt asked for. It's just like image models, once you get over how unbelievable it is that a computer produced this you realize the result isn't actually what you want.
xnx 6 hours ago||
Would be great if this would power the Google Books app feature. The voice system there is pretty out of date.
laweijfmvo 5 hours ago||
the ratio of new voice models i see on hackernews to the number actually deployed in any product i use is approximately infinity.
mamudo 6 hours ago||
Yes, I am quite disappointed by seeing all this cool AI stuff and yet the same Play Books. Come on, it is the best place to apply AI, in my opinion.
hatingisok 3 hours ago||
My ROFLcopter goes: SOISOISOISOISOISOISOISOISOISOISOISOISOISOISOISOI
tantalor 2 hours ago||
Would pay any amount of money for a zombo.com voice.

The original is the best: https://youtu.be/qxWwEPeUuAg

burkaman 5 hours ago||
Seems like voice actors are safe for now. This is technologically incredible, but the results are really not very good, and usually not particularly close to the prompt. In basically all of these examples some core part of the prompt is completely ignored.
kanbankaren 3 hours ago||
Probably.

I listened to some of the voices. The male voices are believable while all the female voices sound the same and artificial. For some reason, it also reminds me of the voices in Toy Story movies.

Bias in the training data?

avazhi 5 hours ago||
> Seems like voice actors are safe for now. This is technologically incredible, but the results are really not very good,

Um, what?

burkaman 5 hours ago||
Technologically incredible as in "I cannot believe it's possible for a computer to do this" and not very good as in "these examples are not what was prompted and I can't think of a use case where these would be acceptable".

Imagine someone showing you that they've trained their dog to hold a paintbrush and paint. There would be no contradiction between "this is incredible" and "these paintings suck".

yipinwong 3 hours ago||
Google is spreading too thin, as gemini isn't really that intelligent.

They are creating gemini SOTA (not really any more), flash versions, text-to-speech, video (omni), etc.

I can see they want to create an ecosystem, but I see no focus in any one area.

Melatonic 2 hours ago||
Honestly I think their approach is the long term most useful one. Being SOTA is probably very costly and difficult. Optimising all the smaller stuff that's real world useful for people seems like it would have better long term use.

On top of that I would imagine the research side of Google might disk over things from one modality being useful to another - something like an audio optimising or memory optimising for text to speech could maybe also be useful for translation or world models. Etc etc

yipinwong 56 minutes ago||
I can see where you are coming from. Vendor lock-in.

Basically spread many AI's everywhere, get people develop/user their ecosystem, and lock them in eventually.

But I find their models' intelligence lacking still.

I don't ever use their built-in gemini features (i have paid gmail) because they don't work well.

e.g. I ask gemini to format my Google docs per Google's material design spec with spacing, etc. It does a real bad job. Many times it does it line by line, and when I finaly get it do it for the whole doc, it does it sloppy, and extremely slow (takes 5 minutes for 10 page doc)

mewse-hn 1 hour ago||
The gemma 4 models are pretty great
yipinwong 59 minutes ago||
not as good as other chinese openweight models.

At best, good, but not great.

sgc 6 hours ago||
Sorry if this is in that article, but I am on my phone and can't see it. How much would this cost to batch generate an audiobook? Right now I just listen to things in the 11 labs app which is free, but I would rather just generate audio files.
thevinter 6 hours ago|
Roughly 5-10$ for 10h, assuming you few-shot it.

Price per hour:

- 3.8 Flash TTS, standard: $0.81

- 3.8 Flash TTS, batch: $0.41

- 3.8 Flash‑Lite TTS, standard: $0.54

- 3.8 Flash‑Lite TTS, batch: $0.27

sgc 5 hours ago|||
That is the biggest difference here for me. I have not looked at every solution, but many. They are either much more expensive or garbage quality. For example my next target audiobook is a monster 250k words, 1.5m characters, so about $15 here or $75 using elevenlabs (0.05 per 1k characters). For me that is the difference between I will or I won't use it.
eis 4 hours ago|||
Until the end of the year, then double that. And that's only the audi output though the text input shouldn't cost much in comparison.

Official pricing can be seen here: https://ai.google.dev/gemini-api/docs/pricing#gemini-3.8-fla...

nitroedge 5 hours ago||
Couldn't see this in the article, does it support API calls for one-shot conversation type responses like ElevenLabs offers?

$0.50 per hour pricing could last a long time with back and forth conversation use.

AyanamiKaine 3 hours ago|
Still, all voices sound like they are missing something only real human speech can sound like. But many people will not notice the difference between AI and normal voices.
pixl97 3 hours ago|
So you're saying they are up to support tech afterb5 hours on the same call... the soul has left their body.
More comments...