Top
Best
New

Posted by k9294 12 hours ago

Gemini-3.5-Transcribe(blog.google)
210 points | 59 commentspage 2
dbbk 7 hours ago|
Still more expensive and worse performing than ElevenLabs Scribe, unfortunately. Not sure who's the target audience for this.
mariano54 3 hours ago||
Just added this to my benchmark site: https://multilingualsttbench.com/

It doesn't reach the frontier in either latency or accuracy for ai multilingual conversations.

adamgoodapp 3 hours ago||
Thanks for this, really helpful.

I would also like to see benchmark for translation. I'm looking for live translated subtitles so my Japanese wife can enjoy any show with out waiting months for official VOD streams to release them.

Kokouane 3 hours ago||
I'm confused, doesn't your leaderboard clearly show it is the most accurate model? It's number one in the leaderboard. Am I missing something?
Kokouane 2 hours ago||
Figured it out. 3.5 Flash and 3.5 Transcribe are different models
jeffbee 9 hours ago||
I am not sure if "Word Error Rate" captures what has always been wrong with transcription. My biggest complaint is that it inserts sentence breaks in random places, then fails to evaluate the result, even though it is obviously wrong. Then I have to go fix it which can be harder than having just typed it myself, due to the difficulty of positioning the Android cursor, the fact that it automatically capitalizes if you delete a capital letter, etc. And much of the time I fail to notice the errors until later.
verdverm 8 hours ago|
have another model do a pass to clean it up, saw a demo of local STT where someone did this, can fix a lot of things, especially with gotchas for the STT model in a clean-transcript.md
coder543 8 hours ago|||
I haven't tried it, but this looked promising for that exact task: https://huggingface.co/superwhisper/s1-mini
jeffbee 8 hours ago|||
I think the model can even evaluate itself. If it looks afterward at an output like "do you. Want to get lunch?" in the absence of affirmative evidence that the user wanted it that way, it should be able to see that it goofed.
iAMkenough 8 hours ago||
Hopefully YouTube automatic captions improve with this
mythz 3 hours ago||
Annoying that they don't include pricing info in new models, here it is [1]:

Gemini 3.5 Transcribe Live (Per 1M tokens in USD):

    Input: $3.50 or $0.005/min* (audio)
    Output: 21.00 or $0.004/min* (text)
Gemini 3.5 Transcribe:

    Input: $2.00 or $0.003/min* (audio)
    Output: $12.00 or $0.002/min* (text)
[1] https://ai.google.dev/gemini-api/docs/pricing#gemini-3.5-tra...
ElijahLynn 7 hours ago||
Very impressive, including the ability to hit fn in any text field and say "generate an image ...".
Freedom2 9 hours ago||
I'd love to know how this handles proper subtitle formatting. I'm in the process of learning many languages, and being able to cross check my own understanding with film and video would be fantastic.
hypfer 8 hours ago||
Where does the compute happen?

I suppose it's a cloud thing?

k9294 8 hours ago|
Yep.
HappyPanacea 9 hours ago||
Does somebody knows what top locales list is sampled from? their own usage data? Also when they will use their AI to give better directions in Waze?
hkjhkjhj 5 hours ago|
[flagged]