Short and long sentences, with and without numbers, abbreviations and English terms. All mixed up, so you can get an honest picture. All samples are in German – the sentence spoken is shown next to each row.
All examples are unedited, only the volume was normalized. No picking of "best takes", no fine-tuning. I don't need to impress anyone – I want to show you what you actually get.
The generation times were deliberately measured on a MacBook M1 without a high-end graphics card, so you get a realistic impression of how it feels on a regular computer. The time shown next to each play button is the pure generation time, not the length of the audio file.
The test sentences were created by an AI, not hand-picked by me – so that I don't unconsciously choose sentences I already know will sound good.
The same sentences, three engines. This way you can hear the difference in sound and generation time right away.
The same sentences, three engines — Piper, Kokoro and CosyVoice, all with a Hessian dialect.
One and the same sentence, six emotions — only available with Piper.
With CosyVoice, you'll notice that numbers are currently pronounced in English instead of German. If you use CosyVoice, you should convert numbers, dates and times into spelled-out German words yourself beforehand (text normalization) — otherwise the result sounds wrong in exactly those places.
That's exactly why the uncomfortable examples are here too, not just the ones that sound best. A voice that seems perfect everywhere wouldn't be honest.
If you want to normalize numbers, dates or times yourself: the Python library num2words reliably converts digits into spelled-out German words and fits nicely into your own preprocessing. German abbreviations like "z. B." (e.g.) or "bzw." (or rather) currently still have to be replaced yourself using a dictionary. There's no universal automatic solution for this. Alternatively, you can also use german transliterate.