0.00 sfirst sound
0.0×real time
Or try one of mine
Soprano reads your words out loud. A small language model writes the sound as a stream of audio tokens, and a second network turns them into a voice, a fraction of a second at a time, so it starts talking before it has finished the sentence.
The words light up roughly in time with the voice. That timing is my estimate from the length of each word, not something the model reports.
Runs on your device. Your text is never uploaded. English only, one voice.