1Reconstruction Results
1.1Short-form Reconstruction
Original audio and reconstructions from different codecs and tokenizers at comparable bitrates.
| Sample | Original | TASTE-S (ours) | Text-only | TASTE | TaDiCodec | Encodec1500 bps | DM-Codec1000 bps | Mimi1000 bps | SpeechTokenizer500 bps | BigCodec1040 bps | WavTokenizer480 bps |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Sample 1 | |||||||||||
| Sample 2 | |||||||||||
| Sample 3 |
1.2Long-form Reconstruction
Long-form speech reconstruction with TASTE-S and TASTE.
| Sample | Original | TASTE-S (ours)w/ built-in ASR | TASTE-S (ours)w/ external ASR | TASTEw/ external ASR |
|---|---|---|---|---|
| Sample 1 |
2Spoken Language Modeling
2.1Spoken Dialogue
A multi-turn spoken conversation. Each turn shows the text and the corresponding speech.
What’s the weather like today?
The weather app on my phone says it's going to be sunny with a high of 70 degrees!
Can you suggest some outdoor activities?
Sure thing! If it‘s sunny, you could go for a hike or have a picnic. If there are clouds, maybe a indoor game night or a movie marathon might be more fun.
2.2Emotional TTS
Speech synthesized with different target emotions and speaking styles.
| Surprise | Sad | Happy | Angry | Whisper |
|---|---|---|---|---|
2.3Lyrics-to-Rap Synthesis
The same lyrics read out plainly and performed in rap style, plus another rap example. Lyrics are shown under each clip.
| Reading-out | Rap-style | Rap-style (another example) |
|---|---|---|
|
The clock keeps ticking, but I never lose my pace. |
The clock keeps ticking, but I never lose my pace. |
I keep a steady flame alive when cold winds blow. |