Demo · Expressive Speech-to-Speech Translation

Holistic Parallel Supervision for Expressive Speech-to-Speech Translation

Hao Zhang, Ismail Rasim Ulgen, Bismarck Bamfo Odoom, Rui Liu, Zihan Zhang, Berrak Sisman, Philipp Koehn
Center for Language and Speech Processing, Johns Hopkins University

Abstract: Expressiveness is central to speech-to-speech translation (S2ST), yet existing datasets rarely align source and target speech jointly in content, prosody, timing, and speaker identity. Consequently, current systems often rely on indirect conditioning or transfer mechanisms. To obtain this supervision, we introduce a training-data construction strategy that combines professional dubbing, which provides human-performed correspondence in content, prosody, and timing, with multilingual voice conversion to reduce speaker mismatch. We evaluate this supervision using existing expressive S2ST systems and DirectS2ST, an end-to-end method developed in this work that predicts the complete target-speech representation directly from source speech. Experiments on English–German and English–Spanish show that holistic parallel supervision improves expressive realization across all evaluated systems. DirectS2ST exhibits the largest gains in content, prosody, timing, speaker similarity, and human-rated expressiveness. To our knowledge, this work provides the first systematic investigation of holistic parallel supervision for expressive S2ST.
Training Data Construction

VC-Dub: Dubbing + Voice Conversion

Professional dubbing supplies human-performed correspondence in content, prosody, and timing. Multilingual voice conversion (Seed-VC) moves the English dub to the original speaker's voice, so the resulting pairs are parallel in content, prosody, timing, and speaker identity. The original German/Spanish recording is retained as the target after audio cleaning, without TTS or VC resynthesis.

(a) VC-Dub construction with multilingual voice conversion; (b) parallel supervision for direct expressive S2ST
Figure 1 (a, b). (a) VC-Dub construction: the English dub, which imitates the original performance, is voice-converted to the original speaker using the original utterance as the speaker reference. (b) Parallel supervision: the voice-converted English speech becomes the source and the original German/Spanish recording the target, aligned in semantics, prosody, timing, and speaker identity. [PDF]
Data Quality Control

VC-DUB Cleaning and Filtering

We apply denoising, vocal extraction, language identification, speaker filtering, and speech-quality filtering before constructing the final VC-DUB training splits.

Panel A: Filtering Statistics

Stage En–Es En–De
Pairs Total Hrs. Retention Pairs Total Hrs. Retention
Raw dubbing from Speech Vecalign352,947748.29–369,139782.86–
After ClearVoice + Demucs352,947748.29100.0%369,137782.86100.0%
After MMS-LID213,469536.6360.5%298,929693.4081.0%
After Sortformer single-speaker gate164,746349.4346.7%237,743468.2864.4%
After DNSMOSPro quality selection90,000194.5125.5%147,639295.8040.0%
Training split used in Table 179,200171.13–131,399263.22–

Panel B: Filtering Settings

DenoisingClearVoice-Studio with MossFormer2_SE_48K under default inference.
Vocal extractionDemucs htdemucs; we retain the vocal stem and require successful preprocessing on both source and target sides.
Language identificationMMS-LID facebook/mms-lid-126; we retain pairs whose top-1 predicted labels match the expected languages. Expected labels are eng/spa for En–Es and eng/deu for En–De. No confidence threshold is applied.
Speaker filteringSortformer nvidia/diar_sortformer_4spk-v1; audio is converted to 16 kHz mono. Retain pairs only when exactly one active speaker is detected in each utterance.
Quality filteringPairs are selected using source–target DNSMOSPro scores: both utterances are scored, and pairs are ranked by a combined source–target score and retained until the corpus is comparable in scale to CVSS-T.
Observed DNSMOSPro boundaryThe observed retained/dropped boundaries are approximately 3.57 for En–Es and 3.60 for En–De; these are not preset thresholds. See the repository for the available settings and reproducibility limitations.
Split constructionThe retained clean pool is subsequently split for controlled S2ST training. The final row in Panel A reports the training split used in Table 1.

Note. Total hours denote the combined source and target durations, computed using the same VAD-span definition as Table 1. Retention is computed relative to the original raw dubbing pool. The final training-split row is included for comparison with Table 1 and does not represent an additional cleaning stage. The DNSMOSPro boundaries are observed after scale-matched selection, not preset thresholds.

Training Data Analysis

Pair-Level Correspondence

Raw-Dub already gives stronger content, prosody, and timing correspondence than CVSS-T; its main weakness is speaker mismatch. Voice conversion substantially raises Vsim and further increases A.PCP. Shuffling the VC-Dub pairings (negative control) substantially reduces the correspondence scores.

Dir.DatasetBLASER↑A.PCP↑Rate↑Pause↑Vsim↑DNSMOSPro↑ (src | tgt)
En→DeCVSS-T3.4052.5050.3650.5240.2083.160 | 2.797
Raw-Dub3.6342.7590.4160.6820.0603.900 | 3.684
VC-Dub3.6113.0000.4350.6980.4064.104 | 3.684
w/ Shuffled2.2852.218-0.0090.3940.0254.104 | 3.684
En→EsCVSS-T3.4002.5370.3720.5500.1943.173 | 2.780
Raw-Dub3.6322.9200.4170.6760.0463.913 | 3.714
VC-Dub3.6063.1830.4360.6860.3624.091 | 3.714
w/ Shuffled2.1992.345-0.0030.3610.0274.091 | 3.714

Table 2. Pair-level correspondence and speech quality of the training data for CVSS-T, Raw-Dub, and VC-Dub. Shuffled VC-Dub pairs provide a negative control. DNSMOSPro is reported as source (En) | target (De/Es).

Direct Modeling

DirectS2ST

DirectS2ST is trained to predict all 16 target Mimi codec streams from source-derived conditions, with every stream supervised by the paired human recording. Its acoustic and speaker prompts are extracted from the source during both training and inference.

DirectS2ST architecture
Figure 1 (c). DirectS2ST. A frozen w2v-BERT 2.0 encoder provides source memory through cross-attention. The AR temporal decoder generates source text, target text, and the first target codec stream c(0) with a shared output head; its <SEP> embedding is replaced by a mean–std pooled Mimi prompt extracted from the source speech. A NAR depth decoder predicts c(1)–c(15), conditioned on a source codec prompt and a projected source-speaker embedding from a frozen CAMPPlus encoder, and is trained with an additional speaker-identity loss. At inference, beam search (beam 5) decodes the text–c(0) sequence, then the depth decoder greedily fills c(1)–c(15). [PDF]
Evaluation on mExpresso

Results

504 test utterances per direction, evenly distributed across seven expressive conditions. Content: ASR-BLEU (Whisper-large-v3) and BLASER 2.0-QE; prosody and timing: AutoPCP (A.PCP), speaking-rate correlation (Rate) and pause similarity (Pause); speaker similarity: Vsim; naturalness: NISQA-TTS (NAT).

Dir.ModelTrainASR-BLEU↑BLASER↑A.PCP↑Rate↑Pause↑Vsim↑NAT↑
En→DeTransVIPCVSS-T12.253.3972.6090.4700.4890.1263.257
TransVIPVC-Dub13.843.3992.743†0.600†0.549†0.169†3.472†
DirectS2STCVSS-T6.013.2342.1620.3340.6870.1143.339
DirectS2STVC-Dub10.52†3.373†2.905†,‡0.589†0.752†,‡0.217†,‡3.450†
En→EsTransVIPCVSS-T15.533.3542.6640.6010.4940.1383.092
TransVIPVC-Dub18.96†3.433†2.803†0.671†0.562†0.158†3.407†
Dub-S2STCVSS-T11.613.2392.9330.5810.5960.2183.658
Dub-S2STCVSS-TS11.893.2262.9380.6080.5720.2143.631
Dub-S2STVC-Dub15.213.321†3.012†0.5950.6120.238†3.759†
DirectS2STCVSS-T10.003.1532.2380.4330.6520.1053.361
DirectS2STVC-Dub14.22†3.259†3.091†,‡0.669†0.733†,‡0.205†3.466†

Table 3. Expressive S2ST results on mExpresso. † denotes significance between training sets within the same model, and ‡ against the strongest baseline in the same language direction (two-sided paired randomization tests, 10,000 permutations, Holm–Bonferroni correction, padj < 0.05). Bold marks the best score per direction. Dub-S2ST is evaluated only on En→Es: the mHuBERT checkpoint used in this baseline covers English, Spanish, and French. CVSS-TS is its speed-adapted CVSS-T variant. Dub-S2ST denotes our adapted baseline: we replace its original S2U module with the model of Popuri et al. (2022) and retain its CosyVoice-initialized U2S module. Both components are trained on the corresponding supervision dataset.

Dir.TrainASR-BLEU↑BLASER↑A.PCP↑Rate↑Pause↑Vsim↑NAT↑
En→DeVC-Dub10.523.3732.905†0.589†0.7520.217†3.450
Raw-Dub9.373.3392.7120.4550.7410.0653.430
En→EsVC-Dub14.223.2593.091†0.669†0.7330.205†3.466
Raw-Dub14.533.2662.8240.5550.7300.0583.470

Table 4. Contribution of voice conversion: DirectS2ST trained on the same aligned dubbing pairs with (VC-Dub) and without (Raw-Dub) VC. VC mainly improves speaker similarity, with additional gains in A.PCP and Rate.

ModelTraining DataMOS↑95% CI
TransVIPCVSS-T1.82[1.62, 2.01]
TransVIPVC-Dub2.24†[2.03, 2.44]
Dub-S2STCVSS-T2.88[2.63, 3.12]
Dub-S2STVC-Dub3.59†[3.34, 3.84]
DirectS2STCVSS-T2.13[1.93, 2.33]
DirectS2STVC-Dub3.42†[3.20, 3.64]

Table 5. Expressiveness MOS (En→Es, 9 listeners × 10 utterances) with 95% confidence intervals. ICC(A,1) = 0.445, ICC(A,9) = 0.878. † padj < 0.05 (crossed linear mixed-effects model, Holm–Bonferroni).

Audio Demo

Speech Samples

Each system is trained on CVSS-T or on our VC-Dub data. Source utterances come from the mExpresso test set; texts are shown as in the dataset transcripts, with emphasized words in bold italics.

English → Spanish

The 10 utterances used in the expressiveness listening test (Table 5), with all six system configurations.

#1Sadex01_sad_00317
Source (En) What's the emergency?
Reference (Es) ¿Cuál es la emergencia?
Source speech (En)
TransVIP · CVSS-T
TransVIP · VC-Dub
Dub-S2ST · CVSS-T
Dub-S2ST · VC-Dub
DirectS2ST · CVSS-T
DirectS2ST · VC-Dub (ours)
#2Defaultex01_default_00010
Source (En) I don't know. My mom said I was, but my dad told me to just blow her off.
Reference (Es) No sé. Mi mamá me dijo que yo lo estaba, pero mi papá me dijo que solo la ignore.
Source speech (En)
TransVIP · CVSS-T
TransVIP · VC-Dub
Dub-S2ST · CVSS-T
Dub-S2ST · VC-Dub
DirectS2ST · CVSS-T
DirectS2ST · VC-Dub (ours)
#3Confusedex01_confused_00020
Source (En) Hold the elevator, please!
Reference (Es) ¡Paren el ascensor, por favor!
Source speech (En)
TransVIP · CVSS-T
TransVIP · VC-Dub
Dub-S2ST · CVSS-T
Dub-S2ST · VC-Dub
DirectS2ST · CVSS-T
DirectS2ST · VC-Dub (ours)
#4Sadex04_sad_00364
Source (En) OK, now you have a podcast that's available on Audible.
Reference (Es) Bien, ahora tienes un pódcast que está disponible en Audible.
Source speech (En)
TransVIP · CVSS-T
TransVIP · VC-Dub
Dub-S2ST · CVSS-T
Dub-S2ST · VC-Dub
DirectS2ST · CVSS-T
DirectS2ST · VC-Dub (ours)
#5Whisperex01_whisper_00155
Source (En) Effort is all it takes.
Reference (Es) Todo lo que se necesita es hacer un esfuerzo.
Source speech (En)
TransVIP · CVSS-T
TransVIP · VC-Dub
Dub-S2ST · CVSS-T
Dub-S2ST · VC-Dub
DirectS2ST · CVSS-T
DirectS2ST · VC-Dub (ours)
#6Enunciatedex02_enunciated_00170
Source (En) Have you read this?
Reference (Es) ¿Has leído esto?
Source speech (En)
TransVIP · CVSS-T
TransVIP · VC-Dub
Dub-S2ST · CVSS-T
Dub-S2ST · VC-Dub
DirectS2ST · CVSS-T
DirectS2ST · VC-Dub (ours)
#7Happyex04_happy_00230
Source (En) The vampire slayer?
Reference (Es) ¿La cazavampiros?
Source speech (En)
TransVIP · CVSS-T
TransVIP · VC-Dub
Dub-S2ST · CVSS-T
Dub-S2ST · VC-Dub
DirectS2ST · CVSS-T
DirectS2ST · VC-Dub (ours)
#8Emphasisex01_default_emphasis_00120
Source (En) The people deserve to know the truth.
Reference (Es) La gente merece conocer la verdad.
Source speech (En)
TransVIP · CVSS-T
TransVIP · VC-Dub
Dub-S2ST · CVSS-T
Dub-S2ST · VC-Dub
DirectS2ST · CVSS-T
DirectS2ST · VC-Dub (ours)
#9Whisperex01_whisper_00272
Source (En) So, what's it about?
Reference (Es) Entonces, ¿de qué se trata?
Source speech (En)
TransVIP · CVSS-T
TransVIP · VC-Dub
Dub-S2ST · CVSS-T
Dub-S2ST · VC-Dub
DirectS2ST · CVSS-T
DirectS2ST · VC-Dub (ours)
#10Confusedex02_confused_00216
Source (En) She stayed in eleven different foster homes?
Reference (Es) ¿Ella estuvo en once hogares de acogida diferentes?
Source speech (En)
TransVIP · CVSS-T
TransVIP · VC-Dub
Dub-S2ST · CVSS-T
Dub-S2ST · VC-Dub
DirectS2ST · CVSS-T
DirectS2ST · VC-Dub (ours)

English → German

10 utterances from the En→De test set across expressive conditions. Dub-S2ST is not included: the mHuBERT checkpoint used in this baseline covers English, Spanish, and French.

#1Whisperex01_whisper_00257
Source (En) You made a mistake.
Reference (De) Du hast einen Fehler gemacht.
Source speech (En)
TransVIP · CVSS-T
TransVIP · VC-Dub
DirectS2ST · CVSS-T
DirectS2ST · VC-Dub (ours)
#2Happyex02_happy_00042
Source (En) Alright, that's it! I'm gonna be right outside those doors.
Reference (De) Alles klar, das war's! Ich werde jetzt durch diese Tür nach draußen gehen.
Source speech (En)
TransVIP · CVSS-T
TransVIP · VC-Dub
DirectS2ST · CVSS-T
DirectS2ST · VC-Dub (ours)
#3Confusedex03_confused_00269
Source (En) Does the prosecution have a recommendation?
Reference (De) Hat die Staatsanwaltschaft eine Empfehlung?
Source speech (En)
TransVIP · CVSS-T
TransVIP · VC-Dub
DirectS2ST · CVSS-T
DirectS2ST · VC-Dub (ours)
#4Defaultex01_default_00028
Source (En) He wrote Gulliver's Travel.
Reference (De) Er hat Gullivers Reisen geschrieben.
Source speech (En)
TransVIP · CVSS-T
TransVIP · VC-Dub
DirectS2ST · CVSS-T
DirectS2ST · VC-Dub (ours)
#5Sadex01_sad_00010
Source (En) I don't know. My mom said I was, but my dad told me to just blow her off.
Reference (De) Ich weiß nicht. Meine Mutter hat es gesagt, aber mein Vater hat gesagt, ich soll nicht auf sie hören.
Source speech (En)
TransVIP · CVSS-T
TransVIP · VC-Dub
DirectS2ST · CVSS-T
DirectS2ST · VC-Dub (ours)
#6Enunciatedex01_enunciated_00205
Source (En) The people got too greedy.
Reference (De) Die Leute wurden zu gierig.
Source speech (En)
TransVIP · CVSS-T
TransVIP · VC-Dub
DirectS2ST · CVSS-T
DirectS2ST · VC-Dub (ours)
#7Emphasisex01_default_emphasis_00181
Source (En) Well, I'm with Kevin.
Reference (De) Nun ja, ich bin mit Kevin zusammen.
Source speech (En)
TransVIP · CVSS-T
TransVIP · VC-Dub
DirectS2ST · CVSS-T
DirectS2ST · VC-Dub (ours)
#8Happyex04_happy_00087
Source (En) I hoped that help!
Reference (De) Ich hoffe, das hilft!
Source speech (En)
TransVIP · CVSS-T
TransVIP · VC-Dub
DirectS2ST · CVSS-T
DirectS2ST · VC-Dub (ours)
#9Sadex04_sad_00189
Source (En) OK, Barbara Long. And which Evans, Dorothy Evens or Jessica Evans?
Reference (De) Okay, Barbara Long. Und welche Evans, Dorothy Evans oder Jessica Evans?
Source speech (En)
TransVIP · CVSS-T
TransVIP · VC-Dub
DirectS2ST · CVSS-T
DirectS2ST · VC-Dub (ours)
#10Confusedex03_confused_00149
Source (En) Where do your loyalties lie?
Reference (De) Wem bist du loyal?
Source speech (En)
TransVIP · CVSS-T
TransVIP · VC-Dub
DirectS2ST · CVSS-T
DirectS2ST · VC-Dub (ours)
References.
TransVIP: Le et al., TransVIP: Speech to Speech Translation System with Voice and Isochrony Preservation, NeurIPS 2024.
Dub-S2ST: Choi, Kim, and Chung, Dub-S2ST: Textless Speech-to-Speech Translation for Seamless Dubbing, Findings of EMNLP 2025.
S2U model: Popuri et al., Enhanced Direct Speech-to-Speech Translation Using Self-supervised Pre-training and Data Augmentation, Interspeech 2022.
CVSS: Jia et al., CVSS Corpus and Massively Multilingual Speech-to-Speech Translation, LREC 2022.
Speech Vecalign: Meng and Koehn, Speech Vecalign: an Embedding-based Method for Aligning Parallel Speech Documents, EMNLP 2025.
Seed-VC: Liu, Zero-shot Voice Conversion with Diffusion Transformers, arXiv:2411.09943, 2024. Code.
Expresso: Nguyen et al., EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis, Interspeech 2023.
mExpresso / Seamless: Barrault et al., Seamless: Multilingual Expressive and Streaming Speech Translation, arXiv:2312.05187, 2023.
Disclaimer: Audio samples are provided for research demonstration purposes only. Due to licensing constraints, the dubbing dataset itself cannot be shared; the construction scripts are released in VC-DUB/.