Demo

sub11 Link to heading

Sample Original Proposed (rtMRI path) w/ X-vector (audio path)
vcv1_r1
bvt_r1
grandfather1_r1
northwind1_r1
picture3

sub20 Link to heading

Sample Original Proposed (rtMRI path) w/ X-vector (audio path)
vcv1_r1
bvt_r1
grandfather1_r1
northwind1_r1
picture3

sub22 Link to heading

Sample Original Proposed (rtMRI path) w/ X-vector (audio path)
vcv1_r1
bvt_r1
grandfather1_r1
northwind1_r1
picture3

sub55 Link to heading

Sample Original Proposed (rtMRI path) w/ X-vector (audio path)
vcv1_r1
bvt_r1
grandfather1_r1
northwind1_r1
picture3

Speaker Split Link to heading

Split Speakers
train sub01, sub06, sub07, sub08, sub10, sub12, sub13, sub14, sub15, sub16, sub17, sub21, sub23, sub24, sub25, sub30, sub31, sub33, sub34, sub35, sub36, sub37, sub38, sub39, sub40, sub41, sub42, sub43, sub44, sub46, sub47, sub48, sub51, sub52, sub57, sub59, sub61, sub63, sub67, sub69, sub71, sub72, sub74
valid sub05, sub09, sub19, sub70
test sub11, sub20, sub22, sub55

Dataset Link to heading

The speech samples are derived from the USC 75-Speaker Speech MRI Database (CC BY 4.0); the original audio was denoised, and synthesized samples are included.

Lim, Y., Toutios, A., Bliesener, Y. et al. A multispeaker dataset of raw and reconstructed speech production real-time MRI video and 3D volumetric images. Sci Data 8, 187 (2021).