Zero-shot multilingual voice cloning with cross-lingual synthesis, disentangled emotion control, pronunciation guidance, and speaking-speed control.