README: State-of-the-art multilingual text-to-speech (TTS) system, redefining the boundaries of voice generation.
README: On Seed-TTS Eval, S2 achieves the lowest WER among all evaluated models including closed-source systems
README: S2 Pro supports over 80 languages without requiring phonemes or language-specific preprocessing
README: Fish Audio S2 supports accurate voice cloning using short reference samples (typically 10-30 seconds). The model captures timbre, speaking style, and emotional tendencies, produci…