2 Comments
User's avatar
Waeckerlin Federowicz's avatar

Useful roundup. For the Voice & Audio section, one thing I now look for in TTS tools is phrase-level control over emotion and pauses, not only cloning quality or latency.

FlowSpeech (https://flowspeech.io/) is a small example in that direction: context-aware text to speech with emotion/pause control and 30+ voices. It would be interesting to compare that control layer against FishSpeech-style open TTS in a future ThursdAI segment.