From Weights & Biases, we watched OpenAI announce o1 full and o1 Pro, covered 2 decentralized LLMs, open source Video and Audio models that beat proprietary companies and the GA launch of Weave
Useful roundup. For the Voice & Audio section, one thing I now look for in TTS tools is phrase-level control over emotion and pauses, not only cloning quality or latency.
FlowSpeech (https://flowspeech.io/) is a small example in that direction: context-aware text to speech with emotion/pause control and 30+ voices. It would be interesting to compare that control layer against FishSpeech-style open TTS in a future ThursdAI segment.
Useful roundup. For the Voice & Audio section, one thing I now look for in TTS tools is phrase-level control over emotion and pauses, not only cloning quality or latency.
FlowSpeech (https://flowspeech.io/) is a small example in that direction: context-aware text to speech with emotion/pause control and 30+ voices. It would be interesting to compare that control layer against FishSpeech-style open TTS in a future ThursdAI segment.
Thanks for the update