The Frontier
Xiaomi open-sources CocktailASR-1 to pick one voice out of a crowd
Xiaomi just gave away one of the harder pieces of the voice stack. On September 11 the company released and open-sourced Xiaomi-CocktailASR-1, an end-to-end speech recognition model built to solve the cocktail party problem — transcribing one specific person's speech from an audio mix