Loading…
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
Our verdict: worth borrowing ideas from
Loses plain transcription to faster-whisper but owns the one thing it lacks: word-level timestamps + speaker diarization. GRAFT that alignment pipeline if per-speaker output is ever needed.
Filed under speech to text in our directory.