5 ms·
well! (i've been waiting for someone to ask ha) when i first knew i wanted to align audio to text, i reached for aeneas, which was built for this exact problem
by samuelcole 2mo ago
well! (i've been waiting for someone to ask ha)
when i first knew i wanted to align audio to text, i reached for aeneas, which was built for this exact problem/product (https://www.readbeyond.it/aeneas/ https://www.readbeyond.it/aeneas/)
however, it drifted by hundreds of seconds deep into chapters and emits no confidence signal, so i couldn't eval it cleanly
so i switched to a neural acoustic model trained with ctc (meta's massively multilingual speech aligner, it ships inside torchaudio): it outputs a probability for each character at every frame of audio, and viterbi decoding finds the most probable path that spells exactly our text, with a per-word probability
that probability is the gate: a book ships read-along only if its median confidence and coverage clear a bar. moby dick's didn't, so it's honestly text-only
- p-e-w 2mo agoCould the problem be Moby Dick’s infamously large vocabulary, which contains words used essentially nowhere else?
- samuelcole 2mo agoi actually had an out-of-memory bug, which i've since fixed. chewing through it now to see if we can get it
- samuelcole 2mo agogot it working! https://tale.fyi/moby-dick https://tale.fyi/moby-dick it got confused by some of the more uhh complex passages ;-)