KnowledgeHub
Questions
Tags
Users
Search
Alex Rivera
|
Logout
Edit Question
Title
Body
I have a set of audio files that are uploaded by users, and there is no knowing what they contain. I would like to take an arbitrary audio file, and extract each of the instances where someone is speaking into separate audio files. I don't want to detect the actual words, just the "started speaking", "stopped speaking" points and generate new files at these points. (I'm targeting a Linux environment, and developing on a Mac) I've found Sox , which looks promising, and it has a 'vad' mode (Voice Activity Detection). However this appears to find the first instance of speech and strips audio until that point, so it's close, but not quite right. I've also looked at Python's 'wave' library, but then I'd need to write my own implementation of Sox's 'vad'. Are there any command line tools that would do what I want off the shelf? If not, any good Python or Ruby approaches?
Tags (comma-separated)
Save Edits
Cancel