Is there prior open-source work done in the field of 'Audio analysis' to detect human-voice (say in spite of some background noise), determine speaker's gender, possibly determine no. of speakers, age of speaker(s), and the emotion of speakers?

My hunch is that the speech recognition software like CMU Sphinx could be a good place to start, but if there's something better, it'd be great.

Edit
Report