Method and apparatus for searching multimedia data using speech recognition in mobile device

Information

  • Patent Application
  • 20070208561
  • Publication Number
    20070208561
  • Date Filed
    February 09, 2007
    19 years ago
  • Date Published
    September 06, 2007
    18 years ago
Abstract
A method of searching music using speech recognition in a mobile device, the method including: recognizing a speech signal uttered by a user as a phoneme sequence; and searching music information by performing partial symbol matching between the recognized phoneme sequence and a standard pronunciation sequence.
Description

BRIEF DESCRIPTION OF THE DRAWINGS

The above and/or other aspects and advantages of the present invention will become apparent and more readily appreciated from the following detailed description, taken in conjunction with the accompanying drawings of which:



FIG. 1 is a diagram illustrating a configuration of a music search apparatus according to an embodiment of the present invention;



FIG. 2 is a diagram illustrating an example of a music information generation unit includable the music search apparatus of FIG. 1, illustrated in the context of a music delivery system;



FIG. 3 is a diagram illustrating an example of a match of a reference pattern and a recognized symbol sequence by the matching unit of the music search apparatus of FIG. 1;



FIG. 4 is a diagram illustrating an example of a phoneme confusion matrix usable by the matching unit of the music search apparatus of FIG. 1;



FIG. 5 is a diagram illustrating an example of display of a music information search result by the display unit of the music search apparatus of FIG. 1; and



FIG. 6 is a flowchart illustrating a music search method according to an embodiment of the present invention.


Claims
  • 1. A method of searching music using speech recognition, the method comprising: recognizing as a phoneme sequence a speech signal uttered by a user; andsearching music information by performing partial symbol matching between the recognized phoneme sequence and a standard pronunciation sequence.
  • 2. The method of claim 1, wherein the searching music information comprises: performing the partial symbol matching between the recognized phoneme sequence and the standard pronunciation sequence;calculating a match score according to a result of the partial symbol matching; anddisplaying a music information search result according to the match score.
  • 3. The method of claim 1, wherein the recognizing as a phoneme sequence a speech signal uttered by a user comprises: extracting a feature of the speech signal uttered by the user; anddecoding a phoneme according to the extracted feature of the speech signal.
  • 4. The method of claim 2, wherein the match score is calculated by a phoneme confusion matrix.
  • 5. The method of claim 2, wherein, in the displaying a music information search result according to the match score, only a music information search result having the match score greater than a predetermined reference value is displayed.
  • 6. The method of claim 1, further comprising extracting a recognition target vocabulary from a predetermined music file and generating the music information with respect to the extracted recognition target vocabulary.
  • 7. The method of claim 6, further comprising: generating a pronunciation dictionary with the recognition target vocabulary; andsorting the generated pronunciation dictionary.
  • 8. A computer-readable recording medium in which a program for executing a method of searching music using speech recognition is recorded, the method comprising: recognizing as a phoneme sequence a speech signal uttered by a user; andsearching music information by performing partial symbol matching between the recognized phoneme sequence and a standard pronunciation sequence.
  • 9. A music search apparatus comprising: a music database storing a pronunciation dictionary with respect to music and music information;a phoneme decoding unit decoding a speech signal into a candidate phoneme sequence;a matching unit matching the candidate phoneme sequence with a reference phoneme pattern in the pronunciation dictionary with respect to the music information;a calculation unit calculating a match score according to a result of the matching; anda display unit displaying a music information search result according to the calculated match score.
  • 10. The apparatus of claim 9, wherein the matching unit matches the candidate phoneme sequence with the reference phoneme pattern in the pronunciation dictionary, with respect to the music information, using a phoneme confusion matrix and language boundary information.
  • 11. The apparatus of claim 9, wherein the matching unit converts a pronunciation sequence of a part of the candidate phoneme sequence exhibiting an effect of palatalization into an original pronunciation sequence in an isolated speech form and matches the converted pronunciation sequence with the reference phoneme pattern of the pronunciation dictionary.
  • 12. The apparatus of claim 9, wherein the display unit displays only music information search results having the match score greater than a predetermined reference value.
  • 13. The apparatus of claim 9, wherein the display unit arranges and displays music information search results according to a predetermined criteria when the match score of the music information search result is the same as another match score of another search.
  • 14. The apparatus of claim 9, further comprising a feature extraction unit extracting a speech feature from the speech signal before the phoneme decoding unit decodes the speech signal into the candidate phoneme sequence.
  • 15. The apparatus of claim 9, further comprising a music information generation unit extracting a recognition target vocabulary from a predetermined music file, and generating the music information with respect to the extracted recognition target vocabulary.
  • 16. A music search apparatus comprising: a feature extraction unit extracting a feature vector sequence of a speech signal of an input speech query;a phoneme decoding unit decoding the extracted feature vector sequence into at least one candidate phoneme sequences;a matching unit partially matching a candidate phoneme sequence with a reference pattern included in a stored lexicon by matching the candidate phoneme sequence with the reference pattern using a phoneme confusion matrix and linguistic constraints and, after the partial matching, matching a converted pronunciation sequence with a reference phoneme pattern of the lexicon so as to overcome an inconsistency due to a difference in pronunciation caused by palatalization; anda calculation unit calculating a match score according to the match score using a probability value of the phoneme confusion matrix and considering probabilities of insertion and deletion of the phoneme.
  • 17. The apparatus of claim 16, further comprising a music database storing music, music information, and the lexicon, the lexicon being for the music information and corresponding to a reference pronunciation pattern for comparing a speech query with a recognized phoneme sequence.
  • 18. The apparatus of claim 16, wherein the feature extraction unit extracts a feature vector sequence of a speech signal of an input speech query by reducing background noise of the speech signal of the speech query, extracting a speech interval from the speech signal, and extracting a feature vector sequence usable in speech recognition from the detected speech interval
  • 19. The apparatus of claim 16, wherein the phoneme decoding unit decodes the extracted feature vector sequence into the at least one candidate phoneme sequence using a phoneme or a tri-phoneme acoustic model and applies connectivity between contexts when using the tri-phoneme acoustic model.
  • 20. The apparatus of claim 16, wherein the phoneme decoding unit applies a phoneme-level grammar when converting the extracted feature vector sequence into the at least one candidate phoneme sequence.
  • 21. The apparatus of claim 16, wherein the matching unit obtains the converted pronunciation sequence used to overcome an inconsistency due to a difference in pronunciation caused by palatalization by converting a pronunciation sequence of a part of the candidate phoneme sequence exhibiting an effect of palatalization into an original pronunciation sequence in an isolated speech form.
  • 22. The apparatus of claim 16, wherein the conversion of the pronunciation sequence into the original pronunciation sequence enables regularization by back-tracking from a pronunciation rule.
  • 23. The apparatus of claim 16, wherein the matching a converted pronunciation sequence with a reference phoneme pattern of the lexicon is achieved by Viterbi alignment with respect to a matched phoneme segment of a candidate recognition list obtained from the partial matching.
Priority Claims (1)
Number Date Country Kind
10-2006-0020089 Mar 2006 KR national