The appended drawings illustrate exemplary embodiments of the invention and, as such, should not be considered as limiting the scope of the invention that may admit to other equally effective embodiments. It is contemplated that features or steps of one embodiment may beneficially be incorporated in other embodiments without further recitation.
The present invention applies to systems and methods of searching a database of images or video using pictorial and/or semantic queries.
In a preferred embodiment, the queries may be parsed into tokens, or sub-images, of general classes of objects such as, but not limited to, faces, people, cars, guns and other objects. The tokens in the query are first checked for proper syntax. In the context of image and video searching, syntax includes the geometric and temporal relationships between the objects in the query. After the tokens are found to have a correct or acceptable syntax, a check may be made of the semantics, i.e., the meaning of the query. In the context of image searching, the meaning may be defined to be the similarities of the sub-images with all other images in the vocabulary. The vocabulary is a set of archetypical images representing the basic objects in the “Pictorial Language.” The tokens are similar to keywords in text based queries. The objects may be considered as binary large objects (BloBs), i.e., as a collection of binary data stored as a single entity in a database management system.
A preferred embodiment of the invention will now be described in detail by reference to the accompanying drawings in which, as far as possible, like elements are designated by like numbers.
Although every reasonable attempt is made in the accompanying drawings to represent the various elements of the embodiments in relative scale, it is not always possible to do so with the limitations of two-dimensional paper. Accordingly, in order to properly represent the relationships of various features among each other in the depicted embodiments and to properly demonstrate the invention in a reasonably simplified fashion, it is necessary at times to deviate from absolute scale in the attached drawings. However, one of ordinary skill in the art would fully appreciate and acknowledge any such scale deviations as not limiting the enablement of the disclosed embodiments.
The iconic GUI 14 allows the user to interact with the PLUTO system 10. The user may, for instance, “drag-and-drop” images as key-images for specifying particular instances of a person, object, event or other item. The user may also type in keywords for generic information such as, but not limited to, a class of objects. In addition to specific keywords and key-images, the PLUTO system 10 allows the user to specify relationships, such as, but not limited to, Boolean, temporal, and spatial relationships, between the keywords and key-images. In a preferred embodiment, the relationships are interpreted before the query is passed to the next level of processing.
The user may also use the GUI 14 to modify or preprocess the images before initiating the search. In a preferred embodiment, the additional functionalities supported by the GUI 14 may include, but are not limited to, morphing of two or more images, adding or subtracting the effect of ageing on a facial image, “ANDing” images by overlaying one image on another, and the ability to select regions in one or multiple image(s) indicating items that are to be present or absent in the search results. The GUI 14 also allows the user to specify time, geographic locations, or even camera settings for the search. The system may also, or instead extract, this information from the metadata tags attached to the images by the cameras taking the images such as the well known Exchangeable image file format (Eiff) metadata tags. These additional features available via the GUI 14 give the user the flexibility to input complex search queries to the system. After the search is performed, the GUI 14 may also display the results of the search in the order of relevance of the results.
The semantic input of the query is interpreted by the GUI 14 and passed on to the class detection module 16. The class detection module 16 uses the semantic input, i.e., the keyword and the Boolean relationships between keywords to find relevant images using a learned Support vector machines (SVM) or a Similarity Inverse Matrix (SIM) machine. The trained SVM or SIM machine is a trained classifier associated with each of the available keywords that detects objects relevant to the keyword. For instance, a user may want to search images with “face AND car.” The SVM or SIM would extract all images containing the two generic classes of faces and cars from the database. The class detection generates a list of objects relevant to the keywords input by the user. The SIM is described in detail in, for instance, co-pending U.S. patent application Ser. No. 11/619,121 filed on Jan. 2, 2007 by C. Podilchuk entitled “System and Method for Machine Learning using a Similarity Inverse Matrix,” the contents of which are hereby incorporated by reference.
The search results obtained from the conventional keyword search for tagged images/video are added to the list of objects found by the text search module 18. A rank ordered subset of these image templates are given to the file manager 22 by the object listing module 20 for performing image based search.
The image input from the user is passed from the GUI 14 to the file manager 22. A set of pre-selected images including images corresponding to the subset of objects detected by the class detection module 16 are used by the file manager 22 to compare with the query or key-image. A fast search technique may be employed by the file manager 22 that uses a pre-computed similarity matrix index and the identity transformation function to find the most relevant objects in the database. Such a fast search technique is described in detail, for instance, in co-pending U.S. patent application Ser. No. 11/619,104 filed on Jan. 2, 2007 by C. Podilchuk entitled “System and Method for Rapidly Searching a Database,” the contents of which are hereby incorporated by reference.
A similarity matrix is a pre-computed matrix of similarity scores between images using P-edit distances and the identity transformation function is required to map the objects of interest to their corresponding location in the similarity matrix.
Edit distances or Levenshtein distances were first used for string matching in 1965. Similar distance metrics have been used in the Smith-Waterman algorithm for local sequence alignment between two nucleotide or protein sequences. Dynamic Time Warping is used to align speech signals of different length. We apply this idea to image recognition and call it P-edit distance.
The P-edit or Picture-edit distances for calculating distances between images. The similarity matrix uses a similarity score based on P-edit distances. We obtain a vector field or an image disparity map between a query and a gallery image using the block matching algorithm. The properties of this mapping are translated into a P-edit distance which is used to compute the similarity score. The P-edit distance is described in detail in, for instance, co-pending U.S. patent application Ser. No. 11/619,092 filed on Jan. 2, 2007 by C. Podilchuk entitled “System and Method for Comparing Images using an Edit Distance,” the contents of which are hereby incorporated by reference.
The file manager 22 combines the ranked search results based on the key-image and the keywords input by the user and requests the multimedia database 12 that contains the relevant video and/or images to return the most relevant objects to the user via the GUI 14.
Long video streams containing objects of interest would be hard to search if every frame with the objects presence is used as a template. In a preferred embodiment of the invention, this problem is solved using a video tracker 30 that allows objects in video sequences to be tracked. In this way only a small number of templates need to be retained and searched to allow objects in the video to be identified. These templates may then used to populate the similarity matrix. When a query is made, the file manager 22 finds relevant images from the similarity matrix module 26 and refers the iconic database 23 to select the relevant video streams.
The identification module 28 contains the file locations in the multimedia database 12 of the actual images that are represented by similarity scores in the similarity matrix module 26.
Additional particulars of the method and system for searching multimedia databases using exemplar images are described in detail in, for instance, co-pending U.S. patent application Ser. No. 11/619,133 filed Jan. 2, 2007 by C. Podilchuk entitled “System and Method for Searching Multimedia using Exemplar Images,” the contents of which are incorporated herein by reference.
Although the invention has been described in language specific to structural features and/or methodological acts, it is to be understood that the invention defined in the appended claims is not necessarily limited to the specific features or acts described. Rather, the specific features and acts are disclosed as exemplary forms of implementing the claimed invention. Modifications may readily be devised by those ordinarily skilled in the art without departing from the spirit or scope of the present invention.
This application is related to, and claims priority from, the following: U.S. provisional patent application no. 60/861,686 filed on Nov. 29, 2006 by C. Podilchuk entitled “Method for multimedia information retrieval using a combination of text and exemplar images in the query,” U.S. provisional patent application no. 60/861,685 filed on Nov. 29, 2006 by C. Podilchuk entitled “New object/target recognition algorithm based on edit distances of images,” U.S. provisional patent application no. 60/861,932 filed on Nov. 30, 2006, by C. Podilchuk entitled “New learning machine based on the similarity inverse matrix (SIM),” U.S. provisional patent application no. 60/873,179 filed on Dec. 6, 2006 by C. Podilchuk entitled “Fast search paradigm of large databases using similarity or distance measures” and U.S. provisional patent no. 60/814,611 filed by C. Podilchuk on Jun. 16, 2006 entitled “Target tracking using adaptive target updates and occlusion detection and recovery,” the contents of all of which are hereby incorporated by reference.
| Number | Date | Country | |
|---|---|---|---|
| 60812646 | Jun 2006 | US | |
| 60816686 | Jun 2006 | US | |
| 60861685 | Nov 2006 | US | |
| 60861932 | Nov 2006 | US | |
| 60873179 | Dec 2006 | US | |
| 60814611 | Jun 2006 | US |