Claims
- 1. A method in a computer system for generating a weight for a term within one of a plurality of documents, the method comprising:generating a term frequency that represents a number of times that the term occurs in the one document; generating a total term frequency that represents a total number of times the term occurs in the plurality of documents; calculating a normalized term frequency by factoring the generated term frequency by a normalizing function of the generated total term frequency, wherein the normalizing function substantially equalizes result of term frequency and reciprocal total term frequency on the weight of the term; and combining the calculated normalized term frequency with a document frequency to generate the weight for the term.
- 2. The method of claim 1 wherein the normalizing function comprises a reciprocal of a root of the generated total term frequency.
- 3. The method of claim 2 wherein the root is a square root.
- 4. The method of claim 2 wherein the root is a cube root.
- 5. The method of claim 1 wherein the normalizing function comprises a reciprocal of a power of the generated total term frequency.
- 6. The method of claim 5 wherein the power is between 0 and 1.
- 7. The method of claim 1 wherein the normalizing function comprises a reciprocal of a logarithm of the generated total term frequency.
- 8. The method of claim 1 wherein the combining of the calculated normalized term frequency with a document frequency multiplies a logarithm of the calculated normalized term frequency by a logarithm of the document frequency.
- 9. The method of claim 8 wherein bases of the logarithms are different.
- 10. The method of claim 8 wherein bases of the logarithms are calculated so that on average the logarithms of the calculated normalized term frequency and the logarithms of the document frequency are equal.
- 11. The method of claim 8 wherein α=⌈β (logβ(1+a TF)logβNan)⌉where α s the base of the logarithm of the calculated normalized term frequency, where β is the base of the logarithm of the document frequency, where aTF is an average of the term frequencies for each term, where N is the number of documents, and where an is an average of the number of documents in which each term is contained.
- 12. The method of claim 8 wherein α=⌈β(logβ(β-1+aTF*(an*aTF)-1/2)logβNan)⌉where α is the base of the logarithm of the calculated normalized term frequency, where β is the base of the logarithm of the document frequency, where aTF is an average of the term frequencies for each term, where N is the number of documents, and where an is an average of the number of documents in which each term is contained.
- 13. The method of claim 8 wherein α=⌈β(logβ(β-1+aTF*logγ(γ-1+an*aTF)-1)logβNan)⌉where α is the base of the logarithm of the calculated normalized term frequency, where β is the base of the logarithm of the document frequency, where γ is the base of the logarithm for normalizing the term frequency, where aTF is an average of the term frequencies for each term, where N is the number of documents, and where an is an average of the number of documents in which each term is contained.
- 14. The method of claim 1 wherein the generated term frequency is an improved term frequency.
- 15. The method of claim 1 wherein the generated term frequency is enhanced based on factors that are independent of other terms in the document.
- 16. The method of claim 1 wherein the generated term frequency is adjusted based on factors related to the term and factors related to other terms in the document.
- 17. A computer-readable medium containing computer-readable instructions for performing the method of claim 1.
- 18. A method in a computer system for selecting a formula for weighting terms within one of a plurality of documents, the method comprising:generating an average term frequency that represents an average of term frequencies for at least one of a plurality of terms within each document, the term frequency being the number of times that a term occurs in a document; generating an average inverse document frequency that represents an average of inverse document frequencies for the at least one of a plurality of terms, the inverse document frequency of the term being the number of documents divided by the number of documents in which the term occurs; and selecting a first function and a second function so that the result of the first function of the generated average term frequency is approximately equal to the result of the second function of the generated average inverse document frequency.
- 19. The method of claim 18 wherein the first and second functions are a logarithmic functions and the identifying of the functions includes calculating a first base for a logarithm and a second base for a logarithm so that the logarithm of the first base of the generated average term frequency is approximately equal to the logarithm of the second base of the generated average inverse document frequency.
- 20. The method of claim 19 wherein the calculating of the second base calculates the second base to be equal to a multiplication factor by which a change in the generated average inverse document frequency results in a change of one in the logarithm of the second base of the generated average inverse document frequency.
- 21. The method of claim 19 wherein α=⌈β (logβ(1+a TF)logβNan)⌉where α is the first base, where β is the second base, where aTF is an average of the term frequencies for each term, where N is the number of documents, and where an is an average of the number of documents in which each term is contained.
- 22. The method of claim 19 wherein α=⌈β(logβ(β-1+aTF*(an*aTF)-1/2)logβNan)⌉where α is the first base, where β is the second base, where aTF is an average of the term frequencies for each term, where N is the number of documents, and where an is an average of the number of documents in which each term is contained.
- 23. The method of claim 19 wherein α=⌈β(logβ(β-1+aTF*logγ(γ-1+an*aTF)-1)logβNan)⌉where α is the first base, where β is the second base, where γ is the base of the logarithm for normalizing the term frequency, where aTF is an average of the term frequencies for each term, where N is the number of documents, and where an is an average of the number of documents in which each term is contained.
- 24. The method of claim 18 wherein the generated term frequency is an improved term frequency.
- 25. The method of claim 18 wherein the generated term frequency is enhanced based on factors that are independent of other terms in the document.
- 26. The method of claim 18 wherein the generated term frequency is adjusted based on factors related to the term and factors related to other terms in the document.
- 27. A system for generating a weight for a term within one of a plurality of documents, the system comprising:a first term frequency generator that computes a number of times that the term occurs in the one of the plurality of documents; a second term frequency generator that computes a total number of times the term occurs in the plurality of documents; a normalizer that calculates a normalized term frequency by factoring the number of times that the term occurs in the one of the plurality of documents by a normalizing function of the total number of times the term occurs in the plurality of documents, the normalizing function substantially equalizing the result of term frequency and reciprocal total term frequency on the weight of the term; and a combiner that combines the calculated normalized term frequency with a document frequency to generate the weight for the term.
CROSS-REFERENCE TO RELATED APPLICATION
This application claims the benefit of pending U.S. Provisional Application No. 60/103,718, filed Oct. 9, 1998, which application is incorporated by reference in its entirety.
US Referenced Citations (9)
Non-Patent Literature Citations (1)
| Entry |
| A Simple Example—Logarithms, Jun. 17, 1997, www.math.utah.edu/˜alfeld/math/sexample.html, pp. 1-2. |
Provisional Applications (1)
|
Number |
Date |
Country |
|
60/103718 |
Oct 1998 |
US |