The present invention relates to the field of computer systems. More specifically, the present invention relates to computer systems for visualizing analysis results.
Devices and computer systems for forming and using arrays of materials on a substrate are known. For example, PCT Publication No. WO 92/10588, incorporated herein by reference for all purposes, describes techniques for sequencing or sequence checking nucleic acids and other materials. Arrays for performing these operations may be formed according to the methods of, for example, the pioneering techniques disclosed in U.S. Pat. No. 5,143,854 and U.S. Pat. No. 5,593,839 both incorporated herein by reference for all purposes.
According to one aspect of the techniques described therein, an array of nucleic acid probes is fabricated at known locations on a substrate or chip. A fluorescently labeled nucleic acid is then brought into contact with the chip and a scanner generates an image file (which is processed into a cell file) indicating the locations where the labeled nucleic acids bound to the chip. Based upon the cell file and identities of the probes at specific locations, it becomes possible to extract information such as the monomer sequence of DNA or RNA. Such systems have been used to form, for example, arrays of DNA that may be used to study and detect mutations relevant to cystic fibrosis, the P53 gene (relevant to certain cancers), HIV, and other genetic characteristics.
Computer aided techniques for monitoring gene expression using such arrays of probes have also been developed as disclosed in U.S. patent application Ser. No. 08/828,952 (Attorney Docket No. 16528X-028900US) and PCT Publication No. WO 97/10365 (Attorney Docket No. 16528X-017110PC), the contents of which are herein incorporated by reference. Many disease states are characterized by differences in the expression levels of various genes either through changes in the copy number of the genetic DNA or through changes in levels of transcription (e.g., through control of initiation, provision of RNA precursors, RNA processing, etc.) of particular genes. For example, losses and gains of genetic material play an important role in malignant transformation and progression. Furthermore, changes in the expression (transcription) levels of particular genes (e.g., oncogenes or tumor suppressors), serve as signposts for the presence and progression of various cancers.
It is desirable to identify genes having expression levels relevant to diagnosis of a diseased state by analyzing the expression levels of large numbers of genes in both diseased and normal individuals. Methods for collecting the expression level information have been developed. However, the user interfaces for gene expression monitoring systems that have been developed until now are designed to clearly present the expression of particular pre-selected genes. A user seeking to identify, e.g., an oncogene or a tumor suppressor gene, must individually review the expression level of large numbers of genes and compare the expression levels between diseased and normal individuals. What is needed is a user interface that takes advantage of collected gene expression information to help the user to identify particular genes of interest.
The present invention provides innovative systems and methods for visualizing information collected from analyzing samples. The samples may include nucleic acids, proteins, or other polymers. Gene expression level as determined from analysis of a nucleic acid sample is one possible analysis result that may be visualized. In one embodiment, a computer system may display the expression levels of multiple genes simultaneously in a way that facilitates user identification of genes whose expression is significant to a characteristic such as disease or resistance to disease. Additionally, the computer system may facilitate display of further information about relevant genes once they are identified.
A first aspect of the invention provides a computer implemented method for presenting expression level information as collected from first and second samples. The method includes steps of: displaying a first axis corresponding to expression level in the first sample, and displaying a second axis substantially perpendicular to the first axis, the second axis corresponding to expression level in the second sample. The method further includes a step of: for a selected expressed sequence, displaying a mark at a position. The position is selected relative to the first axis in accordance with an expression level of the selected expressed sequence in the first sample and relative to the second axis in accordance with an expression level of the selected expressed sequence in the second sample. A particularly useful application is displaying many marks simultaneously for many selected genes to discover which ones of the selected genes may be relevant to the characteristic.
A second aspect of the invention provides a computer-implemented method of presenting sample analysis information. The method includes steps of: displaying a first axis corresponding to a concentration of a compound in a first sample as determined by monitoring binding of the compound to a selected polymer having binding affinity to the compound, and displaying a second axis substantially perpendicular to the first axis. The second axis corresponds to a concentration of the compound in the second sample as determined by monitoring binding of the compound to the selected polymer. The method further preferably includes a step of displaying a mark at a position. The position is selected relative to the first axis in accordance with the concentration in the first sample and relative to the second axis in accordance with the concentration in the second sample.
A further understanding of the nature and advantages of the inventions herein may be realized by reference to the remaining portions of the specification and the attached drawings.
The present invention provides innovative methods of monitoring visualizing gene expression. In the description that follows, the invention will be described in reference to preferred embodiments. However, the description is provided for purposes of illustration and not for limiting the spirit and scope of the invention.
Arrows such as 66 represent the system bus architecture of computer system 1. However, these arrows are illustrative of any interconnection scheme serving to link the subsystems. For example, display adapter 56 may be connected to central processor 50 through a local bus or the system may include a memory cache. Computer system 1 shown in
The VLSIPS™ and GeneChip™ technologies provide methods of making and using very large arrays of polymers, such as nucleic acids, on very small chips. See U.S. Pat. No. 5,143,854 and PCT Patent Publication Nos. WO 90/15070 and 92/10092, each of which is hereby incorporated by reference for all purposes. Nucleic acid probes on the chip are used to detect complementary nucleic acid sequences in a sample nucleic acid of interest (the “target” nucleic acid).
It should be understood that the probes need not be nucleic acid probes but may also be other receptors, such as antibodies, or polymers such as peptides. Peptide probes may be used to detect the concentration of other peptides, proteins, or other compounds in a sample. The probes must be carefully selected to have bonding affinity to the compound whose concentration they are to be used to measure.
In one embodiment, the present invention provides methods of visualizing information relating to the concentration of compounds in a sample as measured by monitoring affinity of the compounds to probes. In a particular application, the concentration information is generated by analysis of hybridization intensity files for a chip containing hybridized nucleic acid probes. The hybridization of a nucleic acid sample to certain probes may represent the expression level of one more genes or expressed sequence tags (ESTs). The expression level of a gene or EST is herein understood to be the concentration within a sample of mRNA or protein that would result from the transcription of the gene or EST.
Expression level information visualized by virtue of the present invention need not be obtained from probes but may originate from any source. If the expression information is collected from a probe array, the probe array need not meet any particular criteria for size and density. Furthermore, the present invention is not limited to visualizing fluorescent measurements of bondings such as hybridizations but may be readily utilized to visualize other measurements.
Concentration of compounds other than nucleic acids may be visualized according to one embodiment of the present invention. For example, a probe array may include peptide probes which may be exposed to protein samples, polypeptide samples, or other compounds which may or may not bond to the peptide probes. By appropriate selection of the peptide probes, one may detect the presence or absence of particular compounds which would bond to the peptide probes.
For purposes of illustration, the present invention is described as being part of a system that designs a chip mask, synthesizes the probes on the chip, labels nucleic acids from a target sample, and scans the hybridized probes. Such a system is set forth in U.S. Pat. No. 5,571,639 which is hereby incorporated by reference for all purposes. However, the present invention may be used separately from the overall system for analyzing data generated by such systems, such as at remote locations, or for visualizing the results of other systems for generating expression information, or for visualizing concentrations of polymers other than nucleic acids.
The chip design files are provided to a system 106 that designs the lithographic masks used in the fabrication of arrays of molecules such as DNA. The system or process 106 may include the hardware necessary to manufacture masks 110 and also the necessary computer hardware and software 108 necessary to lay the mask patterns out on the mask in an efficient manner. As with the other features in
The masks 110, as well as selected information relating to the design of the chips from system 100, are used in a synthesis system 112. Synthesis system 112 includes the necessary hardware and software used to fabricate arrays of polymers on a substrate or chip 114. For example, synthesizer 112 includes a light source 116 and a chemical flow cell 118 on which the substrate or chip 114 is placed. Mask 110 is placed between the light source and the substrate/chip, and the two are translated relative to each other at appropriate times for deprotection of selected regions of the chip. Selected chemical reagents are directed through flow cell 118 for coupling to deprotected regions, as well as for washing and other operations. All operations are preferably directed by an appropriately programmed computer 119, which may or may not be the same computer as the computer(s) used in mask design and mask making.
The substrates fabricated by synthesis system 112 are optionally diced into smaller chips and exposed to marked targets. The targets may or may not be complementary to one or more of the molecules on the substrate. The targets are marked with a label such as a fluorescein label (indicated by an asterisk in
The image file 124 is provided as input to an analysis system 126 that incorporates the visualization and analysis methods of the present invention. Again, the analysis system may be any one of a wide variety of computer system. The present invention provides various methods of analyzing and visualizing the chip design files and the image files, providing appropriate output 128. The chip design need not include any particular number of probes. It should be understood that the present invention does not require any particular source of expression level information.
At step 204 the system evaluates the sequences of interest to determine or assist the user in determining which probes would be desirable on the chip, and provides an appropriate “layout” on the chip for the probes. The process of selecting probes for an expression level analysis is explained in PCT Publication No. WO 97/10365, the contents of which are herein incorporated by reference. An alternative probe selection process that does not require prior knowledge of sequences of interest is explained in PCT Publication No. WO97/27317 (Attorney Docket No. 18547-019410PC), the contents of which are herein incorporated by reference. Further general background on probe selection is found in PCT Publication No. WO95/11995 (Attorney Docket No. 18547-004111PC) and PCT Publication No. WO97/29212 (Attorney Docket No. 18547-018540PC), the contents of which are herein incorporated by reference. The term “perfect match probe” refers to a probe that has a sequence that is perfectly complementary to a particular target sequence. The test probe is typically perfectly complementary to a portion (subsequence) of the target sequence. The term “mismatch control” or “mismatch probe” refer to probes whose sequence is deliberately selected not to be perfectly complementary to a particular target sequence. For each mismatch (MM) control in an array there typically exists a corresponding perfect match (PM) probe that is perfectly complementary to the same particular target sequence.
The process compares hybridization intensities of pairs of perfect match and mismatch probes that are preferably covalently attached to the surface of a substrate or chip. Most preferably, the nucleic acid probes have a density greater than about 60 different nucleic acid probes per 1 cm2 of the substrate.
Initially, nucleic acid probes are selected that are complementary to the target sequence. These probes are the perfect match probes. Another set of probes is specified that are intended to be not perfectly complementary to the target sequence. These probes are the mismatch probes and each mismatch probe includes at least one nucleotide mismatch from a perfect match probe. Accordingly, a mismatch probe and the perfect match probe to which it is identical except for one base make up a pair. As mentioned earlier, the nucleotide mismatch is preferably near the center of the mismatch probe.
The probe lengths of the perfect match probes are typically chosen to exhibit detectably greater hybridization with the target sequence relative to the mismatch probes. For example, the nucleic acid probes may be all 20-mers. However, probes of varying lengths may also be synthesized on the substrate for any number of reasons including resolving ambiguities.
Again referring to
At step 212 a computer system utilizes the layout information and the fluorescence information to evaluate the hybridized nucleic acid probes on the chip. Among the important pieces of information obtained from DNA chips are the relative fluorescent intensities obtained from the perfect match probes and mismatch probes. These intensity levels are used to estimate an expression level for a gene or EST. The computer system used for analysis will preferably have available other details of the experiment including possibly the gene name, gene sequence, probe sequences, probe locations on the substrate, and the like.
According to the present invention, at step 214, the same computer system used for analysis or another one displays the expression level information in a format useful for identifying genes of interest. The visualized expression level information may include information collected from multiple applications of one or more previous steps of
Hybridization intensities for a pair of probes are retrieved at step 954. The background signal intensity is subtracted from each of the hybridization intensities of the pair at step 956. Background subtraction can also be performed on all the raw scan data at the same time.
At step 958, the hybridization intensities of the pair of probes are compared to a difference threshold (D) and a ratio threshold (R). It is determined if the difference between the hybridization intensities of the pair (Ipm−Imm) is greater than or equal to the difference threshold AND the quotient of the hybridization intensities of the pair (Ipm/Imm) is greater than or equal to the ratio threshold. The difference thresholds are typically user defined values that have been determined to produce accurate expression monitoring of a gene or genes. In one embodiment, the difference threshold is 20 and the ratio threshold is 1.2.
If Ipm−Imm>=D and Ipm/Imm>=R, the value NPOS is incremented at step 960. In general, NPOS is a value that indicates the number of pairs of probes which have hybridization intensities indicating that the gene is likely expressed. NPOS is utilized in a determination of the expression of the gene.
At step 962, it is determined if Imm−Ipm>=D and Imm/Ipm>=R. If these expressions are true, the value NNEG is incremented at step 964. In general, NNEG is a value that indicates the number of pairs of probes which have hybridization intensities indicating that the gene is likely not expressed. NNEG, like NPOS, is utilized in a determination of the expression of the gene.
For each pair that exhibits hybridization intensities either indicating the gene is expressed or not expressed, a log ratio value (LR) and intensity difference value (IDIF) are calculated at step 966. LR is calculated by the log of the quotient of the hybridization intensities of the pair (Ipm/Imm). The IDIF is calculated by the difference between the hybridization intensities of the pair (Ipm−Imm). If there is a next pair of hybridization intensities at step 968, they are retrieved at step 954.
At step 972, a decision matrix is utilized to indicate if the gene is expressed. The decision matrix utilizes the values N, NPOS, NNEG, LR (multiple LRs), and IDIF (multiple IDIFs). The following four assignments are performed:
P1=NPOS/NNEG
P2=NPOS/N
P3=SUM(LR)/N
P4=SUM(IDIF)/N
These P values are then utilized to determine if the gene is expressed and if the expression level should be displayed. In a preferred embodiment, the expression level of a gene should be displayed if:
P1>2.2
P2>0.3
P3>0.8
P4>30
Once all the pairs of probes have been processed and the expression of the gene indicated, an average of the IDIF values for the probes that incremented NPOS or NNEG is calculated at step 975, which is utilized as an expression level. Of course, other values including one of P1 through P4 could be used to indicate expression level.
For simplicity,
The expression levels used for determining the position of marks 1006 are preferably taken from the result of step 975. The position of each of marks 1006 depends on two iterations of the steps of
In the depicted representative screen display, the first tissue sample is a cancerous tissue sample and the second tissue sample is a normal tissue sample. The individual marks represent the expression levels of selected genes in both cancerous and normal tissue. A first group of marks 1008 represent genes that are neither tumor suppressors nor oncogenes since their expression levels are roughly similar for both normal and cancerous tissue. These marks 1008 fall roughly along a line which is rotated 45 degrees from each of the axes. A second group of marks 1010 represent genes that are likely oncogenes since their expression levels are found to be significantly higher in cancerous tissue than in normal tissue. A third group of marks 1012 represent genes that are likely tumor suppressors since their expression levels are found to be significantly higher in normal tissue than in cancerous tissue. It will be appreciated that expression levels for large numbers of genes can be reviewed at once to discover the oncogenes and tumor suppressors.
Although in the depicted display, the two types of tissue are normal tissue and cancerous tissue, the present invention would aid in the discovery of genes whose expression is associated with any characteristic that varies among tissue samples. For example, once can compare expression results from tissue from individuals who have been exposed to HIV but remain infected to tissue obtained from infected individuals to identify genes conferring resistance to HIV. One can compare expression results between tissue from plants that survive drought to plants that do not. One can compare expression levels among tissue samples at successive stages or severity levels of the same disease, among tissue samples where different ultimate outcomes of the disease (e.g., patient death or remission) are known, among diseased tissue samples that have been subject to different treatment regimes including e.g., chemotherapy, antisense RNA, etc. For cancers, one can compare expression levels between malignant cells and non malignant cells. Also expression levels can be compared among different organs, between species, and among different stages of development of an organ.
It will be appreciated that the present invention also encompasses displays with more than two dimensions. A third visual dimension can be used to illustrate expression level from a third tissue sample. The time dimension can also be used to illustrate successive groups of two or three tissue samples at successive time periods.
The time dimension can be also used to correspond to tissue samples obtained at, e.g., successive stages of a disease.
Other interface methods corresponding to human senses other than sight can also be incorporated within the presentation system of the present invention. The senses may correspond to additional dimensions. For example, marks can be displayed in succession accompanies by a sound having characteristics corresponding to expression level in another tissue sample.
The user can employ a cursor 1014 to identify a particular mark as being of interest. Cursor 1014 can be moved to a particular mark by use of, e.g., mouse 11. Once cursor 1014 is over a mark of interest, the mark can be selected by, e.g., depression of one of mouse buttons 13. Selection of a particular mark can be facilitated by use of a zoom display feature (not shown). Once a particular mark is selected, further information is displayed about the gene represented by the mark. A special mouse can transmit a tactile sensation back to the user corresponding to expression level in a tissue sample as the user passes the mouse over a corresponding mark.
It will be appreciated that the display of
By selecting GenBank accession number 704 with another cursor (not shown), the user can direct retrieval of the GenBank information for the selected gene. If the GenBank information is not available locally, the retrieval process can include formulating a query and transmitting the query to a GenBank web site. Once the GenBank information is retrieved, it can also be displayed.
In the foregoing specification, the invention has been described with reference to specific exemplary embodiments thereof. It will, however, be evident that various modifications and changes may be made thereunto without departing from the broader spirit and scope of the invention as set forth in the appended claims and their full scope of equivalents.
This application is a divisional application of U.S. application Ser. No. 10/028,748, which is a continuation application of U.S. application Ser. No. 09/020,743, which claims priority to U.S. Provisional No. 60/069,436. U.S. application Ser. Nos. 10/028,748 and 09/020,743 are incorporated by reference herein, and U.S. application Ser. No. 09/020,743 has been issued into U.S. Pat. No. 6,420,108.
Number | Date | Country | |
---|---|---|---|
60069436 | Dec 1997 | US |
Number | Date | Country | |
---|---|---|---|
Parent | 10028748 | Dec 2001 | US |
Child | 11489292 | US |
Number | Date | Country | |
---|---|---|---|
Parent | 11489292 | Jul 2006 | US |
Child | 13626773 | US | |
Parent | 09020743 | Feb 1998 | US |
Child | 10028748 | US |