Saturation mutagenesis [1, 2]—coupled to an appropriate biological assay—represents a fundamental means of achieving a high-resolution understanding of regulatory [3] and protein-coding [4] nucleic acid sequences of interest. However, mutagenized sequences introduced in trans on episomes or via random or “safe-harbour” integration fail to capture the native context of the endogenous chromosomal locus [5]. This shortcoming markedly limits the interpretability of the resulting measurements of mutational impact.
Functional consequences of genetic variants are studied by manipulating the endogenous locus, which provides the native chromosomal context with respect to DNA sequence and epigenetic milieu, and for proteins, endogenous levels and patterns of expression [6]. Programmable endonucleases, e.g. zinc-finger nucleases (ZFNs), transcription activator-like effector nucleases (TALENs) or clustered regularly interspaced short palindromic repeat (CRISPR)/Cas-based RNA-guided DNA endonucleases, enable direct genome editing with increasing practicality [7]. However, genome editing has primarily been applied to introduce single changes to one or a few genomic loci [8], rather than many programmed changes to a single genomic locus. There remains a need for a genomic editing method for introducing multiple programmed changes into a single genomic locus in a single experiment.
In one aspect, this application relates to a method for introducing a plurality of programmed nucleotide modifications into a single locus of a desired genomic DNA sequence. The method entails the steps of: (a) synthesizing a homology-directed repair (HDR) library comprising a plurality of oligonucleotides, and (b) co-transfecting a population of cells with (i) an expression system capable of expressing Cas9 and a single guide RNA (sgRNA) and (ii) the HDR library, wherein the expression system is capable of introducing the plurality of oligonucleotides having the programmed nucleotide modifications to the locus of the desired genomic DNA sequence in one or more cells of the population. This method is carried out in a single experiment, i.e., in a single culture dish during a series of reactions carried out at the same time or within a single experimental protocol. In step (a), each oligonucleotide includes a programmed nucleotide modification in the locus of the desired genome.
In some embodiments, each programmed nucleotide modification is a single nucleotide variant. In some embodiments, the HDR library is constructed using an oligonucleotide including a degenerate sequence and optionally, a selective PCR site. For example, the degenerate sequence is between 1 and 100 nucleotides in length. In some embodiments, the HDR library contains a set of oligonucleotides having at least 100 unique programmed nucleotide modifications, at least 200 unique programmed nucleotide modifications, at least 300 unique programmed nucleotide modifications, at least 400 unique programmed nucleotide modifications, at least 500 unique programmed nucleotide modifications, at least 600 unique programmed nucleotide modifications, at least 700 unique programmed nucleotide modifications, at least 800 unique programmed nucleotide modifications, at least 900 unique programmed nucleotide modifications, at least 1,000 unique programmed nucleotide modifications, at least 3,000 unique programmed nucleotide modifications, at least 4,000 unique programmed nucleotide modifications, at least 5,000 unique programmed nucleotide modifications, at least 6,000 unique programmed nucleotide modifications, at least 7,000 unique programmed nucleotide modifications, at least 8,000 unique programmed nucleotide modifications, at least 9,000 unique programmed nucleotide modifications, at least 10,000 unique programmed nucleotide modifications, at least 12,000 unique programmed nucleotide modifications, at least 14,000 unique programmed nucleotide modifications, at least 16,000 unique programmed nucleotide modifications, at least 18,000 unique programmed nucleotide modifications, at least 20,000 unique programmed nucleotide modifications, at least 25,000 unique programmed nucleotide modifications, at least 30,000 unique programmed nucleotide modifications, at least 40,000 unique programmed nucleotide modifications, or at least 50,000 unique programmed nucleotide modifications.
In some embodiments, the plurality of programmed nucleotide modifications that are introduced to the locus of the desired genomic DNA sequence results in a saturating set of programmed nucleotide modifications. In some embodiments, the plurality of oligonucleotides are synthesized on a microarray. In other embodiments, the plurality of oligonucleotides are synthesized in column-based synthesis.
In some embodiments, the plurality of oligonucleotides that are synthesized as described above are used directly without any additional amplification or cloning steps. Alternatively, the plurality of oligonucleotides are amplified or cloned cloned to an HDR library before being used to introduce programmed nucleotide modifications.
In some embodiments, the expression system includes a Cas9 expression cassette that includes a nucleotide sequence which encodes a Cas9 nuclease, an sgRNA expression cassette, and a species-specific promoter that is specific to the population of cells.
In certain embodiments, each oligonucleotide of the HDR library includes a pair of homology arms.
In some embodiments, the method for introducing a plurality of programmed nucleotide modifications into a single locus of a desired genomic DNA sequence further entails the steps of: (c) harvesting the population of cells, (d) selectively amplifying a genomic DNA and RNA sample, wherein the edited sequences are amplified and the non-edited sequence are not amplified, and (e) sequencing the genomic DNA and RNA sample that has been selectively amplified, resulting in a set of genomic transcripts which include the plurality of programmed nucleotide modifications. Optionally, the method further entails functionally analyzing the set of genomic transcripts using a functional assay. For example, the functional assay is selected from the group consisting of targeted RNA sequencing to measure transcript abundance, targeted DNA sequencing to measure reduced cellular fitness, targeted chromatin immunoprecipitation-sequencing (CHiP-seq) of co-activators to assay enhancers, increased cellular growth rate to assay cancer drivers or drug resistance, and FACS-based phenotypic sorting for cellular assays.
In another aspect, this application relates to a method for analyzing the functional consequence of a genomic mutation. The method entails the steps of: (a) synthesizing a homology-directed repair (HDR) library including a plurality of oligonucleotides, wherein each oligonucleotide contains a programmed nucleotide modification in the locus of the desired genome, (b) co-transfecting a population of cells with (i) an expression system capable of expressing Cas9 and a single guide RNA (sgRNA) and (ii) the HDR library, wherein the expression system is capable of introducing the plurality of oligonucleotides having the programmed nucleotide modifications to the locus of the desired genomic DNA sequence in one or more cells of the population, (c) harvesting the population of cells, (d) selectively amplifying a genomic DNA and RNA sample, wherein the edited sequences are amplified and the non-edited sequence are not amplified, (e) sequencing the genomic DNA and RNA sample that has been selectively amplified, resulting in a set of genomic transcripts which include the plurality of programmed nucleotide modifications, and (f) functionally analyzing the set of genomic transcripts using a functional assay. This method is carried out in a single experiment, i.e., during a series of reactions carried out at the same time or within a single experimental protocol.
In some embodiments, the HDR library is constructed using an oligonucleotide containing a degenerate sequence and optionally, a selective PCR site. The degenerate sequence is between 1 and 100 nucleotides in length. In some embodiments, the plurality of oligonucleotides are synthesized on a microarray. In other embodiments, the plurality of oligonucleotides are synthesized in column-based synthesis. In some embodiments, the plurality of oligonucleotides are used directly without any cloning step to introduce programmed nucleotide modifications. Alternatively, the plurality of oligonucleotides are cloned to an HDR library before being used to introduce programmed nucleotide modifications.
In some embodiments, the expression system includes a Cas9 expression cassette having a nucleotide sequence which encodes a Cas9 nuclease, an sgRNA expression cassette, and a species promoter that is specific to the population of cells
In some embodiments, each oligonucleotide of the HDR library comprises a pair of homology arms.
In some embodiments, the functional assay is selected from the group consisting of targeted RNA sequencing to measure transcript abundance, targeted DNA sequencing to measure reduced cellular fitness, targeted chromatin immunoprecipitation-sequencing (CHiP-seq) of co-activators to assay enhancers, increased cellular growth rate to assay cancer drivers or drug resistance, and FACS-based phenotypic sorting for cellular assays.
In another aspect, this application relates to a method for genomic screening. The method entails the steps of: (a) introducing a plurality of programmed nucleotide modifications to a single genomic locus, wherein the plurality of programmed nucleotide modifications are introduced in a single experiment, i.e., during a series of reactions carried out at the same time or within a single experimental protocol, (b) sequencing the genomic DNA or cDNA of the edited locus, and (c) quantifying the transcript abundance of each mutation.
In some embodiments, step (a) includes (1) synthesizing a homology-directed repair (HDR) library comprising a plurality of oligonucleotides, wherein each oligonucleotide includes a programmed nucleotide modification in the locus of the desired genome, and (2) co-transfecting a population of cells with (i) an expression system capable of expressing Cas9 and a guide RNA (sgRNA) and (ii) the HDR library, wherein the expression system is capable of introducing the plurality of oligonucleotides having the programmed nucleotide modifications to the locus of the desired genomic DNA sequence in one or more cells of the population. This step is carried out in a single experiment, i.e., during a series of reactions carried out at the same time or within a single experimental protocol. In some embodiments, step (c) includes calculating an enrichment score for each mutation.
a-1c show saturation genome editing and multiplex functional analysis of a hexamer region influencing BRCA1 splicing.
a-2c show that multiplex homology-directed repair reveals effects of single nucleotide variants on transcript abundance. Three separate HDR libraries (R, R2, and L) containing a 3% mutation rate (97% wt, 1% each non-wt base) in either half of BRCA1 exon 18 were introduced to the genome via co-transfection with pCas9-sgBRCA1x18. Enrichment scores were calculated for each haplotype observed at least 10 times in the gDNA, and effect sizes of SNVs were determined by weighted linear regression modeling. ‘Sense’ includes both missense and synonymous SNVs.
a-3c show saturation genome editing and multiplex functional analysis at an essential gene, DBR1, in Hap1 cells. An HDR library targeting a highly conserved region of DBR1 exon 2 was used with pCas9-EGFP-sgDbr1x2 to introduce point mutations across 75 base pair (bp) and all possible codon substitutions at three residues believed to participate at the enzyme's active site.
a-4d show the distribution and pair-wise correlations of hexamer abundances.
a-5d show correlations for hexamer genome editing efficiency and enrichment scores between replicates.
a-6b show comparison of genome-based hexamer enrichment scores to plasmid-based hexamer scores.
a-8c show positional SNV editing rates and replication of effect sizes.
a-10c show correlation between effect sizes and predicted disruption of splicing motifs and indel effects.
a-12b show DBR1 editing rates by position and comparison of haplotype abundances between D5 and the HDR library, D8, and D11.
a-13c show performance of computational predictions of deleterious DBR1 mutations and reproducibility between biological replicates.
Methods for introducing multiple programmed nucleotide modifications into a single locus of a desired genomic DNA sequence are provided herein. The methods described herein are carried out in a single experiment, i.e., during a series of reactions carried out concurrently within a single experimental protocol in a single culture dish. Such methods may be used to analyze the functional consequence of a genomic mutation or for genomic screening.
To overcome the limitations of previously used methods, a method was developed to generate and functionally analyze hundreds to thousands of programmed genome edits at a single locus in a single experiment. The method allows a more accurate and scalable measurement of the functional consequences of genetic variations. Measurement of the functional consequences of large numbers of mutations with saturation genome editing potentially facilitates high-resolution functional dissection of both cis-regulatory elements and trans-acting factors, as well as the interpretation of variants of uncertain significance observed in clinical sequencing.
According to the embodiments described herein, this application relates to a method for introducing a plurality of programmed nucleotide modifications into a single locus of a desired genomic DNA sequence. Saturation editing of genomic regions may be achieved by coupling CRISPR/Cas9 RNA-guided cleavage [10] with multiplex homology-directed repair (HDR) using a complex library of donor templates. “Saturation editing” as used herein means that for a particular sequence, each nucleotide position of that sequence is systematically modified with each of all four traditional bases, A, T, G and C. For example, a hexamer substituted at each position would have 4,096 possible single nucleotide variants (four possible substitutions at each of the six nucleotide positions of the hexamer, or 46=4,096).
The multiple programmed nucleotide modifications may be introduced during a single experiment. The phrase “a single experiment” means that multiple programmed edits are introduced to a region of a particular size within a single culture dish, during the course of one experiment or a series of reactions within a single experimental protocol. In other words, the programmed edits are introduced concurrently in a single culture dish. In certain embodiments, the single experimental protocol includes one or more concurrent reactions, i.e., the multiple programmed edits are introduced at approximately the same time. This is in contrast to programmed edits being introduced one-by-one or several at a time in physically separated reactions or experiments carried out using one or more multi-well or otherwise separated culture dishes. The region to be introduced with the programmed edits may have an optimal size that allows efficient multiplex editing in one experiment, for example, the region is about between 1 and 100 base pairs in size. The window associated with HDR mechanisms in mammalian cells [11] may affect the size of the region that can be subjected to multiple editing in one experiment. Therefore, it is within the purview of one skilled in the art to determine the size of the region for multiplex editing. By the same token, saturation genome editing of a full gene—e.g. to measure functional consequences of all possible variants of uncertain significance—will likely require multiple experiments tiling along its exons.
The terms “programmed modifications,” “programmed gene edits,” and “programmed edits” as used herein are interchangeable, meaning that for a particular oligonucleotide, one or more nucleotides at a particular position is changed, for example, from A to T, G, or C. In some embodiments, each programmed nucleotide modification or edit is a single nucleotide variant (SNV). The programmed changes may result in a deletion, substitution, insertion, or other type of mutation to the gene.
The method includes a step of synthesizing a homology-directed repair (HDR) library comprising a plurality of oligonucleotides, each of which includes a programmed nucleotide modification. The plurality of oligonucleotides may be synthesized on a microarray or in column-based synthesis. In some embodiments, the HDR library is constructed using an oligonucleotide having a degenerate sequence and optionally, a selective PCR site.
The degenerate sequence may be between 1 and 100 nucleotides in length, and the constructed library using the degenerate sequence may have a set of oligonucleotides having at least 100 unique programmed nucleotide modifications, at least 200 unique programmed nucleotide modifications, at least 300 unique programmed nucleotide modifications, at least 400 unique programmed nucleotide modifications, at least 500 unique programmed nucleotide modifications, at least 600 unique programmed nucleotide modifications, at least 700 unique programmed nucleotide modifications, at least 800 unique programmed nucleotide modifications, at least 900 unique programmed nucleotide modifications, at least 1,000 unique programmed nucleotide modifications, at least 3,000 unique programmed nucleotide modifications, at least 4,000 unique programmed nucleotide modifications, at least 5,000 unique programmed nucleotide modifications, at least 6,000 unique programmed nucleotide modifications, at least 7,000 unique programmed nucleotide modifications, at least 8,000 unique programmed nucleotide modifications, at least 9,000 unique programmed nucleotide modifications, at least 10,000 unique programmed nucleotide modifications, at least 12,000 unique programmed nucleotide modifications, at least 14,000 unique programmed nucleotide modifications, at least 16,000 unique programmed nucleotide modifications, at least 18,000 unique programmed nucleotide modifications, at least 20,000 unique programmed nucleotide modifications, at least 25,000 unique programmed nucleotide modifications, at least 30,000 unique programmed nucleotide modifications, at least 40,000 unique programmed nucleotide modifications, or at least 50,000 unique programmed nucleotide modifications.
The method further includes a step of co-transfecting a population of cells with (i) an expression system capable of expressing Cas9 and a guide RNA (sgRNA) and (ii) the HDR library. In such embodiments, the expression system acts to introduce a plurality of oligonucleotides (each of which includes a programmed nucleotide modification) to the locus of the desired genomic DNA sequence in one or more cells of the population. In certain embodiments, the expression system includes a plasmid which includes a Cas9 expression cassette that includes a nucleotide sequence which encodes a Cas9 nuclease, an sgRNA expression cassette, and a species-specific promoter that is specific to the population of cells. In certain aspects, each oligonucleotide member of the HDR library includes a pair of homology arms in order to target the desired genomic DNA sequence.
The method described herein may further include one or more steps of harvesting the population of cells after culturing the transfected cells, selectively amplifying a genomic DNA and RNA sample wherein the edited sequences are amplified and the non-edited sequence are not amplified, and sequencing the genomic DNA and RNA sample that has been selectively amplified, resulting in a set of genomic transcripts which include the plurality of programmed nucleotide modifications. In some embodiments, the method includes functionally analyzing the set of genomic transcripts using a functional assay, such as targeted RNA sequencing to measure transcript abundance, targeted DNA sequencing to measure reduced cellular fitness, targeted chromatin immunoprecipitation-sequencing (CHiP-seq) of co-activators to assay enhancers, increased cellular growth rate to assay cancer drivers or drug resistance, and FACS-based phenotypic sorting for cellular assays.
In one aspect, this application also relates to analyzing the functional consequence of a genomic mutation by carrying out the above steps, including: (a) synthesizing a homology-directed repair (HDR) library comprising a plurality of oligonucleotides, wherein each oligonucleotide comprises a programmed nucleotide modification in the locus of the desired genome; (b) co-transfecting a population of cells with (i) an expression system capable of expressing Cas9 and a guide RNA (sgRNA) and (ii) the HDR library, wherein the expression system is capable of introducing the plurality of oligonucleotides having the programmed nucleotide modifications to the locus of the desired genomic DNA sequence in one or more cells of the population; (c) harvesting the population of cells; (d) selectively amplifying a genomic DNA and RNA sample, wherein the edited sequences are amplified and the non-edited sequence are not amplified; (e) sequencing the genomic DNA and RNA sample that has been selectively amplified, resulting in a set of genomic transcripts which include the plurality of programmed nucleotide modifications; and (f) functionally analyzing the set of genomic transcripts using a functional assay.
In some embodiments, the functional assay is biologically relevant and technically viable. In some embodiments, the functional assay directly links genotype to phenotype. For example, the functional assay is a targeted RNA sequencing to measure transcript abundance or targeted DNA sequencing to measure reduced cellular fitness. In other embodiments, the functional assay is targeted ChIP-seq of co-activators to assay enhancers, increased cellular growth rate to assay cancer drivers or drug resistance [31], or FACS-based phenotypic sorting for cellular assays [32].
Also described herein is a method for genomic screening. The method includes a first step of introducing a plurality of programmed nucleotide modifications to a single genomic locus, for example, as described above. The genomic screening method further includes sequencing the genomic DNA or cDNA of the edited locus, and quantifying the transcript abundance of each mutation, e.g., by calculating an enrichment score for each mutation.
For illustration purposes, the saturation genome edits were introduced to exon 18 of BRCA1 and to a well-conserved coding region of an essential gene, DBR1, respectively. By no means the scope of this application is limited to these particular genes. It is within the purview of one skilled in the art to introduce genome edits including saturation genome edits to any gene of interest by carrying out the methods disclosed herein.
In exon 18 of BRCA1, a six base-pair (bp) genomic region was replaced with all possible hexamers, or the full exon was replaced with all possible single nucleotide variants (SNVs), and the effects on transcript abundance attributable to nonsense-mediated decay and exonic splicing elements were measured. Saturation genome edits were introduced to DBR1 in a similar fashion and the relative effects on growth that correlate with functional impact were measured.
In one embodiment, the methods described herein are exemplified by leveraging CRISPR/Cas9 [10, 12, 13] to introduce saturating sets of programmed edits to a specific locus via multiplex HDR. As illustrated in
pCas9-sgBRCA1x18 and the HDR library were co-transfected into ˜800,000 HEK293T cells, achieving 3.33% HDR efficiency. Two independent transfections were performed with the same HDR library (biological replicates' 1, 2), and cells were split on day 3 (‘D3 replicates’ a, b).
Genomic DNA (gDNA) and cDNA from bulk cells on D5 were prepared. PCR reactions were primed on the ‘handle’ uniquely present within successfully edited genomes. Amplification was observed in HDR library/pCas9-sgBRCA1x18 transfected samples, but not in HDR library-only controls. Amplicons derived from gDNA and cDNA were deeply sequenced (
The effect of introducing each hexamer to these genomic coordinates on transcript abundance was estimated by calculating enrichment scores (cDNA divided by gDNA counts, calibrated to wild-type). These enrichment scores were well correlated between biological replicates (
To maximize precision, data across all four replicates for 4,048 hexamers were merged (
In some embodiments,
Using data from all edited exons with ≧1 mutation and ≧10 gDNA counts, effect sizes (beta values) of all possible SNVs were estimated using a weighted linear model. Estimated effect sizes were reproducible (R=0.846 (R), 0.853 (R2), and 0.686 (L);
The estimated effect sizes reflect empirically measured changes in transcript abundance resulting from programmed edits (
In another embodiment,
An optimized single guide RNA (sgRNA) sequence [23, 24] was cloned into a bicistronic sgRNA/Cas9-2A-EGFP vector (pCas9-EGFP-sgDbr1x2). Five million haploid human cells [25] (Hap1) were co-transfected with the DBR1 HDR library and pCas9-EGFP-sgDbr1x2. On D2, ˜250,000 EGFP+ cells were FACS sorted and further cultured, taking samples on D5, D8 and D11 (1.14% HDR efficiency, estimated on D8). Following gDNA isolation and selective PCR, deep sequencing was performed to quantify the relative abundance of edited haplotypes in each sample.
The relative proportions of mutation classes at each time point were first examined (
Amino-acid level enrichment scores were well correlated between D11 biological replicates (R=0.752; P=2.6×10-40;
Genome editing efficiency may be affected by factors such as bottlenecking complexity, limiting reproducibility and in some cases, necessitating the optional selective PCR sites. However, selective PCR sites are not necessarily required in all cases. In some embodiments, a variety of techniques, e.g. transient hypothermia [27] or oligonucleotide-based HDR [28], can be used to improve editing efficiency. In some embodiments, ZFNs and TALENs may improve efficiencies up to 50% [29, 30].
In some embodiments, haploid cells for DBR1 mutagenesis can be used to improve editing efficiency. In other embodiments, mutagenesis can be performed in diploid cells by knocking out one allele via NHEJ and then knocking in the HDR library to the other allele.
The following examples are intended to illustrate various embodiments of the invention. As such, the specific embodiments discussed are not to be construed as limitations on the scope of the invention. It will be apparent to one skilled in the art that various equivalents, changes, and modifications may be made without departing from the scope of invention, and it is understood that such equivalent embodiments are to be included herein. Further, all references cited in the disclosure are hereby incorporated by reference in their entireties, as if fully set forth herein.
An exon in a clinically relevant gene in which known mutations cause aberrant splicing was chosen to be targeted. Previous molecular studies of a G to T nonsense mutation occurring naturally in cancer patients at chr17:41,215,963 suggested exon skipping [14] was secondary to the creation of an exonic splicing silencer site [33]. It is hypothesized that saturation genome editing of this exon could result in a wide range of splicing outcomes.
When performing parallel functional analysis of complex allelic series, how to associate each of many mutations with the biological effects they produce should be considered. It is more difficult when attempting such approaches at the endogenous genomic locus, and with limited editing efficiencies. By performing these experiments in an exon and focusing on the effects of mutations on transcript abundance, genotype and phenotype are directly linked by observing the frequency of each genome edit in the transcript pool, relative to its frequency in genomic DNA. This design is advantageous because it requires no specialized (i.e. gene-specific) functional assay, thus making it amenable to interrogation of transcribed variants' effects on splicing/transcript abundance in any gene.
Inclusion of Selective PCR Sites
Given the modest proportion of HDR-edited loci in a given experiment and the high number of variants to be interrogated (i.e. hundreds to thousands), it would require a large amount of sequencing to sufficiently sample every variant in gDNA and cDNA pools from a population of cells that are predominantly unedited or harboring products of NHEJ. Furthermore, at such efficiencies, the rate of error in high-throughput sequencing is high enough to obscure signal from single nucleotide variants (SNVs) (unpublished observations). Therefore, in certain embodiments, until better methods are developed, techniques to selectively sequence molecules derived from edited cells are likely to be advantageous to isolate populations of cells that have been successfully edited with HDR techniques. In some embodiments, selective PCR sites are present regardless how the HDR libraries are generated, for example, by degenerate oligonucleotide or by programmed edits via microarray-based synthesis. In other embodiments, selective PCR sites are not used.
The HDR libraries were designed to include short, fixed edits to serve as unique priming sites in genomes that successfully undergo HDR. PCR reactions primed at this site, therefore, should only amplify material from edited cells, thus reducing both the noise associated with error from sequencing unedited material and the cost of sequencing in each experiment. Additionally, selective PCR sites that mutate the PAM and protospacer sequences could prevent Cas9 from re-cutting HDR-edited genomes. This should have the effect of increasing the proportion of cells bearing experimentally informative edits, and given the bottleneck imposed by limitations on how many successfully edited cells can be sampled, should result in more robust experimental signal.
DBR1 Experimental Design
To demonstrate that saturation genome editing can be used to explore effects of mutations on protein function and cellular fitness, DBR1, a well-conserved gene that scored highly in a human haploid cell genome-wide loss-of-function screen for essentiality [19] was targeted. Using haploid cells prevents gene compensation from an unedited copy [25]. Without knowing how sensitive the cells would be to mutations, it was chosen to target a region of exon 2 that was highly conserved, included in all transcript annotations on the UCSC Genome Browser, and coded for at least 2 residues (N84, H85) predicted to participate at the enzyme's active site [26]. Selection against edited cells in culture allows phenotype to be linked to genotype from sequencing of the gDNA pool over a series of time points. During HDR library construction, a selective PCR site in a downstream intron was designed to minimize any effect on gene function, and two synonymous mutations to abrogate Cas9 re-cutting were used.
Given the lower transfection efficiency of Hap1 cells (˜4% for the plasmids used here), a DBR1-targeting CRISPR construct that expressed EGFP with Cas9 was cloned and FACS was used to sort a population of successfully transfected cells. The sgRNA was designed using the Zhang Lab tool [described at http://crispr.mit.edu/], and selected to minimize off-target effects that could potentially impair cellular fitness [23].
HDR Library and Cas9-sgRNA Cloning
A homology-directed repair (HDR) library containing all possible 4,096 DNA hexamers substituted at positions +5 to +10 of BRCA1 exon 18 (chr17:41,215,962-41,215,967; CCDS11453.1) was constructed using a partially degenerate oligonucleotide (IDT DNA; “BRCA1ex18NNNNNN5—10 selPCR”) containing a 7 bp selective PCR site/EcoRV restriction digest site at position +17 to +23 (
The DBR1 HDR library was cloned as above except with the following differences. HDR library variants were derived from 388 oligonucleotides synthesized on a microarray (CustomArray) to include all possible single base pair changes in a 75 bp region comprising part of DBR1 exon 2 (chr3:137,892,342-137,892,416), all codon variants at the first three residues of the 75 bp region (chr3:137,892,408-137,892,416), and the reference 75 bp sequence. All DBR1 HDR library sequences also included two synonymous mutations designed to prevent re-cutting of edited genomes by disrupting PAM and protospacer sequences (chr3:137,892,424 and chr3:137,892,421), and a 6 bp selective PCR site in intron 2 of DBR1 (chr3:137,892,331-137,892,336). The library was cloned into a pUC19-DBR1ex2 backbone, a vector containing the surrounding DBR1 sequence cloned from Hap1 gDNA (chr3:137,891,573-137,893,293).
A bicistronic Cas9-sgRNA vector designed to cleave within BRCA1 exon 18 (“pCas9-sgBRCA1x18”) was cloned according to a published protocol[24] by ligating annealed oligonucleotides into a human codon-optimized S. pyogenous Cas9-sgRNA vector from the lab of Feng Zhang (pX330-U6-Chimeric_BB-CBh-hSpCas9; Addgene plasmid #42230). The same protocol was followed to create pCas9-EGFP-sgDbr1x2 from a similar Zhang lab vector that allows for fluorescent identification of Cas9-expressing cells (pSpCas9(BB)-2A-GFP (pX458); Addgene plasmid #48138).
Cell Culture and Transfection
For BRCA1 experiments, HEK293T cells were cultured in Dulbecco's Modified Eagle Medium (Life Technologies) supplemented with 10% FBS (AATC) and 100 U/ml penicillin+100 ug/ml streptomycin (Life Technologies). One day prior to transfection, cells were split to ˜40% confluency in 12-well plates with antibiotic-free media. The next day, 0.5-1.0 ug of each library was co-transfected (Lipofectamine 2000, Invitrogen) with an equivalent amount of pCas9-sgBRCA1x18. Cells were expanded to 6-well plates, then split 1:4 on day 3 into two pools, and DNA and RNA were harvested on D5 (AllPrep DNA/RNA Mini Kit, Qiagen). Biological replicates of each transfection and negative control transfections of each library without pCas9-sgBRCA1x18 were also performed.
For the DBR1 experiment, Hap1 cells (Haplogen) were cultured in Iscove's Modified Dulbecco's Medium supplemented with 10% FBS and 100 U/ml penicillin+100 ug/ml streptomycin. ˜3×106 Hap1 cells were passaged to a 60 mm dish in antibiotic-free media one day prior to co-transfection with 3 ug each of pCas9-EGFP-sgDbr1x2 and the DBR1 HDR library via Turbofectin 8.0 (OriGene) according to protocol. On D2, FACS was performed (BD FACSAria III) to isolate ˜250,000 EGFP+ cells, which were then expanded in culture with samples taken of ˜1×106 cells on D5, and 4-8×106 on D8 and D11. gDNA was isolated according to protocol with the QiaAmp Kit (Qiagen). A biological replicate was performed, as well as negative controls in which the HDR library was transfected with the empty pSpCas9(BB)-2A-GFP construct (to enable FACS of transfected cells without editing).
RT, Selective PCR and Sequencing
For BRCA1 experiments, reverse transcription (RT) was performed using SuperScriptIII (Invitrogen) with a gene-specific primer located in either BRCA1 exon 19 (hexamer experiments) or exon 21 (whole exon experiments). Initial rounds of PCR were performed on large quantities of sample gDNA (8-12 ug gDNA, 100-150 ng/reaction) and cDNA (25 ug total RNA reverse transcribed and split into 45-47 reactions) using the KAPA HiFi HotStart ReadyMix PCR kit. In the first gDNA PCR, a primer external to the HDR library was used to prevent amplification of plasmid DNA. cDNA reactions were either primed from exons 16 and 18 (hexamer experiment; Library L) or exons 18 and 20 (Libraries R, R2). After the initial gDNA and cDNA reactions, all PCR products from a single sample were pooled and purified using the QIAquick PCR Purification Kit (Qiagen).
For both cDNA and gDNA reactions, a primer designed to selectively amplify edited molecules bearing the selective PCR site was used either in the first or second reaction. Optimal annealing temperatures for each primer pair were determined via gradient PCR, and negative control reactions were performed using input from HDR library-only transfections to ensure products were derived from edited genomes as opposed to the HDR library. Negative controls failed to amplify for all experiments. Two subsequent PCRs were performed to add sequencing adaptors (“PU1L” and “PU1R”), sample indices, and flow cell adaptors.
For the DBR1 experiment, 30 cycles of selective PCR were performed on gDNA (300 ng per reaction) from D5 (3 ug), D8 and D11 (27 ug each). Wells from each sample were pooled, PCR purified, and then re-amplified for 15 additional cycles. The 1,055 bp product was gel-purified (QIAquick Gel Extraction Kit, Qiagen), and two subsequent PCRs were performed to incorporate sequencing and flow cell adaptors prior to sequencing as above.
After final reactions were purified (AMPure XP beads, Agencourt), paired-end sequencing was performed on all samples with the Illumina MiSeq to quantify gDNA and/or cDNA abundances for each edited haplotype. All primer sequences for RT, selective PCR, and sequencing library preparation are provided in Table 1.
HDR efficiencies were estimated for all experiments via deep sequencing of target loci by performing PCR on 150-300 ng of gDNA using primers external to the region of editing and the selective PCR site. Reported HDR efficiencies were conservatively calculated as the fraction of sequencing reads containing the selective PCR site and bearing at least one variant represented in the HDR library.
Analysis of Sequencing Data
For quality control, fully overlapping paired-end reads were merged with PEAR [34] (Paired-End reAd mergeR) and discordant pairs were eliminated. By design, the mutagenized region is covered by both the forward and reverse reads on the Illumina platform, resulting in high-confidence calls per site.
For BRCA1 hexamer reads to be included, the six bases on either side of the hexamer were required to match the reference sequence, and every base call in the hexamer required a quality score of at least Q30. For BRCA1 whole-exon mutagenesis, the full read was required to be the correct length and match the library consensus sequence outside of the mutagenized region, every base quality score inside the mutagenized region was required to be at least Q30, and no indels were tolerated in alignment with BWA-MEM [35]. cDNA reads not matching any gDNA haplotype with at least 10 reads were eliminated. After normalizing for sequencing coverage, enrichment scores were calculated as cDNA read counts incremented by one pseudocount divided by gDNA reads, calibrated to the wild-type hexamer.
For DBR1 mutagenesis, reads were subjected to the same requirements of the sequence outside the mutagenized bases matching the consensus and every quality score in the mutagenized region exceeding Q30. Only reads matching programmed haplotypes were analyzed, and haplotypes below a D5 relative abundance of 5E-5 were excluded from analysis. After incrementing all read counts by one pseudocount and dividing by the total number of reads, the abundance of each haplotype on D8 or D11 was divided by the corresponding abundance on D5, and the fold change relative to the wild type sequence was taken to calculate an enrichment score. Based on the bimodal distribution observed in each replicate, mutations with log2-transformed enrichment scores less than −2 were considered “deleterious”; otherwise, mutations were considered “tolerated”. Discordant effects between replicates were defined as mutations “tolerated” in one replicate but “deleterious” in the other. Amino acid level enrichment scores were calculated as the median of SNV enrichment scores for programmed edits resulting in the same change (or lack of change, for synonymous edits).
SNV Effect Size Linear Modeling and Replicate Pooling
To determine effects of SNVs in the BRCA1 whole-exon experiments, cDNA and gDNA read counts were converted into percentages (number of reads for a given haplotype divided by the total number of reads for a given replicate) after discarding haplotypes with fewer than 10 gDNA reads. Because we had variance in the number of reads for each haplotype, the null expectation of equal variance (σ2) for each cDNA/gDNA ratio was violated. Because each effect size (yij) was the average of nij observations (reads), then var(yij)=var εij=σ2/nij, suggesting that the weight for each variable should be nij. To predict single nucleotide effect size across exon 18 of BRCA1, we then fit the weighted linear model:
y
ij=β0+wijβijXij
where yij is the log2 enrichment score for a given haplotype, is the number of gDNA reads for a given haplotype, βij is the effect of nucleotide i at position j relative to the wild-type allele, and Xij is a dummy variable indicating the presence or absence of a particular nucleotide change i at position j relative to the wild type allele. Regression analyses were performed in R 3.0.0 using the lm( ) function. The resulting coefficients of the model adjusted for the model intercepts (β0+βij) were taken as effect sizes of the individual SNVs on exon splicing/stability. To merge data across replicates, effect sizes were averaged (including across overlapping bases between libraries L and R in BRCA1 exon).
Comparisons to Other Metrics of Functional Impact
For comparison to plasmid studies, ESR-seq scores were taken from Ke et al. (2011) [9]. Hexamers with positive ESR-seq scores are deemed exonic splicing enhancers, whereas negative ESR-seq scores denote exonic splicing silencers. For comparison of BRCA1 exon 18's SNV effect sizes to an in silico method, all SNVs were queried on MutPredSplice's web server (http://mutdb.org/mutpredsplice/submit.htm). MutPredSplice reports a single score estimating the likelihood that a variant will disrupt splicing at any genomic locus. Absolute values of BRCA1 exon 18 splicing effect sizes were then correlated with MutPredSplice scores to determine concordance between our data and predicted effects on splicing.
For DBR1, calculated enrichment scores were compared to BLOSUM62 substitution scores [20] (obtained from NCBI), PolyPhen-2 [21], and CADD [22] (PolyPhen-2 and CADD scores obtained from querying genomic coordinates from CADD's precomputed genomic annotations (http://cadd.gs.washington.edu/download). Whereas BLOSUM62 is derived from evolutionary conservation and PolyPhen-2 predicts changes in protein function, CADD is an integrated measure of deleteriousness that incorporates many functional annotations (including PolyPhen-2).
Reproducibility of Saturation Genome Editing Experiments
The correlations between replicates for each of the experiments suggest that while this technique reproducibly measures effects of many concurrent programmed genome edits, there are also sources of noise.
The noise observed may relate to the fact that modest editing efficiencies lead to relatively few cells in each experiment harboring each specific edit. In the BRCA1 hexamer experiment and the DBR1 experiment, a bimodal distribution of gDNA read counts is observed (
Consistent with lowly sampled edits being more prone to noise, hexamers that are more highly represented in gDNA counts are more reproducible. For example, whereas R=0.659 between two biological replicates overall, hexamers falling into the top third with respect to gDNA count correlated much more highly (R=0.857). Furthermore, considering the two BRCA1 experiments, because there were far fewer possible SNVs (n=234; experiment in
Whether the noise represents biological variability (for instance, two cells with the same edit producing transcripts at different rates) or technical variability (stochastic effects inherent to sample prep) it is reasoned that by pooling or averaging replicates, the number of successfully edited cells sampled is effectively increased, and therefore noise attributable to low sampling is reduced. Consistent with this, pooling read counts from D3 replicates in the BRCA1 hexamer experiment improved correlation between biological replicates.
For the DBR1 experiment, the overall reproducibility of D11 enrichment scores is reasonable (R=0.752;
This experiment, subject to bottlenecking at the editing step, generates clonal populations possibly expanded from a single edited cell. Falsely tolerated edits (i.e. nonsense mutations not selected against) in a given replicate could be explained by Hap1 cells' reversion to diploidy prior to editing occurring, as noted by Haplogen (the cell line's source). Falsely deleterious edits in a given replicate could be observed due to off-target CRISPR cutting in other essential regions, or random dropout when half the sample is split on D5.
These findings suggest that while the technique is sensitive enough to measure effects from very few edited cells, noise associated with sampling such small populations mandates the necessity of replicating data sets to improve confidence in the measurements associated with individual genome edits. The data also suggest that increased reproducibility may be achievable by a) transfecting and analyzing a higher number of cells, b) limiting complexity of HDR libraries, or c) improving HDR efficiency to allow for sampling of more edited cells.
Potential Applications of Saturation Genome Editing
In the experiments disclosed herein, genotype is directly linked to phenotype to assay pools of multiplex HDR-derived variants. Targeted RNA and DNA sequencing of the edits themselves via selective PCR are well suited to catalog variants' effects on splicing and cellular fitness, respectively. However, with relatively simple adaptations of the method, complex pools of genome edits can be subjected to many additional assays that measure diverse aspects of biology.
First, the approach illustrated in the BRCA1 experiments is broadly applicable to study how genomic variation within virtually any transcribed element affects its own RNA abundance. Specifically, this approach could readily be adopted to study how other transcribed elements contribute to expression levels (e.g. the influence of 5′- and 3′ UTRs sequence on RNA stability, etc.). In this context, enhancers are transcribed at low levels (eRNA), suggesting an approach for studying enhancer activity, as well.
Additionally, assays such as targeted ChIP-Seq could be performed to characterize how libraries of genomic edits affect epigenetic states in coding or non-coding regions. By taking large quantities of DNA from expanded populations of edited cells and functionally separating edits based on biochemical interactions (i.e. transcription factor binding, associated histone modification, nucleosome positioning, etc.), genotype-phenotype associations would be preserved.
Apart from the molecular assays described above, the DBR1 experiment is just one example of a cell-based assay that can be read out with high-throughput sequencing. In addition to essentiality (in haploid cells or diploid cells made functionally haploid through previous gene disruption), gain-of-function (such as drug resistance or growth gain), haploid insufficiency and dominant negative effects could be measured with appropriate selection assays. In fact, any well-customized assay that allows functionally-based separation of cell populations (e.g., with FACS) is amenable to downstream sequencing of edited populations of assayed cells as a readout. For instance, reporter cell lines engineered to express fluorescently tagged genes of interest could be used to assay multiplex HDR-edited transcription factors or enhancers.
Given the relative ease of targeted nuclease production and mutagenesis library cloning, the methods disclosed herein are readily scalable. Exons could be tiled to functionally assess each coding SNV across entire genes. Therefore, the methods disclosed herein provide a valuable approach for determining functional effects of large numbers of programmed genomic mutations in many biological contexts.
The references, patents and published patent applications listed below, and all references cited in the specification above are hereby incorporated by reference in their entirety, as if fully set forth herein.
This application claims priority to U.S. Provisional Patent Application No. 62/032,734, filed Aug. 4, 2014, and U.S. Provisional Patent Application No. 62/046,074, filed Sep. 4, 2014, the subject matters of which are hereby incorporated by reference in their entireties as if fully set forth herein.
This invention was made with government support under Grant No. DP1HG007811, awarded by the National Institutes of Health. The government has certain rights in the invention.
| Number | Date | Country | |
|---|---|---|---|
| 62032734 | Aug 2014 | US | |
| 62046074 | Sep 2014 | US |