BIOSYNTHESIS OF CANNABINOIDS AND CANNABINOID PRECURSORS

REFERENCE TO AN ELECTRONIC SEQUENCE LISTING

The contents of the electronic sequence listing (G091970085WO00-SEQ-KVC.xml; Size: 345,085 bytes; and Date of Creation: Sep. 29, 2022) is herein incorporated by reference in its entirety.

FIELD OF INVENTION

The present disclosure relates to the biosynthesis of cannabinoids and cannabinoid precursors, such as in recombinant cells.

BACKGROUND

Cannabinoids are chemical compounds that may act as ligands for endocannabinoid receptors and have multiple medical applications. Traditionally, cannabinoids have been isolated from plants of the genus Cannabis. The use of plants for producing cannabinoids is inefficient, however, with isolated products often limited to the two most prevalent endogenous cannabinoids. THC and CBD, as other cannabinoids are typically produced in very low concentrations in Cannabis plants. Further, the cultivation of Cannabis plants is restricted in many jurisdictions. In addition, in order to obtain consistent results, Cannabis plants are often grown in a controlled environment, such as indoor grow rooms without windows, to provide flexibility in modulating growing conditions such as lighting, temperature, humidity, airflow, etc. Growing Cannabis plants in such controlled environments can result in high energy usage per gram of cannabinoid produced, especially for rare cannabinoids that the plants produce only in small amounts. For example, lighting in such grow rooms is provided by artificial sources, such as high-powered sodium lights. As many species of Cannabis have a vegetative cycle that requires 18 or more hours of light per day, powering such lights can result in significant energy expenditures. It has been estimated that between 0.88-1.34 kWh of energy is required to produce one gram of THC in dried Cannabis flower form (e.g., before any extraction or purification). Additionally, concern has been raised over agricultural practices in certain jurisdictions, such as California, w % here the growing season coincides with the dry season such that the water usage may impact connected surface water in streams (Dillis, Christopher, Connor McIntee, Van Butsic, Lance Le, Kason Grady, and Theodore Grantham. “Water storage and irrigation practices for cannabis drive seasonal patterns of water extraction and use in Northern California.” Journal of Environmental Management 272 (2020): 110955). See, also, Summers, H. M., Sproul, E. & Quinn, J. C. The greenhouse gas emissions of indoor cannabis production in the United States. Nat Sustain 4, 644-650 (2021).; and Zheng, Z., Fiddes, K. & Yang, L. A narrative review on environmental impacts of cannabis cultivation. J Cannabis Res 3, 35 (2021).

Cannabinoids can be produced through chemical synthesis (see, e.g., U.S. Pat. No. 7,323,576 to Souza et al). However, such methods suffer from low yields and high cost. Production of cannabinoids, cannabinoid analogs, and cannabinoid precursors using engineered organisms may provide an advantageous approach to meet the increasing demand for these compounds.

SUMMARY

Aspects of the present disclosure provide methods for production of cannabinoids and cannabinoid precursors from fatty acid substrates using genetically modified host cells.

Aspects of the present disclosure provide methods for producing a cannabinoid compound, comprising contacting olivetol and geranyl pyrophosphate with a prenyltransferase (PT), wherein the PT comprises an amino acid sequence that is at least 90% identical to the sequence of SEQ ID NO: 34.

Further aspects of the disclosure provide methods for producing a cannabinoid compound, comprising contacting 5-substituted resorcinol and a prenyl moiety with a prenyltransferase (PT), wherein the PT comprises an amino acid sequence that is at least 90% identical to the sequence of SEQ ID NO: 34.

Further aspects of the disclosure provide methods for producing a cannabinoid compound, comprising contacting 5-substituted 1,3-benzenediol and a prenyl moiety with a prenyltransferase (PT), wherein the PT comprises an amino acid sequence that is at least 90% identical to the sequence of SEQ ID NO: 34.

In some embodiments, the method occurs in vitro. In some embodiments, the method occurs within a host cell that expresses a heterologous polynucleotide encoding the PT.

Further aspects of the disclosure provide methods for producing a cannabinoid compound, comprising culturing a host cell in the presence of olivetol, wherein the host cell comprises a heterologous polynucleotide encoding a prenyltransferase (PT), wherein the PT comprises an amino acid sequence that is at least 90% identical to the sequence of SEQ ID NO: 34.

In some embodiments, the PT comprises the sequence of SEQ ID NO: 34 or a conservatively substituted version thereof. In some embodiments, the heterologous polynucleotide comprises a sequence that is at least 90% identical to the sequence of SEQ ID NO: 35. In some embodiments, the heterologous polynucleotide comprises the sequence of SEQ ID NO: 35.

In some embodiments, the heterologous polynucleotide is integrated into the genome of the host cell.

In some embodiments, the heterologous polynucleotide is expressed from a plasmid.

In some embodiments, the cannabinoid compound is CBG.

In some embodiments, the host cell produces at least 5, 10, 15, 20 or more than 20 fold more CBG than a host cell that expresses a heterologous polynucleotide encoding a PT that comprises the amino acid sequence of SEQ ID NO: 8. In some embodiments, the host cell produces at least 5, 10, 15, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500 or more than 500% more CBG than a host cell that expresses a heterologous polynucleotide encoding a PT that comprises the amino acid sequence of SEQ ID NO: 8. In some embodiments, the host cell produces at least 1000, 2000, 3000, 4000, 5000, 6000 or 7000 μg/L CBG.

In some embodiments, the host cell is capable of producing cannabichromene (CBC), tetrahydrocannabinol (THC) and/or cannabidiol (CBD). In some embodiments, the host cell comprises a heterologous polynucleotide encoding a terminal synthase (TS), wherein the TS comprises an amino acid sequence that is at least 90% identical to any one of SEQ ID NOs: 27, 38, 44, and 50. In some embodiments, the TS comprises the sequence of any one of SEQ ID NOs: 27, 38, 44, and 50. In some embodiments, the heterologous polynucleotide encoding the TS comprises a sequence that is at least 90% identical to any one of SEQ ID NOs: 28, 39, 45, and 51. In some embodiments, the heterologous polynucleotide comprises the sequence of any one of SEQ ID NOs: 28, 39, 45, and 51.

In some embodiments, the host cell is a plant cell, an algal cell, a yeast cell, a bacterial cell, or an animal cell. In some embodiments, the host cell is a yeast cell. In some embodiments, the yeast cell is a Saccharomyces cell, a Yarrowia cell, a Komagataella cell, or a Pichia cell. In some embodiments, the Saccharomyces cell is a Saccharomyces cerevisiae cell. In some embodiments, the yeast cell is Yarrowia cell. In some embodiments, the host cell is a bacterial cell. In some embodiments, the bacterial cell is an E. coli cell.

Further aspects of the disclosure provide methods for producing a cannabinoid compound, comprising contacting cannabigerol (CBG) with a terminal synthase (TS), wherein the TS comprises an amino acid sequence that is at least 90% identical to the sequence of any one of SEQ ID NOs: 27, 38, 44, and 50.

Further aspects of the disclosure provide methods for producing a cannabinoid compound, comprising contacting 5-substituted 2-prenyl-1,3-benzendiol with a terminal synthase (TS), wherein the TS comprises an amino acid sequence that is at least 90% identical to the sequence of any one of SEQ ID NOs: 27, 38, 44, and 50.

In some embodiments, the method occurs in vitro. In some embodiments, the method occurs within a host cell that expresses a heterologous polynucleotide encoding the TS.

Further aspects of the disclosure provide methods for producing a cannabinoid compound, comprising culturing a host cell in the presence of cannabigerol (CBG), wherein the host cell comprises a heterologous polynucleotide encoding a TS, wherein the TS comprises an amino acid sequence that is at least 90% identical to the sequence of any one of SEQ ID NOs: 27, 38, 44, and 50.

In some embodiments, the TS comprises the sequence of any one of SEQ ID NOs: 27, 38, 44, and 50, or a conservatively substituted version thereof. In some embodiments, the heterologous polynucleotide comprises a sequence that is at least 90% identical to the sequence of any one of SEQ ID NOs: 28, 39, 45, and 51. In some embodiments, the heterologous polynucleotide comprises the sequence of any one of SEQ ID NOs: 28, 39, 45, and 51.

In some embodiments, the heterologous polynucleotide is integrated into the genome of the host cell. In some embodiments, the heterologous polynucleotide is expressed from a plasmid.

In some embodiments, the cannabinoid compound is CBC. In some embodiments, the host cell is capable of producing at least 40.000 μg/L, at least 50,000 μg/L, at least 60,000 μg/L or at least 64,000 μg/L CBC.

In some embodiments, the cannabinoid compound is tetrahydrocannabinol (THC). In some embodiments, the host cell is capable of producing at least 1.500 μg/L, at least 2,000 μg/L or at least 2,500 μg/L THC.

In some embodiments, the cannabinoid compound is cannabidiol (CBD). In some embodiments, the host cell is capable of producing at least at least 500 μg/L, at least 750 μg/L or at least 1,000 μg/L CBD.

In some embodiments, the host cell further comprises one or more heterologous polynucleotides encoding one or more of: an acyl activating enzyme (AAE), a polyketide synthase (PKS), a polyketide cyclase (PKC), a bifunctional PKS-PKC, a prenyltransferase (PT) and/or a terminal synthase (TS). In some embodiments, the PKS is an olivetol synthase (OLS). In some embodiments, the PKS comprises a sequence that is at least 90% identical to the sequence of SEQ ID NO: 58. In some embodiments, the PKS comprises the sequence of SEQ ID NO: 58. In some embodiments, the PT comprises a sequence that is at least 90% identical to the sequence of SEQ ID NO: 34. In some embodiments, the PT comprises the sequence of SEQ ID NO: 34. In some embodiments, the heterologous polynucleotide encoding the PT comprises a sequence that is at least 90% identical to the sequence of SEQ ID NO: 35. In some embodiments, the heterologous polynucleotide encoding the PT comprises the sequence of SEQ ID NO: 35.

Further aspects of the disclosure provide compositions comprising olivetol and a heterologous polynucleotide encoding a prenyltransferase (PT), wherein the PT comprises an amino acid sequence that is at least 90% identical to the sequence of SEQ ID NO: 34, and wherein the PT is capable of utilizing olivetol as a substrate for producing cannabigerol (CBG). Further aspects of the disclosure provide host cells comprising such compositions, wherein the host cell is capable of producing cannabigerol (CBG).

Further aspects of the disclosure provide host cells that comprise olivetol and a heterologous polynucleotide encoding a prenyltransferase (PT), wherein the PT comprises an amino acid sequence that is at least 90% identical to the sequence of SEQ ID NO: 34, and wherein the host cell is capable of producing cannabigerol (CBG).

Further aspects of the disclosure provide host cells comprising 5-substituted 1,3-benzenediol and a heterologous polynucleotide encoding a prenyltransferase (PT), wherein the PT comprises an amino acid sequence that is at least 90% identical to the sequence of SEQ ID NO: 34, and wherein the PT is capable of utilizing 5-substituted 1,3-benzenediol as a substrate for producing cannabigerol (CBG).

Further aspects of the disclosure provide host cells comprising 5-alkyl-1,3-benzenediol and a heterologous polynucleotide encoding a prenyltransferase (PT), wherein the PT comprises an amino acid sequence that is at least 90% identical to the sequence of SEQ ID NO. 34, and wherein the PT is capable of utilizing 5-alkyl-1,3-benzenediol as a substrate for producing cannabigerol (CBG).

In some embodiments, the host cell produces at least 5, 10, 15, 20 or more than 20 fold more CBG than a control host cell, wherein the control host cell expresses a heterologous polynucleotide encoding a PT that comprises the amino acid sequence of SEQ ID NO: 8, and wherein the control host cell does not express a PT that comprises the sequence of SEQ ID NO: 34. In some embodiments, the host cell produces at least 5, 10, 15, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500 or more than 500% more CBG than a control host cell, wherein the control host cell expresses a heterologous polynucleotide encoding a PT that comprises the amino acid sequence of SEQ ID NO: 8, and wherein the control host cell does not express a PT that comprises the sequence of SEQ ID NO: 34. In some embodiments, the host cell produces at least 1000, 2000, 3000, 4000, 5000, 6000 or 7000 μg/L CBG.

Further aspects of the disclosure provide compositions comprising cannabigerol (CBG) and a heterologous polynucleotide encoding a terminal synthase (TS), wherein the TS is a fungal TS, and wherein TS is capable of producing cannabichromene (CBC).

Further aspects of the disclosure provide compositions comprising 5-substituted 2-prenyl-1,3-benzendiol and a heterologous polynucleotide encoding a terminal synthase (TS), wherein the TS is a fungal TS, and wherein TS is capable of producing cannabichromene (CBC).

Further aspects of the disclosure provide host cells comprising such compositions.

Further aspects of the disclosure provide compositions comprising cannabigerol (CBG) and a heterologous polynucleotide encoding a terminal synthase (TS), wherein the TS comprises an amino acid sequence that is at least 90% identical to the sequence of any one of SEQ ID NOs: 27, 38, 44, and 50, wherein the TS is capable of utilizing CBG as a substrate to produce a cannabinoid compound.

In some embodiments, the host cell is capable of producing a cannabinoid compound. In some embodiments, the TS comprises the sequence of any one of SEQ ID NOs: 27, 38, 44, and 50, or a conservatively substituted version thereof. In some embodiments, the heterologous polynucleotide comprises a sequence that is at least 90% identical to the sequence of any one of SEQ ID NOs: 28, 39, 45, and 51. In some embodiments, the heterologous polynucleotide comprises the sequence of any one of SEQ ID NOs: 28, 39, 45, and 51.

In some embodiments, the heterologous polynucleotide is integrated into the genome of the host cell. In some embodiments, the heterologous polynucleotide is expressed from a plasmid.

In some embodiments, the cannabinoid compound is CBC. In some embodiments, the host cell is capable of producing at least 40,000 μg/L, at least 50.000 μg/L, at least 60,000 μg/L or at least 64,000 μg/L CBC.

In some embodiments, the cannabinoid compound is tetrahydrocannabinol (THC). In some embodiments, the host cell is capable of producing at least 1,500 μg/L, at least 2,000 μg/L or at least 2,500 μg/L THC.

In some embodiments, the heterologous polynucleotide encoding the PT comprises a sequence that is at least 90% identical to the sequence of SEQ ID NO: 35. In some embodiments, the heterologous polynucleotide encoding the PT comprises the sequence of SEQ ID NO: 35.

Further aspects of the disclosure provide bioreactors for producing a cannabinoid compound, wherein the bioreactor contains olivetol and a prenyltransferase (PT) comprising an amino acid sequence that is at least 90% identical to the sequence of SEQ ID NO: 34.

Further aspects of the disclosure provide bioreactors for producing a cannabinoid compound, wherein the bioreactor contains CBG and a terminal synthase (TS) comprising a sequence that is at least 90% identical to the sequence of any one of SEQ ID NOs: 27, 38, 44, and 50.

Further aspects of the disclosure provide bioreactors for producing a cannabinoid compound, wherein the bioreactor contains: (i) a prenyltransferase (PT) comprising an amino acid sequence that is at least 90% identical to SEQ ID NO: 34, and (ii) a terminal synthase (TS) comprising a sequence that is at least 90% identical to the sequence of any one of SEQ ID NOs: 27, 38, 44, and 50.

In some embodiments, the cannabinoid compound is cannabigerol (CBG). In some embodiments, the cannabinoid compound is cannabichromene (CBC), tetrahydrocannabinol (THC) and/or cannabidiol (CBD).

Each of the limitations of the invention can encompass various embodiments of the invention. It is, therefore, anticipated that each of the limitations of the invention involving any one element or combinations of elements can be included in each aspect of the invention. This disclosure is not limited in its application to the details of construction and the arrangement of components set forth in the following description or illustrated in the drawings. The invention is capable of other embodiments and of being practiced or of being carried out in various ways. Also, the phraseology and terminology used in this application is for the purpose of description and should not be regarded as limiting. The use of “including,” “comprising,” or “having,” “containing,” “involving.” and variations thereof, is meant to encompass the items listed thereafter and equivalents thereof as well as additional items.

BRIEF DESCRIPTION OF DRAWINGS

The accompanying drawings are not intended to be drawn to scale. In the drawings, each identical or nearly identical component that is illustrated in various figures is represented by a like numeral. For purposes of clarity, not every component may be labeled in every drawing. In the drawings:

FIG. 1 is a schematic depicting the native Cannabis biosynthetic pathway for production of cannabinoid compounds, including five enzymatic steps mediated by: (R1a) acyl activating enzymes (AAE); (R2a) olivetol synthase enzymes (OLS); (R3a) olivetolic acid cyclase enzymes (OAC); (R4a) prenyltransferase enzymes (PT); and (R5a) terminal synthase enzymes (TS). Formulae 1a-11a correspond to hexanoic acid (1a), hexanoyl-CoA (2a), malonyl-CoA (3a), 3,5,7-trioxododecanoyl-CoA (4a), olivetol (5a), olivetolic acid (6a), geranyl pyrophosphate (7a), cannabigerolic acid (8a), cannabidiolic acid (9a), tetrahydrocannabinolic acid (10a), and cannabichromenic acid (11a). Hexanoic acid is an exemplary carboxylic acid substrate; other carboxylic acids may also be used (e.g., butyric acid, isovaleric acid, octanoic acid, decanoic acid, etc.; see e.g., FIG. 3 below). The enzymes that catalyze the synthesis of 3,5,7-trioxododecanoyl-CoA and olivetolic acid are shown in R2a and R3a, respectively, and can include multi-functional enzymes that catalyze the synthesis of 3,5,7-trioxododecanoyl-CoA and olivetolic acid. The enzymes cannabidiolic acid synthase (CBDAS), tetrahydrocannabinolic acid synthase (THCAS), and cannabichromenic acid synthase (CBCAS) that catalyze the synthesis of cannabidiolic acid, tetrahydrocannabinolic acid, and cannabichromenic acid, respectively, are shown in step R5a. FIG. 1 is adapted from Carvalho et al. “Designing Microorganisms for Heterologous Biosynthesis of Cannabinoids” (2017) FEMS Yeast Research June 1; 17(4), which is incorporated by reference in its entirety.

FIG. 2 is a schematic depicting a heterologous biosynthetic pathway for production of cannabinoid compounds, including five enzymatic steps mediated by: (R1) acyl activating enzymes (AAE); (R2) polyketide synthase enzymes (PKS) or bifunctional polyketide synthase-polyketide cyclase enzymes (PKS-PKC); (R3) polyketide cyclase enzymes (PKC) or bifunctional PKS-PKC enzymes: (R4) prenyltransferase enzymes (PT); and (R5) terminal synthase enzymes (TS). Any carboxylic acid of varying chain lengths, structures (e.g., aliphatic, alicyclic, or aromatic) and functionalization (e.g., hydroxylic-, keto-, amino-, thiol-, aryl-, or alogeno-) may also be used as precursor substrates (e.g., thiopropionic acid, hydroxy phenyl acetic acid, norleucine, bromodecanoic acid, butyric acid, isovaleric acid, octanoic acid, decanoic acid, etc).

FIG. 3 is a non-exclusive representation of select putative precursors for the cannabinoid pathway in FIG. 2.

FIG. 4 is a schematic depicting the biosynthetic pathway for production of varin cannabinoid compounds, including five enzymatic steps mediated by: (R1) acyl activating enzymes (AAE); (R2) polyketide synthase enzymes (PKS) or bifunctional polyketide synthase-polyketide cyclase enzymes (PKS-PKC); (R3) polyketide cyclase enzymes (PKC) or bifunctional PKS-PKC enzymes; (R4) prenyltransferase enzymes (PT); and (R5) terminal synthase enzymes (TS). The compounds of Formulae 1b-11b correspond to butyric acid (1b), butyroyl-CoA (2b), malonyl-CoA (3b), 3,5,7-trioxodecanoyl-CoA (4b), divarinol (5b), divaric acid (6b), geranyl pyrophosphate (7b), cannabigerovarinic acid (8b), cannabidivarinic acid (9b), tetrahydrocannabivarinic acid (10b), and cannabichromevarinic acid (11b). Butyric acid is an exemplary carboxylic acid substrate; other carboxylic acids may also be used (e.g., hexanoic acid, isovaleric acid, octanoic acid, decanoic acid, etc.: see e.g., FIG. 3 above). The enzymes that catalyze the synthesis of 3,5,7-trioxodecanoyl-CoA and divaric acid are shown in R2 and R3, respectively, and can include multi-functional enzymes that catalyze the synthesis of 3,5,7-trioxodecanoyl-CoA and divaric acid. The enzymes cannabigerovarinic acid synthase (CBGVAS), tetrahydrocannabivarinic acid synthase (THCVAS), and cannabichromevarinic acid synthase (CBCVAS) that catalyze the synthesis of the varinolic cannabinoids cannabigerovarinic acid, tetrahydrocannabivarinic acid, and cannabichromevarinic acid, respectively, and their catalytic function thereof, are shown in step R5.

FIGS. 5A-5B are schematics showing reactions catalyzed by a PT enzyme. FIG. 5A is a schematic showing a reaction wherein olivetolic acid (OA, Formula (6a)) and geranyl pyrophosphate (GPP, Formula (7a)) are condensed to form either the major cannabinoid cannabigerolic acid (CBGA, Formula (8a)) or 2-O-geranyl olivetolic acid (OGOA, Formula (8b)). FIG. 5B is a schematic showing a reaction wherein a prenyl group is added to olivetol (OL, Formula (5a)) to form the cannabinoid cannabigerol (CBG. Formula 8a-1).

FIGS. 6A-6B are schematics showing reactions catalyzed by a TS enzyme. FIG. 6A is a schematic showing a reaction wherein the geranyl moiety of cannabigerolic acid (Formula (8a)) is cyclized to yield cannabidiolic acid, tetrahydrocannabinolic acid, or cannabichromenic acid. FIG. 6B is a schematic showing a reaction wherein the geranyl moiety of cannabigerol (Formula (8a-1)) is cyclized to yield cannabidiol, tetrahydrocannabinol, or cannabichromene.

FIG. 7 is a schematic showing a plasmid used to express candidate PT enzymes in S. cerevisiae described in Example 1. The coding sequence for the candidate PT enzymes (labeled “Library gene”) was driven by the GAL1 promoter. The plasmid contains markers for both yeast (URA3) and bacteria (ampR), as well as origins of replication for yeast (2 micron), and bacteria (pBR322).

FIG. 8 is a schematic showing a plasmid used to express TS enzymes in S. cerevisiae described in Example 2. The coding sequence for the TS enzymes (labeled “Library gene”) was driven by the GAL1 promoter.

FIGS. 9A-9B depict graphs showing activity data of a PT enzyme identified in Example 1, expressed in strain t913655, for cannabigerol (CBG) production based on an in vivo activity assay in S. cerevisiae. FIG. 9A depicts olivetol utilization and FIG. 9B depicts cannabigerol (CBG) production. Strain t935014, expressing GFP, was used as a negative control. Strain t914495, comprising NphB from Streptomyces sp. (SEQ ID NO: 8), was included in the library as a positive control for enzyme activity. The data represent the average of four bioreplicates±one standard deviation of the mean.

FIGS. 10A-10B depict graphs showing activity data of a PT enzyme identified in Example 1, expressed in strain t913655, for cannabigerovarin (CBGV) production based on an in vivo activity assay in S. cerevisiae. FIG. 10A depicts divarinol utilization and FIG. 10B depicts cannabigerovarin (CBGV) production. Strain t935014, expressing GFP, was used as a negative control. Strain t914495, comprising NphB from Streptomyces sp. (SEQ ID NO: 8), was included in the library as a positive control for enzyme activity. The data represent the average of four bioreplicates±one standard deviation of the mean.

FIGS. 11A-11B depict MS/MS data for a CBG standard (FIG. 11A) and for products produced by the PT enzyme expressed in strain t913655, identified in Example 1 (referred to in FIG. 11B as “A0A132B7I1”).

FIGS. 12A-12B depict graphs showing screening data of TSs for cannabichromene (CBC) production based on an in vivo activity assay in S. cerevisiae as described in Example 2. Strain t865842, expressing GFP, was used as a negative control. Strain t876606, expressing a C. sativa THCAS, and strain t876607, expressing C. sativa CBDAS, were used as positive controls for enzyme activity. Both the C. sativa THCAS and C. sativa CBDAS enzymes were expressed with an N-terminally fused MFα2 signal peptide and a C-terminally fused HDEL signal peptide. FIG. 12A depicts utilization of CBG as a substrate and FIG. 12B depicts CBC production. Four library strains expressing TSs were found to utilize CBG to produce CBC: strain t870557, which comprises a CBCAS from Aspergillus vadensis (corresponding to UniProt Accession No. A0A319B6X5, the protein sequence for which is provided as SEQ ID NO: 38): strain t870559, which comprises a CBCAS from Aspergillus awamori (corresponding to UniProt Accession No. A0A401KY63, the protein sequence for which is provided by SEQ ID NO: 44); strain 1878476, which comprises a CBCAS from Aspergillus lacticoffeatus (corresponding to UniProt Accession No. A0A319AGI5, the protein sequence for which is provided by SEQ ID NO: 50); and strain t887304, which comprises a CBCAS from Aspergillus niger (corresponding to UniProt Accession No. A0A254UC34, the protein sequence for which is provided by SEQ ID NO: 27). Strains depicted in FIGS. 12A-12B and their corresponding activity are shown in Table 7.

FIGS. 13A-13B depict graphs showing production of tetrahydrocannabinol (THC) and cannabidiol (CBD) by the strains described above in FIG. 12 based on an in vivo activity assay in S. cerevisiae as described in Example 2. Strain t865842, expressing GFP, was used as a negative control. Strain t876606, expressing a C. sativa THCAS, and strain t876607, expressing C. sativa CBDAS, were used as positive controls for enzyme activity. Strains 1870557, t870559, t878476 and t887304 were observed to produce THC (FIG. 13A) and CBD (FIG. 13B). Strains depicted in FIGS. 13A-13B and their corresponding activity are shown in Table 7.

DETAILED DESCRIPTION

This disclosure provides methods for production of cannabinoids and cannabinoid precursors from fatty acid substrates using genetically modified host cells. Methods include heterologous expression of a prenyltransferase (PT) and/or a terminal synthase (TS), such as a cannabichromenic acid synthase (CBCAS). The application describes PTs and TSs that can be functionally expressed in host cells such as S. cerevisiae. As demonstrated in the Examples, a PT was identified that is capable of using olivetol as a substrate to produce cannabigerol (CBG) in a host cell. As further demonstrated in the Examples, multiple non-Cannabis CBCASs were identified that were capable of utilizing CBG to produce the cannabinoid cannabichromene (CBC) in a host cell, as well as other cannabinoids such as THC and CBD. The PT and TSs provided in this disclosure may provide several advantages in the biosynthesis of cannabinoids over native Cannabis enzymes; for example, the enzymatic prenylation of olivetol to produce CBG provides a route to the valorization of an otherwise unused by-product of the cannabinoid pathway and/or the reduction of toxicity to a host cell performing such biosynthesis.

Definitions

While the following terms are believed to be well understood by one of ordinary skill in the art, the following definitions are set forth to facilitate explanation of the disclosed subject matter.

The term “a” or “an” refers to one or more of an entity, i.e., can identify a referent as plural. Thus, the terms “a” or “an,” “one or more” and “at least one” are used interchangeably in this application. In addition, reference to “an element” by the indefinite article “a” or “an” does not exclude the possibility that more than one of the elements is present, unless the context clearly requires that there is one and only one of the elements.

The terms “microorganism” or “microbe” should be taken broadly. These terms are used interchangeably and include, but are not limited to, the two prokaryotic domains. Bacteria and Archaea, as well as certain eukaryotic fungi and protists. In some embodiments, the disclosure may refer to the “microorganisms” or “microbes” of lists/tables and figures present in the disclosure. This characterization can refer to not only the identified taxonomic genera of the tables and figures, but also the identified taxonomic species, as well as the various novel and newly identified or designed strains of any organism in the tables or figures. The same characterization holds true for the recitation of these terms in other parts of the specification, such as in the Examples.

The term “prokaryotes” is recognized in the art and refers to cells that contain no nucleus or other cell organelles. The prokaryotes are generally classified in one of two domains, the Bacteria and the Archaea.

“Bacteria” or “eubacteria” refers to a domain of prokaryotic organisms. Bacteria include at least 11 distinct groups as follows: (1) Gram-positive (gram+) bacteria, of which there are two major subdivisions: (a) high G+C group (Actinomycetes, Mycobacteria, Micrococcus, others) and (b) low G+C group (Bacillus, Clostridia, Lactobacillus, Staphylococci, Streptococci, Mycoplasmas); (2) Proteobacteria, e.g., Purple photosynthetic+non-photosynthetic Gram-negative bacteria (includes most “common” Gram-negative bacteria); (3) Cyanobacteria, e.g., oxygenic phototrophs: (4) Spirochetes and related species: (5) Planctomyces; (6) Bacteroides, Flavobacteria; (7) Chlamydia; (8) Green sulfur bacteria; (9) Green non-sulfur bacteria (also anaerobic phototrophs): (10) Radioresistant micrococci and relatives; and (11) Thermotoga and Thermosipho thermophiles.

The term “Archaea” refers to a taxonomic classification of prokaryotic organisms with certain properties that make them distinct from Bacteria in physiology and phylogeny.

The term “Cannabis” refers to a genus in the family Cannabaceae. Cannabis is a dioecious plant. Glandular structures located on female flowers of Cannabis, called trichomes, accumulate relatively high amounts of a class of terpeno-phenolic compounds known as phytocannabinoids (described in further detail below). Cannabis has conventionally been cultivated for production of fibre and seed (commonly referred to as “hemp-type”), or for production of intoxicants (commonly referred to as “drug-type”). In drug-type Cannabis, the trichomes contain relatively high amounts of tetrahydrocannabinolic acid (THCA), which can convert to tetrahydrocannabinol (THC) via a decarboxylation reaction, for example upon combustion of dried Cannabis flowers, to provide an intoxicating effect. Drug-type Cannabis often contains other cannabinoids in lesser amounts. In contrast, hemp-type Cannabis contains relatively low concentrations of THCA, often less than 0.3% THC by dry weight. Hemp-type Cannabis may contain non-THC and non-THCA cannabinoids, such as cannabidiolic acid (CBDA), cannabidiol (CBD), and other cannabinoids. Presently, there is a lack of consensus regarding the taxonomic organization of the species within the genus. Unless context dictates otherwise, the term “Cannabis” is intended to include all putative species within the genus, such as, without limitation, Cannabis sativa, Cannabis indica, and Cannabis ruderalis and without regard to whether the Cannabis is hemp-type or drug-type.

The term “cyclase activity” in reference to a polyketide synthase (PKS) enzyme (e.g., an olivetol synthase (OLS) enzyme) or a polyketide cyclase (PKC) enzyme (e.g., an olivetolic acid cyclase (OAC) enzyme), refers to the activity of catalyzing the cyclization of an oxo fatty acyl-CoA (e.g. 3,5,7-trioxododecanoyl-COA, 3,5,7-trioxodecanoyl-COA) to the corresponding intramolecular cyclization product (e.g., olivetolic acid, divarinic acid). In some embodiments, the PKS or PKC catalyzes the C2-C7 aldol condensation of an acyl-COA with three additional ketide moieties added thereto.

A “cytosolic” or “soluble” enzyme refers to an enzyme that is predominantly localized (or predicted to be localized) in the cytosol of a host cell.

A “eukaryote” is any organism whose cells contain a nucleus and other organelles enclosed within membranes. Eukaryotes belong to the taxon Eukarya or Eukaryota. The defining feature that sets eukaryotic cells apart from prokaryotic cells (i.e., bacteria and archaea) is that they have membrane-bound organelles, especially the nucleus, which contains the genetic material, and is enclosed by the nuclear envelope.

The term “host cell” refers to a cell that can be used to express a polynucleotide, such as a polynucleotide that encodes an enzyme used in biosynthesis of cannabinoids or cannabinoid precursors. The terms “genetically modified host cell,” “recombinant host cell,” and “recombinant strain” are used interchangeably and refer to host cells that have been genetically modified by, e.g., cloning and transformation methods, or by other methods known in the art (e.g., selective editing methods, such as CRISPR). Thus, the terms include a host cell (e.g., bacterial cell, yeast cell, fungal cell, insect cell, plant cell, mammalian cell, human cell, etc.) that has been genetically altered, modified, or engineered, so that it exhibits an altered, modified, or different genotype and/or phenotype, as compared to the naturally-occurring cell from which it was derived. It is understood that in some embodiments, the terms refer not only to the particular recombinant host cell in question, but also to the progeny or potential progeny of such a host cell.

The term “control host cell,” or the term “control” when used in relation to a host cell, refers to an appropriate comparator host cell for determining the effect of a genetic modification or experimental treatment. In some embodiments, the control host cell is a wild type cell. In other embodiments, a control host cell is genetically identical to the genetically modified host cell, except for the genetic modification(s) differentiating the genetically modified or experimental treatment host cell. In some embodiments, the control host cell has been genetically modified to express a wild type or otherwise known variant of an enzyme being tested for activity in other test host cells.

The term “heterologous” with respect to a polynucleotide, such as a polynucleotide comprising a gene, is used interchangeably with the term “exogenous” and the term “recombinant” and refers to: a polynucleotide that has been artificially supplied to a biological system: a polynucleotide that has been modified within a biological system, or a polynucleotide whose expression or regulation has been manipulated within a biological system. A heterologous polynucleotide that is introduced into or expressed in a host cell may be a polynucleotide that comes from a different organism or species from the host cell, or may be a synthetic polynucleotide, or may be a polynucleotide that is also endogenously expressed in the same organism or species as the host cell. For example, a polynucleotide that is endogenously expressed in a host cell may be considered heterologous when it is situated non-naturally in the host cell: expressed recombinantly in the host cell, either stably or transiently; modified within the host cell; selectively edited within the host cell; expressed in a copy number that differs from the naturally occurring copy number within the host cell; or expressed in a non-natural way within the host cell, such as by manipulating regulatory regions that control expression of the polynucleotide. In some embodiments, a heterologous polynucleotide is a polynucleotide that is endogenously expressed in a host cell but whose expression is driven by a promoter that does not naturally regulate expression of the polynucleotide. In other embodiments, a heterologous polynucleotide is a polynucleotide that is endogenously expressed in a host cell and whose expression is driven by a promoter that does naturally regulate expression of the polynucleotide, but the promoter or another regulatory region is modified. In some embodiments, the promoter is recombinantly activated or repressed. For example, gene-editing based techniques may be used to regulate expression of a polynucleotide, including an endogenous polynucleotide, from a promoter, including an endogenous promoter. See, e.g., Chavez et al., Nat Methods. 2016 July: 13(7): 563-567. A heterologous polynucleotide may comprise a wild-type sequence or a mutant sequence as compared with a reference polynucleotide sequence.

The term “at least a portion” or “at least a fragment” of a nucleic acid or polypeptide means a portion having the minimal size characteristics of such sequences, or any larger fragment of the full length molecule, up to and including the full length molecule. A fragment of a polynucleotide of the disclosure may encode a biologically active portion of an enzyme, such as a catalytic domain. A biologically active portion of a genetic regulatory element may comprise a portion or fragment of a full length genetic regulatory element and have the same type of activity as the full length genetic regulatory element, although the level of activity of the biologically active portion of the genetic regulatory element may vary compared to the level of activity of the full length genetic regulatory element.

A coding sequence and a regulatory sequence are said to be “operably joined” or “operably linked” when the coding sequence and the regulatory sequence are covalently linked and the expression or transcription of the coding sequence is under the influence or control of the regulatory sequence. If the coding sequence is to be translated into a functional protein, the coding sequence and the regulatory sequence are said to be operably joined if induction of a promoter in the 5′ regulatory sequence promotes transcription of the coding sequence and if the nature of the linkage between the coding sequence and the regulatory sequence does not (1) result in the introduction of a frame-shift mutation, (2) interfere with the ability of the promoter region to direct the transcription of the coding sequence, or (3) interfere with the ability of the corresponding RNA transcript to be translated into a protein.

The terms “link,” “linked,” or “linkage” means two entities (e.g., two polynucleotides or two proteins) are bound to one another by any physicochemical means. Any linkage known to those of ordinary skill in the art, covalent or non-covalent, is embraced. In some embodiments, a nucleic acid sequence encoding an enzyme of the disclosure is linked to a nucleic acid encoding a signal peptide. In some embodiments, an enzyme of the disclosure is linked to a signal peptide. Linkage can be direct or indirect.

The terms “transformed” or “transform” with respect to a host cell refer to a host cell in which one or more nucleic acids have been introduced, for example on a plasmid or vector or by integration into the genome. In some instances where one or more nucleic acids are introduced into a host cell on a plasmid or vector, one or more of the nucleic acids, or fragments thereof, may be retained in the cell, such as by integration into the genome of the cell, while the plasmid or vector itself may be removed from the cell. In such instances, the host cell is considered to be transformed with the nucleic acids that were introduced into the cell regardless of whether the plasmid or vector is retained in the cell or not.

The term “volumetric productivity” or “production rate” refers to the amount of product formed per volume of medium per unit of time. Volumetric productivity can be reported in gram per liter per hour (g/L/h).

The term “specific productivity” of a product refers to the rate of formation of the product normalized by unit volume or mass or biomass and has the physical dimension of a quantity of substance per unit time per unit mass or volume [M·T⁻¹·M⁻¹or M·T⁻¹·L⁻³, where M is mass or moles, T is time, L is length].

The term “biomass specific productivity” refers to the specific productivity in gram product per gram of cell dry weight (CDW) per hour (g/g CDW/h) or in mmol of product per gram of cell dry weight (CDW) per hour (mmol/g CDW/h). Using the relation of CDW to OD600 for the given microorganism, specific productivity can also be expressed as gram product per liter culture medium per optical density of the culture broth at 600 nm (OD) per hour (g/L/h/OD). Also, if the elemental composition of the biomass is known, biomass specific productivity can be expressed in mmol of product per C-mole (carbon mole) of biomass per hour (mmol/C-mol/h).

The term “yield” refers to the amount of product obtained per unit weight of a certain substrate and may be expressed as g product per g substrate (g/g) or moles of product per mole of substrate (mol/mol). Yield may also be expressed as a percentage of the theoretical yield. “Theoretical yield” is defined as the maximum amount of product that can be generated per a given amount of substrate as dictated by the stoichiometry of the metabolic pathway used to make the product and may be expressed as g product per g substrate (g/g) or moles of product per mole of substrate (mol/mol).

The term “titer” refers to the strength of a solution or the concentration of a substance in solution. For example, the titer of a product of interest (e.g., small molecule, peptide, synthetic compound, fuel, alcohol, etc.) in a fermentation broth is described as g of product of interest in solution per liter of fermentation broth or cell-free broth (g/L) or as g of product of interest in solution per kg of fermentation broth or cell-free broth (g/Kg).

The term “total titer” refers to the sum of all products of interest produced in a process, including but not limited to the products of interest in solution, the products of interest in gas phase if applicable, and any products of interest removed from the process and recovered relative to the initial volume in the process or the operating volume in the process. For example, the total titer of products of interest (e.g., small molecule, peptide, synthetic compound, fuel, alcohol, etc.) in a fermentation broth is described as g of products of interest in solution per liter of fermentation broth or cell-free broth (g/L) or as g of products of interest in solution per kg of fermentation broth or cell-free broth (g/Kg).

The term “amino acid” refers to organic compounds that comprise an amino group, —NH2, and a carboxyl group, —COOH. The term “amino acid” includes both naturally occurring and unnatural amino acids. Nomenclature for the twenty common amino acids is as follows: alanine (ala or A), arginine (arg or R); asparagine (asn or N); aspartic acid (asp or D); cysteine (cys or C); glutamine (gln or Q); glutamic acid (glu or E); glycine (gly or G); histidine (his or H); isoleucine (ile or I); leucine (leu or L); lysine (lys or K); methionine (met or M); phenylalanine (phe or F); proline (pro or P); serine (ser or S); threonine (thr or T); tryptophan (trp or W); tyrosine (tyr or Y); and valine (val or V). Non-limiting examples of unnatural amino acids include homo-amino acids, proline and pyruvic acid derivatives, 3-substituted alanine derivatives, glycine derivatives, ring-substituted phenylalanine derivatives, ring-substituted tyrosine derivatives, linear core amino acids, amino acids with protecting groups including Fmoc, Boc, and Cbz, β-amino acids (β3 and β2), and N-methyl amino acids.

The term “aliphatic” refers to alkyl, alkenyl, alkynyl, and carbocyclic groups. Likewise, the term “heteroaliphatic” refers to heteroalkyl, heteroalkenyl, heteroalkynyl, and heterocyclic groups.

The term “alkyl” refers to a radical of, or a substituent that is, a straight-chain or branched saturated hydrocarbon group having from 1 to 20 carbon atoms (“C1-20 alkyl”). In certain embodiments, the term “alkyl” refers to a radical of, or a substituent that is, a straight-chain or branched saturated hydrocarbon group having from 1 to 10 carbon atoms (“C_1-10alkyl”). In some embodiments, an alkyl group has 1 to 9 carbon atoms (“C_1-9alkyl”). In some embodiments, an alkyl group has 1 to 8 carbon atoms (“C_1-8alkyl”). In some embodiments, an alkyl group has 1 to 7 carbon atoms (“C_1-7alkyl”). In some embodiments, an alkyl group has 2 to 7 carbon atoms (“C2-7 alkyl”). In some embodiments, an alkyl group has 3 to 7 carbon atoms (“C3-7 alkyl”). In some embodiments, an alkyl group has 1 to 6 carbon atoms (“C_1-6alkyl”). In some embodiments, an alkyl group has 2 to 6 carbon atoms (“C_2-6alkyl”). In some embodiments, an alkyl group has 3 to 5 carbon atoms (“C3-s alkyl”). In some embodiments, an alkyl group has 5 carbon atoms (“C₅alkyl”). In some embodiments, the alkyl group has 3 carbon atoms (“C3 alkyl”). In some embodiments, the alkyl group has 7 carbon atoms (“C7 alkyl”). In some embodiments, an alkyl group has 1 to 5 carbon atoms (“C_1-5alkyl”). In some embodiments, an alkyl group has 1 to 4 carbon atoms (“C_1-4alkyl”). In some embodiments, an alkyl group has 1 to 3 carbon atoms (“C_1-3alkyl”). In some embodiments, an alkyl group has 1 to 2 carbon atoms (“C_1-2alkyl”). In some embodiments, an alkyl group has 1 carbon atom (“C₁alkyl”).

Examples of C_1-6alkyl groups include methyl (C₁), ethyl (C₂), propyl (C₃) (e.g., n-propyl, isopropyl), butyl (C₄) (e.g., n-butyl, tert-butyl, sec-butyl, iso-butyl), pentyl (C₅) (e.g., n-pentyl, 3-pentanyl, amyl, neopentyl, 3-methyl-2-butanyl, tertiary amyl), and hexyl (C₆) (e.g., n-hexyl). Additional examples of alkyl groups include n-heptyl (C₇), n-octyl (C₈), and the like. Unless otherwise specified, each instance of an alkyl group is independently unsubstituted (an “unsubstituted alkyl”) or substituted (a“substituted alkyl”) with one or more substituents (e.g., halogen, such as F). In certain embodiments, the alkyl group is an unsubstituted C_1-10alkyl (such as unsubstituted C_1-6alkyl, e.g., —CH₃(Me), unsubstituted ethyl (Et), unsubstituted propyl (Pr, e.g., unsubstituted n-propyl (n-Pr), unsubstituted isopropyl (i-Pr)), unsubstituted butyl (Bu, e.g., unsubstituted n-butyl (n-Bu), unsubstituted tert-butyl (tert-Bu or t-Bu), unsubstituted sec-butyl (sec-Bu), unsubstituted isobutyl (i-Bu)). In certain embodiments, the alkyl group is a substituted C_1-10alkyl (such as substituted C_1-6alkyl, e.g., —CF₃, benzyl).

The term “acyl” refers to a group having the general formula —C(═O)R^X1, —C(═O)OR^X1, —C(═O)—O—C(═O)R^X1, —C(═O)SR^X1, —C(═O)N(R^X1)₂, —C(═S)R^X1, —C(═S)N(R^X1)₂, and —C(═S)S(R^X1), —C(═NR^X1)R^X1, —C(═NR^X1)OR^X1, —C(═NR^X1)SR^X1, and —C(═NR^X1)N(R^X1)₂, wherein R^X1is hydrogen; halogen; substituted or unsubstituted hydroxyl; substituted or unsubstituted thiol; substituted or unsubstituted amino; substituted or unsubstituted acyl, cyclic or acyclic, substituted or unsubstituted, branched or unbranched aliphatic; cyclic or acyclic, substituted or unsubstituted, branched or unbranched heteroaliphatic; cyclic or acyclic, substituted or unsubstituted, branched or unbranched alkyl; cyclic or acyclic, substituted or unsubstituted, branched or unbranched alkenyl; substituted or unsubstituted alkynyl; substituted or unsubstituted aryl, substituted or unsubstituted heteroaryl, aliphaticoxy, heteroaliphaticoxy, alkyloxy, heteroalkyloxy, aryloxy, heteroaryloxy, aliphaticthioxy, heteroaliphaticthioxy, alkylthioxy, heteroalkylthioxy, arylthioxy, heteroarylthioxy, mono- or di-aliphaticamino, mono- or di-heteroaliphaticamino, mono- or di-alkylamino, mono- or di-heteroalkylamino, mono- or di-arylamino, or mono- or di-heteroarylamino; or two R^X1groups taken together form a 5- to 6-membered heterocyclic ring. Exemplary acyl groups include aldehydes (—CHO), carboxylic acids (—CO₂H), ketones, acyl halides, esters, amides, imines, carbonates, carbamates, and ureas. Acyl substituents include, but are not limited to, any of the substituents described in this application that result in the formation of a stable moiety (e.g., aliphatic, alkyl, alkenyl, alkynyl, heteroaliphatic, heterocyclic, aryl, heteroaryl, acyl, oxo, imino, thiooxo, cyano, isocyano, amino, azido, nitro, hydroxyl, thiol, halo, aliphaticamino, heteroaliphaticamino, alkylamino, heteroalkylamino, arylamino, heteroarylamino, alkylaryl, arylalkyl, aliphaticoxy, heteroaliphaticoxy, alkyloxy, heteroalkyloxy, aryloxy, heteroaryloxy, aliphaticthioxy, heteroaliphaticthioxy, alkylthioxy, heteroalkylthioxy, arylthioxy, heteroarylthioxy, acyloxy, and the like, each of which may or may not be further substituted).

“Alkenyl” refers to a radical of, or a substituent that is, a straight-chain or branched hydrocarbon group having from 2 to 20 carbon atoms, one or more carbon-carbon double bonds, and no triple bonds (“C_2-20alkenyl”). In some embodiments, an alkenyl group has 2 to 10 carbon atoms (“C_2-10alkenyl”). In some embodiments, an alkenyl group has 2 to 9 carbon atoms (“C_2-9alkenyl”). In some embodiments, an alkenyl group has 2 to 8 carbon atoms (“C_2-8alkenyl”). In some embodiments, an alkenyl group has 2 to 7 carbon atoms (“C_2-7alkenyl”). In some embodiments, an alkenyl group has 2 to 6 carbon atoms (“C_2-6alkenyl”). In some embodiments, an alkenyl group has 2 to 5 carbon atoms (“C_2-5alkenyl”). In some embodiments, an alkenyl group has 2 to 4 carbon atoms (“C_2-4alkenyl”). In some embodiments, an alkenyl group has 2 to 3 carbon atoms (“C_2-3alkenyl”). In some embodiments, an alkenyl group has 2 carbon atoms (“C₂alkenyl”). The one or more carbon-carbon double bonds can be internal (such as in 2-butenyl) or terminal (such as in 1-butenyl). Examples of C_2-4alkenyl groups include ethenyl (C₂), 1-propenyl (C₃), 2-propenyl (C₃), 1-butenyl (C₄), 2-butenyl (C₄), butadienyl (C₄), and the like. Examples of C_2-6alkenyl groups include the aforementioned C_2-4alkenyl groups as well as pentenyl (C₅), pentadienyl (C₅), hexenyl (C₆), and the like. Additional examples of alkenyl include heptenyl (C₇), octenyl (C₈), octatrienyl (C₈), and the like. Unless otherwise specified, each instance of an alkenyl group is independently optionally substituted, i.e., unsubstituted (an “unsubstituted alkenyl”) or substituted (a “substituted alkenyl”) with one or more substituents. In certain embodiments, the alkenyl group is unsubstituted C_2-10alkenyl. In certain embodiments, the alkenyl group is substituted C_2-10alkenyl.

“Alkynyl” refers to a radical of, or a substituent that is, a straight-chain or branched hydrocarbon group having from 2 to 20 carbon atoms, one or more carbon-carbon triple bonds, and optionally one or more double bonds (“C_2-20alkynyl”). In some embodiments, an alkynyl group has 2 to 10 carbon atoms (“C_2-10alkynyl”). In some embodiments, an alkynyl group has 2 to 9 carbon atoms (“C_2-9alkynyl”). In some embodiments, an alkynyl group has 2 to 8 carbon atoms (“C_2-8alkynyl”). In some embodiments, an alkynyl group has 2 to 7 carbon atoms (“C_2-7alkynyl”). In some embodiments, an alkynyl group has 2 to 6 carbon atoms (“C_2-6alkynyl”). In some embodiments, an alkynyl group has 2 to 5 carbon atoms (“C_2-5alkynyl”). In some embodiments, an alkynyl group has 2 to 4 carbon atoms (“C_2-4alkynyl”). In some embodiments, an alkynyl group has 2 to 3 carbon atoms (“C_2-3alkynyl”). In some embodiments, an alkynyl group has 2 carbon atoms (“C₂alkynyl”). The one or more carbon-carbon triple bonds can be internal (such as in 2-butynyl) or terminal (such as in 1-butynyl). Examples of C_2-4alkynyl groups include, without limitation, ethynyl (C₂), 1-propynyl (C₃), 2-propynyl (C₃), 1-butynyl (C₄), 2-butynyl (C₄), and the like. Examples of C_2-6alkenyl groups include the aforementioned C_2-4alkynyl groups as well as pentynyl (C₅), hexynyl (C₆), and the like. Additional examples of alkynyl include heptynyl (C₇), octynyl (C₈), and the like. Unless otherwise specified, each instance of an alkynyl group is independently optionally substituted, i.e., unsubstituted (an “unsubstituted alkynyl”) or substituted (a “substituted alkynyl”) with one or more substituents. In certain embodiments, the alkynyl group is unsubstituted C_2-10(alkynyl. In certain embodiments, the alkynyl group is substituted C_2-10alkynyl.

“Carbocyclyl” or “carbocyclic” refers to a radical of a non-aromatic cyclic hydrocarbon group having from 3 to 10 ring carbon atoms (“C_3-10carbocyclyl”) and zero heteroatoms in the non-aromatic ring system. In some embodiments, a carbocyclyl group has 3 to 8 ring carbon atoms (“C_3-8carbocyclyl”). In some embodiments, a carbocyclyl group has 3 to 6 ring carbon atoms (“C_3-6carbocyclyl”). In some embodiments, a carbocyclyl group has 3 to 6 ring carbon atoms (“C_3-6carbocyclyl”). In some embodiments, a carbocyclyl group has 5 to 10 ring carbon atoms (“C_5-10carbocyclyl”). Exemplary C_3-6carbocyclyl groups include, without limitation, cyclopropyl (C₃), cyclopropenyl (C₃), cyclobutyl (C₄), cyclobutenyl (C₄), cyclopentyl (C₅), cyclopentenyl (C₅), cyclohexyl (C₆), cyclohexenyl (C₆), cyclohexadienyl (C₆), and the like. Exemplary C_3-8carbocyclyl groups include, without limitation, the aforementioned C_3-6carbocyclyl groups as well as cycloheptyl (C₇), cycloheptenyl (C₇), cycloheptadienyl (C₇), cycloheptatrienyl (C₇), cyclooctyl (C₈), cyclooctenyl (C₈), bicyclo[2.2.1]heptanyl (C₇), bicyclo[2.2.2]octanyl (C₈), and the like. Exemplary C_3-10carbocyclyl groups include, without limitation, the aforementioned C_3-8carbocyclyl groups as well as cyclononyl (C₉), cyclononenyl (C₉), cyclodecyl (C₁₀), cyclodecenyl (C₁₀), octahydro-1H-indenyl (C₉), decahydronaphthalenyl (C₁₀), spiro[4.5]decanyl (C₁₀), and the like. As the foregoing examples illustrate, in certain embodiments, the carbocyclyl group is either monocyclic (“monocyclic carbocyclyl”) or contain a fused, bridged or spiro ring system such as a bicyclic system (“bicyclic carbocyclyl”) and can be saturated or can be partially unsaturated. “Carbocyclyl” also includes ring systems wherein the carbocyclic ring, as defined above, is fused with one or more aryl or heteroaryl groups wherein the point of attachment is on the carbocyclic ring, and in such instances, the number of carbons continue to designate the number of carbons in the carbocyclic ring system. Unless otherwise specified, each instance of a carbocyclyl group is independently optionally substituted, i.e., unsubstituted (an “unsubstituted carbocyclyl”) or substituted (a “substituted carbocyclyl”) with one or more substituents. In certain embodiments, the carbocyclyl group is unsubstituted C₃to carbocyclyl. In certain embodiments, the carbocyclyl group is a substituted C_3-10carbocyclyl.

In some embodiments. “carbocyclyl” is a monocyclic, saturated carbocyclyl group having from 3 to 10 ring carbon atoms (“C_3-10cycloalkyl”). In some embodiments, a cycloalkyl group has 3 to 8 ring carbon atoms (“C_3-8cycloalkyl”). In some embodiments, a cycloalkyl group has 3 to 6 ring carbon atoms (“C_3-6cycloalkyl”). In some embodiments, a cycloalkyl group has 5 to 6 ring carbon atoms (“C_5-6cycloalkyl”). In some embodiments, a cycloalkyl group has 5 to 10 ring carbon atoms (“C_5-10cycloalkyl”). Examples of C_5-6cycloalkyl groups include cyclopentyl (C₅) and cyclohexyl (C₅). Examples of C_3-6cycloalkyl groups include the aforementioned C_5-6cycloalkyl groups as well as cyclopropyl (C₃) and cyclobutyl (C₄). Examples of C_3-8cycloalkyl groups include the aforementioned C_3-6cycloalkyl groups as well as cycloheptyl (C₇) and cyclooctyl (C₈). Unless otherwise specified, each instance of a cycloalkyl group is independently unsubstituted (an “unsubstituted cycloalkyl”) or substituted (a “substituted cycloalkyl”) with one or more substituents. In certain embodiments, the cycloalkyl group is unsubstituted C_3-10cycloalkyl. In certain embodiments, the cycloalkyl group is substituted C_3-10cycloalkyl.

“Aryl” refers to a radical of a monocyclic or polycyclic (e.g., bicyclic or tricyclic) 4n+2 aromatic ring system (e.g., having 6, 10, or 14 pi electrons shared in a cyclic array) having 6-14 ring carbon atoms and zero heteroatoms provided in the aromatic ring system (“C_6-14aryl”). In some embodiments, an aryl group has six ring carbon atoms (“C₆aryl”; e.g., phenyl). In some embodiments, an aryl group has ten ring carbon atoms (“C₁₀aryl”; e.g., naphthyl such as 1-naphthyl and 2-naphthyl). In some embodiments, an aryl group has fourteen ring carbon atoms (“C₁₄aryl”; e.g., anthracyl). “Aryl” also includes ring systems wherein the aryl ring, as defined above, is fused with one or more carbocyclyl or heterocyclyl groups wherein the radical or point of attachment is on the aryl ring, and in such instances, the number of carbon atoms continue to designate the number of carbon atoms in the aryl ring system. Unless otherwise specified, each instance of an aryl group is independently optionally substituted, i.e., unsubstituted (an “unsubstituted aryl”) or substituted (a “substituted aryl”) with one or more substituents. In certain embodiments, the aryl group is unsubstituted C_6-14aryl. In certain embodiments, the aryl group is substituted C_6-14aryl.

“Aralkyl” is a subset of alkyl and aryl and refers to an optionally substituted alkyl group substituted by an optionally substituted aryl group. In certain embodiments, the aralkyl is optionally substituted benzyl. In certain embodiments, the aralkyl is benzyl. In certain embodiments, the aralkyl is optionally substituted phenethyl. In certain embodiments, the aralkyl is phenethyl. In certain embodiments, the aralkyl is 7-phenylheptanyl. In certain embodiments, the aralkyl is C7 alkyl substituted by an optionally substituted aryl group (e.g., phenyl). In certain embodiments, the aralkyl is a C7-C10 alkyl group substituted by an optionally substituted aryl group (e.g., phenyl).

“Partially unsaturated” refers to a group that includes at least one double or triple bond. A “partially unsaturated” ring system is further intended to encompass rings having multiple sites of unsaturation but is not intended to include aromatic groups (e.g., aryl or heteroaryl groups) as defined in this application. Likewise. “saturated” refers to a group that does not contain a double or triple bond, i.e., contains all single bonds.

The term “optionally substituted” means substituted or unsubstituted.

Alkyl, alkenyl, alkynyl, carbocyclyl, heterocyclyl, aryl, and heteroaryl groups are optionally substituted (e.g., “substituted” or “unsubstituted” alkyl, “substituted” or “unsubstituted” alkenyl, “substituted” or “unsubstituted” alkynyl, “substituted” or “unsubstituted” carbocyclyl, “substituted” or “unsubstituted” heterocyclyl, “substituted” or “unsubstituted” aryl or “substituted” or “unsubstituted” heteroaryl group). In general, the term “substituted.” whether preceded by the term “optionally” or not, means that at least one hydrogen present on a group (e.g., a carbon or nitrogen atom) is replaced with a permissible substituent, e.g., a substituent which upon substitution results in a stable compound, e.g., a compound which does not spontaneously undergo transformation such as by rearrangement, cyclization, elimination, or other reaction. Unless otherwise indicated, a “substituted” group has a substituent at one or more substitutable positions of the group, and when more than one position in any given structure is substituted, the substituent is either the same or different at each position. The term “substituted” is contemplated to include substitution with all permissible substituents of organic compounds, any of the substituents described in this application that results in the formation of a stable compound. The present invention contemplates any and all such combinations in order to arrive at a stable compound. For purposes of this invention, heteroatoms such as nitrogen may have hydrogen substituents and/or any suitable substituent as described in this application which satisfy the valencies of the heteroatoms and results in the formation of a stable moiety.

Exemplary carbon atom substituents include, but are not limited to, halogen, —CN, —NO₂, —N₃, —SO₂H, —SO₃H, —OH, —OR^aa, —ON(R^bb)₂, —N(R^bb)₂, —N(R^bb)₃⁺X⁻, —N(OR^cc)R^bb, —SH, —SR^aa, —SSR^cc, —C(═O)R^aa, —CO₂H, —CHO, —C(OR^cc)₂, —CO₂R^aa, —OC(═O)R^aa, —OCO₂R^aa, —C(═O)N(R^bb)₂, —OC(═O)N(R^bb)₂, —NR^bbC(═O)R^aa, —NR^bbCO₂R^aa, —NR^bbC(═O)N(R^bb)₂, —C(═NR^bb)R^aa, —C(═NR^bb)OR^aa, —OC(═NR^bb)R^aa, —OC(═NR^bb)OR^aa, —C(═NR&)N(R^bb)₂, —OC(═NR^bb)N(R^bb)₂, —NR^bbC(═NR^bb)N(R^bb)₂, —C(═O))NR^bbSO₂R^aa, —NR^bbSO₂R^aa, —SO₂N(R^bb)₂, —SO₂R^aa, —SO₂OR^aa, —OSO₂R^aa, —S(═O)R^aa, —OS(═O)R^aa, —Si(R^aa)₃, —OSi(R^aa)₃—C(═S)N(R^bb)₂, —C(═O)SR^aa, —C(═S)SR^aa, —SC(═S)SR^aa, —SC(═O)SR^aa, —OC(═O)SR^aa, —SC(═O)OR^aa, —SC(═O)R^aa, —P(═O)(R^aa)₂, —P(═O)(OR^aa)₂, —OP(═O)(R^aa)₂, —OP(═O)(OR^cc)₂, —P(═O)(N(R^bb)₂)₂, —OP(═O)(N(R^bb)₂)₂, —NR^bbP(═O)(R^aa)₂, —NR^bbP(═O)(OR^cc)₂, —NR^bbP(═O)(N(R^bb)₂)₂, —P(R^cc)₂, —P(OR^cc)₂, —P(R^cc)₃⁺X⁻, —P(OR^cc)₃⁺X⁻, —P(R^cc)₄, —P(OR^cc)₄, —OP(R^cc)₂, —OP(R^cc)₃⁺X⁻, —OP(OR^cc)₂, —OP(OR^cc)₃⁺X⁻, —OP(R^cc)₄, —OP(OR^cc)₄, —B(R^aa)₂, —B(OR^aa)₂, —BR^aa(OR^cc), C_1-10alkyl, C_1-10perhaloalkyl, C_2-10alkenyl, C_2-10alkynyl, heteroC_1-10alkyl, heteroC_2-10alkenyl, heteroC_2-10alkynyl, C_3-10carbocyclyl, 3-14 membered heterocyclyl, C_6-14aryl, and 5-14 membered heteroaryl; wherein:

- each instance of R^aais, independently, selected from C_1-10alkyl, C_1-10perhaloalkyl, C_2-10alkenyl, C_2-10alkynyl, heteroC_1-10alkyl, heteroC_2-10alkenyl, heteroC_2-10alkynyl, C_3-10carbocyclyl, 3-14 membered heterocyclyl, C_6-14aryl, and 5-14 membered heteroaryl, or two R^aagroups are joined to form a 3-14 membered heterocyclyl or 5-14 membered heteroaryl ring, wherein each alkyl, alkenyl, alkynyl, heteroalkyl, heteroalkenyl, heteroalkynyl, carbocyclyl, heterocyclyl, aryl, and heteroaryl is independently substituted with 0, 1, 2, 3, 4, or 5 R^ddgroups;
- each instance of R^bbis, independently, selected from hydrogen, —OH, —OR^aa, —N(R^cc)₂, —CN, —C(═O)R^aa, —C(═O)N(R^cc)₂, —CO₂R^aa, —SO₂R^aa, —C(═NR^cc)OR^aa, —C(═NR^cc)N(R^cc)₂, —SO₂N(R^cc)₂, —SO₂R^cc, —SO₂OR^cc, —SOR^aa, —C(═S)N(R^cc)₂, —C(═O)SR^cc, —C(═S)SR^cc, —P(═O)(R^aa)₂, —P(═O)OR^cc)₂, —P(═O)(N(R^cc)₂)₂, C_1-10alkyl, C_1-10perhaloalkyl, C_2-10alkenyl, C_2-10alkynyl, heteroC_1-10alkyl, heteroC_2-10alkenyl, heteroC_2-10alkynyl, C_3-10carbocyclyl, 3-14 membered heterocyclyl, C_6-14aryl, and 5-14 membered heteroaryl, or two R^bbgroups are joined to form a 3-14 membered heterocyclyl or 5-14 membered heteroaryl ring, wherein each alkyl, alkenyl, alkynyl, heteroalkyl, heteroalkenyl, heteroalkynyl, carbocyclyl, heterocyclyl, aryl, and heteroaryl is independently substituted with 0, 1, 2, 3, 4, or 5 R^ddgroups; wherein X⁻ is a counterion;
- each instance of R^ccis, independently, selected from hydrogen, C_1-10alkyl, C_1-10perhaloalkyl, C_2-10alkenyl, C_2-10alkynyl, heteroC_1-10alkyl, heteroC_2-10alkenyl, heteroC_2-10alkynyl, C_3-10carbocyclyl, 3-14 membered heterocyclyl, C_6-14aryl, and 5-14 membered heteroaryl, or two R^ccgroups are joined to form a 3-14 membered heterocyclyl or 5-14 membered heteroaryl ring, wherein each alkyl, alkenyl, alkynyl, heteroalkyl, heteroalkenyl, heteroalkynyl, carbocyclyl, heterocyclyl, aryl, and heteroaryl is independently substituted with 0, 1, 2, 3, 4, or 5 R^ddgroups;
- each instance of R^ddis, independently, selected from halogen, —CN, —NO₂, —N₃, —SO₂H, —SO₃H, —OH, —OR^cc, —ON(R^ff)₂, —N(R^ff)₂, —N(R^ff)₃⁺X⁻, —N(OR^ee)R^ff, —SH, —SR^ee, —SSR^ee, —C(═O)R^ee, —CO₂H, —CO₂R^ee, —OC(═O)R^ee, —OCO₂R^ee, —C(═O)N(R^ff)₂, —OC(═O)N(R^ff)₂, —NR^ffC(═O)R^ee, —NR^ffCO₂R^ee, —NR^ffC(═O)N(R^ff)₂, —C(═NR^ff)OR^ee, —OC(═NR^ff)R^ee, —OC(═NR^ff)OR^ee, —C(═NR^ff)N(R^ff)₂, —OC(═NR^ff)N(R^ff)₂, —NR^ffC(═NR^ff)N(R^ff)₂, —NR^ffSO₂R^ee, —SO₂N(R^ff)₂. —SO₂R^ee, —SO₂OR^ee, —OSO₂R^ee, —S(═O)R^ee, —Si(R^ee)₃, —OSi(R^ee)₃, —C(═S)N(R^ff)₂, —C(═O)SR^ee, —C(═S)SR^ee, —SC(═S)SR^ee, —P(═O)(OR^ee)₂, —P(═O)(R^ee)₂, —OP(═O)(R^ee)₂, —OP(═O)(OR^ee)₂, C_1-6alkyl, C_1-6perhaloalkyl, C_2-6alkenyl, C_2-6alkynyl, heteroC_1-6alkyl, heteroC_2-6alkenyl, heteroC_2-6alkynyl, C_3-10carbocyclyl, 3-10 membered heterocyclyl, C_6-10aryl, 5-10 membered heteroaryl, wherein each alkyl, alkenyl, alkynyl, heteroalkyl, heteroalkenyl, heteroalkynyl, carbocyclyl, heterocyclyl, aryl, and heteroaryl is independently substituted with 0, 1, 2, 3, 4, or 5 R^gggroups, or two geminal R^ddsubstituents can be joined to form ═O or ═S; wherein X⁻ is a counterion;
- each instance of R^eeis, independently, selected from C_1-6alkyl, C_1-6perhaloalkyl, C_2-6alkenyl, C_2-6alkynyl, heteroC_1-6alkyl, heteroC_2-6alkenyl, heteroC_2-6alkynyl, C_3-10carbocyclyl, C_6-10aryl, 3-10 membered heterocyclyl, and 3-10 membered heteroaryl, wherein each alkyl, alkenyl, alkynyl, heteroalkyl, heteroalkenyl, heteroalkynyl, carbocyclyl, heterocyclyl, aryl, and heteroaryl is independently substituted with 0, 1, 2, 3, 4, or 5 R^gggroups;
- each instance of R^ffis, independently, selected from hydrogen, C_1-6alkyl, C_1-6perhaloalkyl, C_2-6alkenyl, C_2-6alkynyl, heteroC_1-6alkyl, heteroC_2-6alkenyl, heteroC_2-6alkynyl, C_3-10carbocyclyl, 3-10 membered heterocyclyl, C_6-10aryl and 5-10 membered heteroaryl, or two R^ffgroups are joined to form a 3-10 membered heterocyclyl or 5-10 membered heteroaryl ring, wherein each alkyl, alkenyl, alkynyl, heteroalkyl, heteroalkenyl, heteroalkynyl, carbocyclyl, heterocyclyl, aryl, and heteroaryl is independently substituted with 0, 1, 2, 3, 4, or 5 R^gggroups; and
- each instance of R^ggis, independently, halogen, —CN, —NO₂, —N₃, —SO₂H, —SO₃H, —OH, —OC_1-6alkyl, —ON(C_1-6alkyl)₂, —N(C_1-6alkyl)₂, —N(C_1-6alkyl)₃⁺X⁻, —NH(C_1-6alkyl)₂⁺X⁻, —NH₂(C_1-6alkyl)⁺X⁻, —NH₃⁺X⁻, —N(OC_1-6alkyl)(C_1-6alkyl), —N(OH)(C_1-6alkyl), —NH(OH), —SH, —SC_1-6alkyl, —SS(C_1-6alkyl), —C(═O)(C_1-6alkyl), —CO₂H, —CO₂(C_1-6alkyl), —OC(═O)(C_1-6alkyl), —OCO₂(C_1-6alkyl), —C(═O)NH₂, —C(═O)N(C_1-6alkyl)₂, —OC(═O)NH(C_1-6alkyl), —NHC(═O)(C_1-6alkyl), —N(C_1-6alkyl)C(═O)(C_1-6alkyl), —NHCO₂(C_1-6alkyl), —NHC(═O)N(C_1-6alkyl)₂, —NHC(═O)NH(C_1-6alkyl), —NHC(═O)NH₂, —C(═NH)O(C_1-6alkyl), —OC(═NH)(C_1-6alkyl), —OC(═NH)OC_1-6alkyl, —C(═NH)N(C_1-6alkyl)₂, —C(═NH)NH(C_1-6alkyl), —C(═NH)NH₂, —OC(═NH)N(C_1-6alkyl)₂, —OC(NH)NH(C_1-6alkyl), —OC(NH)NH₂, —NHC(NH)N(C_1-6alkyl)₂, —NHC(═NH)NH₂, —NHSO₂(C_1-6alkyl), —SO₂N(C_1-6alkyl)₂, —SO₂NH(C_1-6alkyl), —SO₂NH₂, —SO₂C_1-6alkyl, —SO₂OC_1-6alkyl, —OSO₂C_1-6alkyl, —SOC_1-6alkyl, —Si(C_1-6alkyl)₃, —OSi(C_1-6alkyl)₃-C(═S)N(C_1-6alkyl)₂, C(═S)NH(C_1-6alkyl), C(═S)NH₂, —C(═O)S(C_1-6alkyl), —C(═S)SC_1-6alkyl, —SC(═S)SC_1-6alkyl, —P(═O)(OC_1-6alkyl)₂, —P(═O)(C_1-6alkyl)₂, —OP(═O)(C_1-6alkyl)₂, —OP(═O)(OC_1-6alkyl)₂, C_1-6alkyl, C_1-6perhaloalkyl, C_2-6alkenyl, C_2-6alkynyl, heteroC_1-6alkyl, heteroC_2-6alkenyl, heteroC_2-6alkynyl, C₃-10 carbocyclyl, C_6-10aryl, 3-10 membered heterocyclyl, 5-10 membered heteroaryl; or two geminal R^ggsubstituents can be joined to form ═O or ═S; wherein X⁻ is a counterion. Alternatively, two geminal hydrogens on a carbon atom are replaced with the group ═O, ═S, ═NN(R^bb)₂, ═NR^bbC(═O)R^aa, ═NNR^bbC(═O)OR^aa, ═NNR^bbS(═O)₂R^aa, ═NR^bb, or ═NOR^cc; wherein each alkyl, alkenyl, alkynyl, heteroalkyl, heteroalkenyl, heteroalkynyl, carbocyclyl, heterocyclyl, aryl, and heteroaryl is independently substituted with 0, 1, 2, 3, 4, or 5 R^ddgroups; wherein X⁻ is a counterion;
  
  wherein:
- each instance of R^aais, independently, selected from C_1-10alkyl, C_1-10perhaloalkyl, C_2-10alkenyl, C_2-10alkynyl, heteroC_1-10alkyl, heteroC_2-10alkenyl, heteroC_2-10alkynyl. C_3-10carbocyclyl, 3-14 membered heterocyclyl, C_6-14aryl, and 5-14 membered heteroaryl, or two R^aagroups are joined to form a 3-14 membered heterocyclyl or 5-14 membered heteroaryl ring, wherein each alkyl, alkenyl, alkynyl, heteroalkyl, heteroalkenyl, heteroalkynyl, carbocyclyl, heterocyclyl, aryl, and heteroaryl is independently substituted with 0, 1, 2, 3, 4, or 5 R^ddgroups;
- each instance of R^bbis, independently, selected from hydrogen, —OH, —OR^aa, —N(R^cc)₂, —CN, —C(═O)R^aa, —C(═O)N(R^cc)₂, —CO₂R^aa, —SO₂R^aa, —C(═NR^cc)OR^aa, —C(═NR^cc)N(R^cc)₂, —SO₂N(R^cc)₂, —SO₂R^cc, —SO₂OR^cc, —SOR^aa, —C(═S)N(R^cc)₂, —C(═O)SR^cc, —C(═S)SR^cc, —P(═O)(R^aa)₂, —P(═O)(OR^cc)₂, —P(═O)(N(R^cc)₂)₂, C_1-10alkyl, C_1-10perhaloalkyl, C_2-10alkenyl, C_2-10alkynyl, heteroC_1-10alkyl, heteroC_2-10alkenyl, heteroC_2-10alkynyl, C_3-10carbocyclyl, 3-14 membered heterocyclyl, C_6-14aryl, and 5-14 membered heteroaryl, or two R^bbgroups are joined to form a 3-14 membered heterocyclyl or 5-14 membered heteroaryl ring, wherein each alkyl, alkenyl, alkynyl, heteroalkyl, heteroalkenyl, heteroalkynyl, carbocyclyl, heterocyclyl, aryl, and heteroaryl is independently substituted with 0, 1, 2, 3, 4, or 5 R^ddgroups; wherein X⁻ is a counterion;
- each instance of R^ccis, independently, selected from hydrogen, C_1-10alkyl, C_1-10perhaloalkyl, C_2-10alkenyl, C_2-10alkynyl, heteroC_1-10alkyl, heteroC_2-10alkenyl, heteroC_2-10alkynyl, C_3-10carbocyclyl, 3-14 membered heterocyclyl, C_6-14aryl, and 5-14 membered heteroaryl, or two R^ccgroups are joined to form a 3-14 membered heterocyclyl or 5-14 membered heteroaryl ring, wherein each alkyl, alkenyl, alkynyl, heteroalkyl, heteroalkenyl, heteroalkynyl, carbocyclyl, heterocyclyl, aryl, and heteroaryl is independently substituted with 0, 1, 2, 3, 4, or 5 R^ddgroups;
- each instance of R^ddis, independently, selected from halogen, —CN, —NO₂, —N₃, —SO₂H, —SO₃H, —OH, —OR^ee, —ON(R^ff)₂, —N(R^ff)₂, —N(R^ff)₃⁺X⁻, —N(OR^ee)R^ff, —SH, —SR^ee, —SSR^ee, —C(═O)R^ee, —CO₂H, —CO₂R^ee, —OC(═O)R^ee, —OCO₂R^ee, —C(═O)N(R^ff)₂, —OC(═O)N(R^ff)₂, —NR^ffC(═O)R^ee, —NR^ffCO₂R^ee, —NR^ffC(═O)N(R^ff)₂, —C(═NR^ff)OR^ee, —OC(═NR^ff)R^ee, —OC(═NR^ff)OR^ee, —C(═NR^ff)N(R^ff)₂, —OC(═NR^ff)N(R^ff)₂, —NR^ffC(═NR^ff)N(R^ff)₂, —NR^ffSO₂R^ee, —SO₂N(R^ff)₂, —SO₂R^ee, —SO₂OR^ee, OSO₂R^ee, —S(═O)R^ee, —Si(R^ee)₃, —OSi(R^ee)₃, —C(═S)N(R^ff)₂, —C(═O)SR^ee, —C(═S)SR^ee, —SC(═S)SR^ee, —P(═O)(OR^ee)₂, —P(═O)(R^ee)₂, —OP(═O)(R^ee)₂, —OP(═O)(OR^ee)₂, C_1-6alkyl, C_1-6perhaloalkyl, C_2-6alkenyl, C_2-6alkynyl, heteroC_1-6alkyl, heteroC_2-6alkenyl, heteroC_2-6alkynyl, C_3-10carbocyclyl, 3-10 membered heterocyclyl. C_6-10aryl, 5-10 membered heteroaryl, wherein each alkyl, alkenyl, alkynyl, heteroalkyl, heteroalkenyl, heteroalkynyl, carbocyclyl, heterocyclyl, aryl, and heteroaryl is independently substituted with 0, 1, 2, 3, 4, or 5 Ru groups, or two geminal R^ddsubstituents can be joined to form ═O or ═S; wherein X is a counterion;
- each instance of R^eeis, independently, selected from C_1-6alkyl, C^1-6perhaloalkyl, C_2-6alkenyl, C_2-6alkynyl, heteroC_1-6alkyl, heteroC_2-6alkenyl, heteroC_2-6alkynyl, C_3-10carbocyclyl, C_6-10aryl, 3-10 membered heterocyclyl, and 3-10 membered heteroaryl, wherein each alkyl, alkenyl, alkynyl, heteroalkyl, heteroalkenyl, heteroalkynyl, carbocyclyl, heterocyclyl, aryl, and heteroaryl is independently substituted with 0, 1, 2, 3, 4, or 5 R^gggroups;
- each instance of R^ffis, independently, selected from hydrogen, C_1-6alkyl, C_1-6perhaloalkyl, C_2-6alkenyl, C_2-6alkynyl, heteroC_1-6alkyl, heteroC_2-6alkenyl, heteroC_2-6alkynyl, C_3-10carbocyclyl, 3-10 membered heterocyclyl, C_6-10aryl and 5-10 membered heteroaryl, or two R^ffgroups are joined to form a 3-10 membered heterocyclyl or 5-10 membered heteroaryl ring, wherein each alkyl, alkenyl, alkynyl, heteroalkyl, heteroalkenyl, heteroalkynyl, carbocyclyl, heterocyclyl, aryl, and heteroaryl is independently substituted with 0, 1, 2, 3, 4, or 5 R^gggroups; and
- each instance of R^ggis, independently, halogen, —CN, —NO₂, —N₃, —SO₂H, —SO₃H, —OH, —OC_1-6alkyl, —ON(C_1-6alkyl)₂, —N(C_1-6alkyl)₂, —N(C_1-6alkyl)₃⁺X⁻, —NH(C_1-6alkyl)₂+X⁻, —NH₂(C_1-6alky)⁺X⁻, —NH₃⁺X⁻, —N(OC_1-6alkyl)(C_1-6alkyl), —N(OH)(C_1-6alkyl), —NH(OH), —SH, —SC_1-6alkyl, —SS(C_1-6alkyl), —C(═O)(C_1-6alkyl), —CO₂H, —CO₂(C_1-6alkyl), —OC(═O)(C_1-6alkyl), —OCO₂(C_1-6alkyl), —C(═O)NH₂, —C(═O)N(C_1-6alkyl)₂, —OC(═O)NH(C_1-6alkyl), —NHC(═O)(C_1-6alkyl), —N(C_1-6alkyl)C(═O)(C_1-6alkyl), —NHCO₂(C_1-6alkyl), —NHC(═O)N(C_1-6alkyl)₂, —NHC(=)NH(C_1-6alkyl), —NHC(═O)NH₂, —C(═NH)O(C_1-6alkyl), —OC(═NH)(C_1-6alkyl), —OC(═NH)OC_1-6alkyl, —C(═NH)N(C_1-6alkyl)₂, —C(═NH)NH(C_1-6alkyl), —C(═NH)NH₂, —OC(═NH)N(C_1-6alkyl)₂, —OC(NH)NH(C_1-6alkyl), —OC(NH)NH₂, —NHC(NH)N(C_1-6alkyl)₂, —NHC(═NH)NH₂, —NHSO₂(C_1-6alkyl), —SO₂N(C_1-6alkyl)₂, —SO₂NH(C_1-6alkyl), —SO₂NH₂, —SO₂C_1-6alkyl, —SO₂OC_1-6alkyl, —OSO₂C_1-6alkyl, —SOC_1-6alkyl, —Si(C_1-6alkyl)₃, —OSi(C_1-6alkyl)₃, —C(═S)N(C_1-6alkyl)₂, C(═S)NH(C_1-6alkyl), C(═S)NH₂, —C(═O)S(C_1-6alkyl), —C(═S)SC_1-6alkyl, —SC(═S)SC_1-6alkyl, —P(═O)(OC_1-6alkyl)₂, —P(═O)(C_1-6alkyl)₂, —OP(═O)(C_1-6alkyl)₂, —OP(═O)(OC_1-6alkyl)₂, C_1-6alkyl, C_1-6perhaloalkyl, C_2-6alkenyl, C_2-6alkynyl, heteroC_1-6alkyl, heteroC_2-6alkenyl, heteroC_2-6alkynyl, C_3-10carbocyclyl, C_6-10aryl, 3-10 membered heterocyclyl, 5-10 membered heteroaryl; or two geminal Ru substituents can be joined to form ═O or ═S; wherein X⁻ is a counterion.

A “counterion” or “anionic counterion” is a negatively charged group associated with a positively charged group in order to maintain electronic neutrality. An anionic counterion may be monovalent (i.e., including one formal negative charge). An anionic counterion may also be multivalent (i.e., including more than one formal negative charge), such as divalent or trivalent. Exemplary counterions include halide ions (e.g., F⁻, Cl⁻, Br⁻, I⁻), NO₃⁻, ClO₄⁻, OH⁻, H₂PO₄⁻, HCO₃⁻, HSO₄⁻, sulfonate ions (e.g., methansulfonate, trifluoromethanesulfonate, p-toluenesulfonate, benzenesulfonate, 10-camphor sulfonate, naphthalene-2-sulfonate, naphthalene-1-sulfonic acid-5-sulfonate, ethan-1-sulfonic acid-2-sulfonate, and the like), carboxylate ions (e.g., acetate, propanoate, benzoate, glycerate, lactate, tartrate, glycolate, gluconate, and the like), BF₄⁻, PF₄⁻, PF₆⁻, AsF₆⁻, SbF₆⁻, B[3,5-(CF₃)₂C₆H₃]₄]⁻, B(C₆F₅)₄⁻, BPh₄⁻, Al(OC(CF₃)₃)₄⁻, and carborane anions (e.g., CB₁₁H₁₂⁻ or (HCB₁₁Me₅Br₆)⁻). Exemplary counterions which may be multivalent include CO₃²⁻, HPO₄²⁻, PO₄³⁻, B₄O₇²⁻, SO₄²⁻, S₂O₃²⁻, carboxylate anions (e.g., tartrate, citrate, fumarate, maleate, malate, malonate, gluconate, succinate, glutarate, adipate, pimelate, suberate, azelate, sebacate, salicylate, phthalates, aspartate, glutamate, and the like), and carboranes.

The term “pharmaceutically acceptable salt” refers to those salts which are, within the scope of sound medical judgment, suitable for use in contact with the tissues of humans and lower animals without undue toxicity, irritation, allergic response and the like, and are commensurate with a reasonable benefit/risk ratio. Pharmaceutically acceptable salts are well known in the art. For example, Berge et al., describe pharmaceutically acceptable salts in detail in J. Pharmaceutical Sciences, 1977, 66, 1-19, incorporated by reference. Pharmaceutically acceptable salts of the compounds disclosed in this application include those derived from suitable inorganic and organic acids and bases. Examples of pharmaceutically acceptable, nontoxic acid addition salts are salts of an amino group formed with inorganic acids such as hydrochloric acid, hydrobromic acid, phosphoric acid, sulfuric acid, and perchloric acid or with organic acids such as acetic acid, oxalic acid, maleic acid, tartaric acid, citric acid, succinic acid, or malonic acid or by using other methods known in the art such as ion exchange. Other pharmaceutically acceptable salts include adipate, alginate, ascorbate, aspartate, benzenesulfonate, benzoate, bisulfate, borate, butyrate, camphorate, camphorsulfonate, citrate, cyclopentanepropionate, digluconate, dodecylsulfate, ethanesulfonate, formate, fumarate, glucoheptonate, glycerophosphate, gluconate, hemisulfate, heptanoate, hexanoate, hydroiodide, 2-hydroxy-ethanesulfonate, lactobionate, lactate, laurate, lauryl sulfate, malate, maleate, malonate, methanesulfonate, 2-naphthalenesulfonate, nicotinate, nitrate, oleate, oxalate, palmitate, pamoate, pectinate, persulfate, 3-phenylpropionate, phosphate, picrate, pivalate, propionate, stearate, succinate, sulfate, tartrate, thiocyanate, p-toluenesulfonate, undecanoate, valerate salts, and the like. Salts derived from appropriate bases include alkali metal, alkaline earth metal, ammonium and N⁺(C_1-4alkyl)₄⁻ salts. Representative alkali or alkaline earth metal salts include sodium, lithium, potassium, calcium, magnesium, and the like. Further pharmaceutically acceptable salts include, when appropriate, nontoxic ammonium, quaternary ammonium, and amine cations formed using counterions such as halide, hydroxide, carboxylate, sulfate, phosphate, nitrate, lower alkyl sulfonate, and aryl sulfonate.

The term “solvate” refers to forms of a compound that are associated with a solvent, usually by a solvolysis reaction. This physical association may include hydrogen bonding. Conventional solvents include water, methanol, ethanol, acetic acid, DMSO, THF, diethyl ether, and the like. The compounds of Formula (1), (9), (10), and (11) may be prepared, e.g., in crystalline form, and may be solvated. Suitable solvates include pharmaceutically acceptable solvates and further include both stoichiometric solvates and non-stoichiometric solvates. In certain instances, the solvate will be capable of isolation, for example, when one or more solvent molecules are incorporated in the crystal lattice of a crystalline solid. “Solvate” encompasses both solution-phase and isolable solvates. Representative solvates include hydrates, ethanolates, and methanolates.

The term “hydrate” refers to a compound that is associated with water. Typically, the number of the water molecules contained in a hydrate of a compound is in a definite ratio to the number of the compound molecules in the hydrate. Therefore, a hydrate of a compound may be represented, for example, by the general formula R·x H₂O, wherein R is the compound and wherein x is a number greater than 0. A given compound may form more than one type of hydrates, including, e.g., monohydrates (x is 1), lower hydrates (x is a number greater than 0 and smaller than 1, e.g., hemihydrates (R·0.5 H₂O)), and polyhydrates (x is a number greater than 1, e.g., dihydrates (R·2 H₂O) and hexahydrates (R·6 H₂O)).

The term “tautomers” refer to compounds that are interchangeable forms of a particular compound structure, and that vary in the displacement of hydrogen atoms and electrons. Thus, two structures may be in equilibrium through the movement of n electrons and an atom (usually H). For example, enols and ketones are tautomers because they are rapidly interconverted by treatment with either acid or base. Another example of tautomerism is the aci- and nitro-forms of phenylnitromethane, which are likewise formed by treatment with acid or base. Tautomeric forms may be relevant to the attainment of the optimal chemical reactivity and biological activity of a compound of interest.

It is also to be understood that compounds that have the same molecular formula but differ in the nature or sequence of bonding of their atoms or the arrangement of their atoms in space are termed “isomers.” Isomers that differ in the arrangement of their atoms in space are termed “stereoisomers.”

Stereoisomers that are not mirror images of one another are termed “diastereomers” and those that are non-superimposable mirror images of each other are termed “enantiomers.” When a compound has an asymmetric center, for example, it is bonded to four different groups, a pair of enantiomers is possible. An enantiomer can be characterized by the absolute configuration of its asymmetric center and described by the R- and S-sequencing rules of Cahn and Prelog. An enantiomer can also be characterized by the manner in which the molecule rotates the plane of polarized light, and designated as dextrorotatory or levorotatory (i.e., as (+) or (−)-isomers respectively). A chiral compound can exist as either an individual enantiomer or as a mixture of enantiomers. A mixture containing equal proportions of the enantiomers is called a “racemic mixture.”

The term “co-crystal” refers to a crystalline structure comprising at least two different components (e.g., a compound described in this application and an acid), wherein each of the components is independently an atom, ion, or molecule. In certain embodiments, none of the components is a solvent. In certain embodiments, at least one of the components is a solvent. A co-crystal of a compound and an acid is different from a salt formed from a compound and the acid. In the salt, a compound described in this application is complexed with the acid in a way that proton transfer (e.g., a complete proton transfer) from the acid to a compound described in this application easily occurs at room temperature. In the co-crystal, however, a compound described in this application is complexed with the acid in a way that proton transfer from the acid to a compound described in this application does not easily occur at room temperature. In certain embodiments, in the co-crystal, there is no proton transfer from the acid to a compound described in this application. In certain embodiments, in the co-crystal, there is partial proton transfer from the acid to a compound described in this application. Co-crystals may be useful to improve the properties (e.g., solubility, stability, and ease of formulation) of a compound described in this application.

The term “polymorphs” refers to a crystalline form of a compound (or a salt, hydrate, or solvate thereof) in a particular crystal packing arrangement. All polymorphs of the same compound have the same elemental composition. Different crystalline forms usually have different X-ray diffraction patterns, infrared spectra, melting points, density, hardness, crystal shape, optical and electrical properties, stability, and solubility. Recrystallization solvent, rate of crystallization, storage temperature, and other factors may cause one crystal form to dominate. Various polymorphs of a compound can be prepared by crystallization under different conditions.

The term “prodrug” refers to compounds, including derivatives of the compounds of Formula (X), (8), (9), (10), or (11), that have cleavable groups and become by solvolysis or under physiological conditions the compounds of Formula (X), (8), (9), (10), or (11) and that are pharmaceutically active in vivo. The prodrugs may have attributes such as, without limitation, solubility, bioavailability, tissue compatibility, or delayed release in a mammalian organism. Examples include, but are not limited to, derivatives of compounds described in this application, including derivatives formed from glycosylation of the compounds described in this application (e.g., glycoside derivatives), carrier-linked prodrugs (e.g., ester derivatives), bioprecursor prodrugs (a prodrug metabolized by molecular modification into the active compound), and the like. Non-limiting examples of glycoside derivatives are disclosed in and incorporated by reference from PCT Publication No. WO2018/208875 and U.S. Patent Publication No. 2019/0078168. Non-limiting examples of ester derivatives are disclosed in and incorporated by reference from U.S. Patent Publication No. US2017/0362195.

Other derivatives of the compounds of this invention have activity in both their acid and acid derivative forms, but the acid sensitive form often offers advantages of solubility, bioavailability, tissue compatibility, or delayed release in a mammalian organism (see, Bundgard, H., Design of Prodrugs, pp. 7-9, 21-24, Elsevier, Amsterdam 1985). Prodrugs include acid derivatives well known to practitioners of the art, such as, for example, esters prepared by reaction of the parent acid with a suitable alcohol, or amides prepared by reaction of the parent acid compound with a substituted or unsubstituted amine, or acid anhydrides, or mixed anhydrides. Simple aliphatic or aromatic esters, amides, and anhydrides derived from acidic groups pendant on the compounds of this invention are particular prodrugs. In some cases it is desirable to prepare double ester type prodrugs such as (acyloxy)alkyl esters or ((alkoxycarbonyl)oxy)alkylesters. C₁-C₈alkyl, C₂-C₈alkenyl, C₂-C₈alkynyl, aryl, C₇-C₁₂substituted aryl, and C₇-C₁₂arylalkyl esters of the compounds of Formula (X), (8), (9). (10), or (11) may be preferred.

Cannabinoids

As used in this application, the term “cannabinoid” includes compounds of Formula (X):

embedded image

or a pharmaceutically acceptable salt, co-crystal, tautomer, stereoisomer, solvate, hydrate, polymorph, isotopically enriched derivative, or prodrug thereof, wherein R1 is hydrogen, optionally substituted acyl, optionally substituted alkyl, optionally substituted alkenyl, optionally substituted alkynyl, optionally substituted carbocyclyl, or optionally substituted aryl; R2 and R6 are, independently, hydrogen or carboxyl; R3 and R5 are, independently, hydroxyl, halogen, or alkoxy; and R4 is a hydrogen or an optionally substituted prenyl moiety; or optionally R4 and R3 are taken together with their intervening atoms to form a cyclic moiety, or optionally R4 and R5 are taken together with their intervening atoms to form a cyclic moiety, or optionally both 1) R4 and R3 are taken together with their intervening atoms to form a cyclic moiety and 2) R4 and R5 are taken together with their intervening atoms to form a cyclic moiety. In certain embodiments, R4 and R3 are taken together with their intervening atoms to form a cyclic moiety. In certain embodiments, R4 and R5 are taken together with their intervening atoms to form a cyclic moiety. In certain embodiments, “cannabinoid” refers to a compound of Formula (X), or a pharmaceutically acceptable salt thereof. In certain embodiments, both 1) R4 and R3 are taken together with their intervening atoms to form a cyclic moiety and 2) R4 and R5 are taken together with their intervening atoms to form a cyclic moiety.

In some embodiments, cannabinoids may be synthesized via the following steps: a) one or more reactions to incorporate three additional ketone moieties onto an acyl-CoA scaffold, where the acyl moiety in the acyl-CoA scaffold comprises between four and fourteen carbons; b) a reaction cyclizing the product of step (a); and c) a reaction to incorporate a prenyl moiety to the product of step (b) or a derivative of the product of step (b). In some embodiments, non-limiting examples of the acyl-CoA scaffold described in step (a) include hexanoyl-CoA and butyryl-CoA. In some embodiments, non-limiting examples of the product of step (b) or a derivative of the product of step (b) include olivetolic acid, divarinic acid, and sphaerophorolic acid.

In some embodiments, a cannabinoid compound of Formula (X) is of Formula (X-A), (X-B), or (X-C):

embedded image

or a pharmaceutically acceptable salt, solvate, hydrate, polymorph, co-crystal, tautomer, stereoisomer, isotopically labeled derivative, or prodrug thereof;

wherein custom-character is a double bond or a single bond, as valency permits;

- R is hydrogen, optionally substituted acyl, optionally substituted alkyl, optionally substituted alkenyl, optionally substituted alkynyl, optionally substituted carbocyclyl, or optionally substituted aryl;
- R^Z1is hydrogen, optionally substituted acyl, optionally substituted alkyl, optionally substituted alkenyl, optionally substituted alkynyl, optionally substituted carbocyclyl, or optionally substituted aryl;
- R^Z2is hydrogen, optionally substituted acyl, optionally substituted alkyl, optionally substituted alkenyl, optionally substituted alkynyl, optionally substituted carbocyclyl, or optionally substituted aryl;
- or optionally, R^Z1and R^Z2are taken together with their intervening atoms to form an optionally substituted carbocyclic ring;
- R^3Ais hydrogen, optionally substituted acyl, optionally substituted alkyl, optionally substituted alkenyl, or optionally substituted alkynyl;
- R^3Bis hydrogen, optionally substituted acyl, optionally substituted alkyl, optionally substituted alkenyl, or optionally substituted alkynyl;
- R^Yis hydrogen, optionally substituted acyl, optionally substituted alkyl, optionally substituted alkenyl, or optionally substituted alkynyl;
- R^Zis hydrogen, optionally substituted acyl, optionally substituted alkyl, optionally substituted alkenyl, or optionally substituted alkynyl.

In certain embodiments, a cannabinoid compound is of Formula (X-A):

embedded image

wherein custom-character is a double bond, and each of R^Z1and R^Z2is hydrogen, one of R^3Aand R^3Bis optionally substituted C_2-6alkenyl, and the other one of R^3Aand R^3Bis optionally substituted C_2-6alkyl. In some embodiments, a cannabinoid compound of Formula (X) is of Formula (X-A), wherein each of R^Z1and R^Z2is hydrogen, one of R^3Aand R^3Bis a prenyl group, and the other one of R^3Aand R^3Bis optionally substituted methyl.

In certain embodiments, a cannabinoid compound of Formula (X) of Formula (X-A) is of Formula (11-z):

embedded image

wherein custom-character is a double bond or single bond, as valency permits; one of R^3Aand R^3Bis C_1-6alkyl optionally substituted with alkenyl, and the other of R^3Aand R^3Bis optionally substituted C_1-6alkyl. In certain embodiments, in a compound of Formula (11-z), is a single bond; one of R^3Aand R^3Bis C_1-6alkyl optionally substituted with prenyl; and the other of one of R^3Aand R^3Bis unsubstituted methyl; and R is as described in this application. In certain embodiments, in a compound of Formula (1-z), custom-character is a single bond; one of R^3Aand R^3Bis

embedded image

and the other of one of R^3Aand R^3Bis unsubstituted methyl; and R is as described in this application. In certain embodiments, a cannabinoid compound of Formula (11-z) is of Formula (11a):

embedded image

In certain embodiments, a cannabinoid compound of Formula (X) of Formula (X-A) is of Formula (11a):

embedded image

In certain embodiments, a cannabinoid compound of Formula (11-z) is of Formula (11b):

embedded image

In certain embodiments, a cannabinoid compound of Formula (X) of Formula (X-A) is of Formula (11b):

embedded image

In certain embodiments, a cannabinoid compound of Formula (X-A) is of Formula (10-z):

embedded image

wherein custom-character is a double bond or single bond, as valency permits. R^Yis hydrogen, optionally substituted acyl, optionally substituted alkyl, optionally substituted alkenyl, or optionally substituted alkynyl; and each of R^3Aand R^3Bis independently optionally substituted C_1-6alkyl. In certain embodiments, in a compound of Formula (10-z), custom-character is a single bond; each of R^3Aand R^3Bis unsubstituted methyl, and R is as described in this application. In certain embodiments, a cannabinoid compound of Formula (10-z) is of Formula (10a):

embedded image

In certain embodiments, a compound of Formula (10a)

embedded image

has a chiral atom labeled with * at carbon 10 and a chiral atom labeled with ** at carbon 6. In certain embodiments, in a compound of Formula (10a)

embedded image

the chiral atom labeled with * at carbon 10 is of the R-configuration or S-configuration; and a chiral atom labeled with ** at carbon 6 is of the R-configuration. In certain embodiments, in a compound of Formula (10a)

embedded image

the chiral atom labeled with * at carbon 10 is of the S-configuration; and a chiral atom labeled with ** at carbon 6 is of the R-configuration or S-configuration. In certain embodiments, in a compound of Formula (10a)

embedded image

the chiral atom labeled with * at carbon 10 is of the R-configuration and a chiral atom labeled with ** at carbon 6 is of the R-configuration. In certain embodiments, a compound of Formula (10a)

embedded image

is of the formula:

embedded image

In certain embodiments, in a compound of Formula (10a)

embedded image

the chiral atom labeled with * at carbon 10 is of the S-configuration and a chiral atom labeled with ** at carbon 6 is of the S-configuration. In certain embodiments, a compound of Formula (10a)

embedded image

is of the formula:

embedded image

In certain embodiments, a cannabinoid compound of Formula (10-z) is of Formula (10b):

embedded image

In certain embodiments, a compound of Formula (10b)

embedded image

has a chiral atom labeled with * at carbon 10 and a chiral atom labeled with ** at carbon 6. In certain embodiments, in a compound of Formula (10b)

embedded image

the chiral atom labeled with * at carbon 10 is of the R-configuration and a chiral atom labeled with ** at carbon 6 is of the R-configuration. In certain embodiments, a compound of Formula (10b)

embedded image

is of the formula:

embedded image

In certain embodiments, in a compound of Formula (10b)

embedded image

the chiral atom labeled with * at carbon 10 is of the S-configuration and a chiral atom labeled with ** at carbon 6 is of the S-configuration. In certain embodiments, a compound of Formula (10b)

embedded image

is of the formula:

embedded image

In certain embodiments, a cannabinoid compound is of Formula (X-B):

embedded image

wherein custom-character is a double bond; R^Yis hydrogen, optionally substituted acyl, optionally substituted alkyl, optionally substituted alkenyl, or optionally substituted alkynyl; and each of R^3Aand R^3Bis independently optionally substituted C_1-6alkyl. In certain embodiments, in a compound of Formula (X-B), R^Yis optionally substituted C_1-6alkyl; one of R^3Aand R^3Bis

embedded image

and the other one of R^3Aand R^3Bis unsubstituted methyl, and R is as described in this application. In certain embodiments, a compound of Formula (X-B) is of Formula (9a):

embedded image

In certain embodiments, a compound of Formula (9a)

embedded image

has a chiral atom labeled with * at carbon 3 and a chiral atom labeled with ** at carbon 4. In certain embodiments, in a compound of Formula (9a)

embedded image

the chiral atom labeled with * at carbon 3 is of the R-configuration or S-configuration; and a chiral atom labeled with ** at carbon 4 is of the R-configuration. In certain embodiments, in a compound of Formula (9a)

embedded image

the chiral atom labeled with * at carbon 3 is of the S-configuration; and a chiral atom labeled with ** at carbon 4 is of the R-configuration or S-configuration. In certain embodiments, in a compound of Formula (9a)

embedded image

the chiral atom labeled with * at carbon 3 is of the R-configuration and a chiral atom labeled with ** at carbon 4 is of the R-configuration. In certain embodiments, a compound of Formula (9a)

embedded image

is of the formula:

embedded image

In certain embodiments, in a compound of Formula (9a)

embedded image

the chiral atom labeled with * at carbon 3 is of the S-configuration and a chiral atom labeled with ** at carbon 4 is of the S-configuration. In certain embodiments, a compound of Formula (9a)

embedded image

is of the formula:

embedded image

In certain embodiments, a compound of Formula (X-B) is of Formula (9b):

embedded image

In certain embodiments, a compound of Formula (9b)

embedded image

has a chiral atom labeled with * at carbon 3 and a chiral atom labeled with ** at carbon 4. In certain embodiments, in a compound of Formula (9b)

embedded image

the chiral atom labeled with * at carbon 3 is of the R-configuration and a chiral atom labeled with ** at carbon 4 is of the R-configuration. In certain embodiments, a compound of Formula (9b)

embedded image

is of the formula:

embedded image

In certain embodiments, in a compound of Formula (9a)

embedded image

the chiral atom labeled with * at carbon 3 is of the S-configuration and a chiral atom labeled with ** at carbon 4 is of the S-configuration. In certain embodiments, a compound of Formula (9b)

embedded image

is of the formula:

embedded image

In certain embodiments, a cannabinoid compound is of Formula (X-C):

embedded image

wherein R^Zis optionally substituted alkyl or optionally substituted alkenyl. In certain embodiments, a compound of Formula (X-C) is of formula:

embedded image

wherein a is 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10. In certain embodiments, a is 1. In certain embodiments, a is 2. In certain embodiments, a is 3. In certain embodiments, a is 1, 2, or 3 for a compound of Formula (X-C). In certain embodiments, a cannabinoid compound is of Formula (X-C), and a is 1, 2, 3, 4, or 5. In certain embodiments, a compound of Formula (X-C) is of Formula (8a):

embedded image

In some embodiments, a cannabinoid compound of Formula (X-1) is of Formula (X-A-1), (X-B-1), or (X-C-1):

embedded image

- R is hydrogen, optionally substituted acyl, optionally substituted alkyl, optionally substituted alkenyl, optionally substituted alkynyl, optionally substituted carbocyclyl, or optionally substituted aryl;
- R^Z1is hydrogen, optionally substituted acyl, optionally substituted alkyl, optionally substituted alkenyl, optionally substituted alkynyl, optionally substituted carbocyclyl, or optionally substituted aryl;
- R^Z2is hydrogen, optionally substituted acyl, optionally substituted alkyl, optionally substituted alkenyl, optionally substituted alkynyl, optionally substituted carbocyclyl, or optionally substituted aryl;
- or optionally, R^Z1and R^Z2are taken together with their intervening atoms to form an optionally substituted carbocyclic ring;
- R^3Ais hydrogen, optionally substituted acyl, optionally substituted alkyl, optionally substituted alkenyl, or optionally substituted alkynyl;
- R^3Bis hydrogen, optionally substituted acyl, optionally substituted alkyl, optionally substituted alkenyl, or optionally substituted alkynyl;
- R^Yis hydrogen, optionally substituted acyl, optionally substituted alkyl, optionally substituted alkenyl, or optionally substituted alkynyl;
- R^Zis hydrogen, optionally substituted acyl, optionally substituted alkyl, optionally substituted alkenyl, or optionally substituted alkynyl.

In certain embodiments, a cannabinoid compound is of Formula (X-A-1):

embedded image

wherein custom-character is a double bond, and each of R^Z1and R^Z2is hydrogen, one of R^3Aand R^3Bis optionally substituted C_2-6alkenyl, and the other one of R^3Aand R^3Bis optionally substituted C_2-6alkyl. In some embodiments, a cannabinoid compound of Formula (X-1) is of Formula (X-A-1), wherein each of R^Z1and R^Z2is hydrogen, one of R^3Aand R^3Bis a prenyl group, and the other one of R^3Aand R^3Bis optionally substituted methyl.

In certain embodiments, a cannabinoid compound of Formula (X-1) of Formula (X-A-1) is of Formula (11-z-1):

embedded image

wherein custom-character is a double bond or single bond, as valency permits; one of R^3Aand R^3Bis C_1-6alkyl optionally substituted with alkenyl, and the other of R^3Aand R^3Bis optionally substituted C_1-6alkyl. In certain embodiments, in a compound of Formula (11-z-1), is a single bond; one of R^3Aand R^3Bis C_1-6alkyl optionally substituted with prenyl; and the other of one of R^3Aand R^3Bis unsubstituted methyl; and R is as described in this application. In certain embodiments, in a compound of Formula (11-z-1), custom-character is a single bond; one of R^3Aand R^3Bis

embedded image

and the other of one of R^3Aand R^3Bis unsubstituted methyl; and R is as described in this application. In certain embodiments, a cannabinoid compound of Formula (11-z-1) is of Formula (11a-1):

embedded image

In certain embodiments, a cannabinoid compound of Formula (X-1) of Formula (X-A-1) is of Formula (11a-1):

embedded image

In certain embodiments, a cannabinoid compound of Formula (X-A-1) is of Formula (10-z-1):

embedded image

wherein custom-character is a double bond or single bond, as valency permits; R^Yis hydrogen, optionally substituted acyl, optionally substituted alkyl, optionally substituted alkenyl, or optionally substituted alkynyl; and each of R^3Aand R^3Bis independently optionally substituted C_1-6alkyl. In certain embodiments, in a compound of Formula (10-z-1), custom-character is a single bond; each of R^3Aand R^3Bis unsubstituted methyl, and R is as described in this application. In certain embodiments, a cannabinoid compound of Formula (10-z-1) is of Formula (10a-1):

embedded image

In certain embodiments, a compound of Formula (10a-1)

embedded image

has a chiral atom labeled with * at carbon 10 and a chiral atom labeled with ** at carbon 6. In certain embodiments, in a compound of Formula (10a-1)

embedded image

the chiral atom labeled with * at carbon 10 is of the R-configuration and a chiral atom labeled with ** at carbon 6 is of the R-configuration. In certain embodiments, a compound of Formula (10a-1)

embedded image

is of the formula:

embedded image

In certain embodiments, in a compound of Formula (10a-1)

embedded image

the chiral atom labeled with * at carbon 10 is of the S-configuration and a chiral atom labeled with ** at carbon 6 is of the S-configuration. In certain embodiments, a compound of Formula (10a-1)

embedded image

is of the formula:

embedded image

In certain embodiments, a cannabinoid compound is of Formula (X-B-1):

embedded image

In certain embodiments, a cannabinoid compound is of Formula (X-B-1):

embedded image

wherein custom-character is a double bond; R^Yis hydrogen, optionally substituted acyl, optionally substituted alkyl, optionally substituted alkenyl, or optionally substituted alkynyl; and each of R^3Aand R^3Bis independently optionally substituted C_1-6alkyl or optionally substituted C_1-6alkenyl. In certain embodiments, in a compound of Formula (X-B-1), R^Yis optionally substituted C_1-6alkyl; one of R^3Aand R^3Bis

embedded image

and the other one of R^3Aand R^3Bis unsubstituted methyl, and R is as described in this application. In certain embodiments, a compound of Formula (X-B-1) is of Formula (9a-1):

embedded image

In certain embodiments, a compound of Formula (9a-1)

embedded image

has a chiral atom labeled with * at carbon 3 and a chiral atom labeled with ** at carbon 4. In certain embodiments, in a compound of Formula (9a-1)

embedded image

the chiral atom labeled with * at carbon 3 is of the R-configuration and a chiral atom labeled with ** at carbon 4 is of the R-configuration. In certain embodiments, a compound of Formula (9a-1)

embedded image

is of the formula:

embedded image

In certain embodiments, in a compound of Formula (9a-1)

embedded image

the chiral atom labeled with * at carbon 3 is of the S-configuration and a chiral atom labeled with ** at carbon 4 is of the S-configuration. In certain embodiments, a compound of Formula (9a-1)

embedded image

is of the formula:

embedded image

In certain embodiments, a cannabinoid compound is of Formula (X-C-1):

embedded image

wherein R^Zis optionally substituted alkyl or optionally substituted alkenyl. In certain embodiments, a compound of Formula (X-C-1) is of formula:

embedded image

wherein a is 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10. In certain embodiments, a is 1. In certain embodiments, a compound of Formula (8′-1) is the same as a compound of Formula (8′-1). In certain embodiments, a is 2. In certain embodiments, a is 3. In certain embodiments, a is 1, 2, or 3 for a compound of Formula (X-C-1). In certain embodiments, a cannabinoid compound is of Formula (X-C-1), and a is 1, 2, 3, 4, or 5. In certain embodiments, a compound of Formula (X-C-1) is of Formula (8a-1):

embedded image

In some embodiments, cannabinoids of the present disclosure comprise cannabinoid receptor ligands. Cannabinoid receptors are a class of cell membrane receptors in the G protein-coupled receptor superfamily. Cannabinoid receptors include the CB₁receptor and the CB₂receptor. In some embodiments, cannabinoid receptors comprise GPR18, GPR55, and PPAR. (See Bram et al. “Activation of GPR18 by cannabinoid compounds: a tale of biased agonism” Br J Pharmcol v171 (16) (2014); Shi et al. “The novel cannabinoid receptor GPR55 mediates anxiolytic-like effects in the medial orbital cortex of mice with acute stress” Molecular Brain 10, No. 38 (2017); and O'Sullvan, Elizabeth. “An update on PPAR activation by cannabinoids” Br J Pharmcol v. 173(12) (2016)).

In some embodiments, cannabinoids comprise endocannabinoids, which are substances produced within the body, and phytocannabinoids, which are cannabinoids that are naturally produced by plants of genus Cannabis. In some embodiments, phytocannabinoids comprise the acidic and decarboxylated acid forms of the naturally-occurring plant-derived cannabinoids, and their synthetic and biosynthetic equivalents.

Over 94 phytocannabinoids have been identified to date (Berman, Paula, et al. “A new ESI-LC/MS approach for comprehensive metabolic profiling of phytocannabinoids in Cannabis.” Scientific reports 8.1 (2018): 14280; El-Alfy et al., 2010, “Antidepressant-like effect of delta-9-tetrahydrocannabinol and other cannabinoids isolated from Cannabis sativa L”, Pharmacology Biochemistry and Behavior 95 (4): 434-42; Rudolf Brenneisen, 2007, Chemistry and Analysis of Phytocannabinoids, Citti, Cinzia, et al. “A novel phytocannabinoid isolated from Cannabis sativa L. with an in vivo cannabimimetic activity higher than Δ9-tetrahydrocannabinol: Δ9-Tetrahydrocannabiphorol.” Sci Rep 9 (2019): 20335, each of which is incorporated by reference in this application in its entirety). In some embodiments, cannabinoids comprise Δ⁹-tetrahydrocannabinol (THC) type (e.g., (−)-trans-delta-9-tetrahydrocannabinol or dronabinol, (+)-trans-delta-9-tetrahydrocannabinol, (−)-cis-delta-9-tetrahydrocannabinol, or (+)-cis-delta-9-tetrahydrocannabinol), cannabidiol (CBD) type, cannabigerol (CBG) type, cannabichromene (CBC) type, cannabicyclol (CBL) type, cannabinodiol (CBND) type, or cannabitriol (CBT) type cannabinoids, or any combination thereof (see, e.g., R Pertwee, ed, Handbook of Cannabis (Oxford, UK: Oxford University Press, 2014)), which is incorporated by reference in this application in its entirety). A non-limiting list of cannabinoids comprises: cannabiorcol-C1 (CBNO), CBND-C1 (CBNDO), Δ⁹-trans-Tetrahydrocannabiorcolic acid-C1 (Δ⁹-THCO), Cannabidiorcol-C1 (CBDO), Cannabiorchromene-C1 (CBCO), (−)-Δ⁸-trans-(6aR,10aR)-Tetrahydrocannabiorcol-C1 (Δ⁸-THCO), Cannabiorcyclol C1 (CBLO). CBG-C1 (CBGO), Cannabinol-C2 (CBN-C2), CBND-C2, Δ⁹-THC-C2, CBD-C2, CBC-C2, Δ⁸-THC-C2, CBL-C2, Bisnor-cannabielsoin-C1 (CBEO), CBG-C2, Cannabivarin-C3 (CBNV), Cannabinodivarin-C3 (CBNDV), (−)-Δ⁹-trans-Tetrahydrocannabivarin-C3 (Δ⁹-THCV), (−)-Cannabidivarin-C3 (CBDV), (t)-Cannabichromevarin-C3 (CBCV), (−)-Δ⁸-trans-THC-C3 (Δ⁷-THCV), (±)-(1aS,3aR,8bR,8cR)-Cannabicyclovarin-C3 (CBLV), 2-Methyl-2-(4-methyl-2-pentenyl)-7-propyl-2H-1-benzopyran-5-ol, Δ⁷-tetrahydrocannabivarin-C3 (Δ⁷-THCV), CBE-C2, Cannabigerovarin-C3 (CBGV), Cannabitriol-C1 (CBTO), Cannabinol-C4 (CBN-C4), CBND-C4, (−)-Δ⁹-trans-Tetrahydrocannabinol-C4 (Δ⁹-THC-C4), Cannabidiol-C4 (CBD-C4), CBC-C4, (−)-trans-Δ⁸-THC-C4, CBL-C4, Cannabielsoin-C3 (CBEV), CBG-C4, CBT-C2, Cannabichromanone-C3, Cannabiglendol-C3 (OH-iso-HHCV-C3), Cannabioxepane-C5 (CBX), Dehydrocannabifuran-C5 (DCBF). Cannabinol-C5 (CBN), Cannabinodiol-C5 (CBND), (−)-Δ⁹-trans-Tetrahydrocannabinol-C5 (Δ⁹-THC), (−)-Δ⁸-trans-(6aR,10aR)-Tetrahydrocannabinol-C5 (Δ⁸-THC), (+)-Cannabichromene-C5 (CBC), (−)-Cannabidiol-C5 (CBD), (+)-(1aS,3aR,8bR,8cR)-CannabicyclolC5 (CBL), Cannabicitran-C5 (CBR), (−)-Δ⁹-(6aS,10aR-cis)-Tetrahydrocannabinol-C5 ((−)-cis-Δ⁹-THC), (−)-Δ⁷-trans-(1R,3R,6R)-Isotetrahydrocannabinol-C5 (trans-isoΔ⁷-THC), CBE-C4, Cannabigerol-C5 (CBG), Cannabitriol-C3 (CBTV), Cannabinol methyl ether-C5 (CBNM), CBNDM-C5, 8-OH—CBN-C5 (OH-CBN), OH-CBND-C5 (OH-CBND), 10-Oxo-Δ^6a(10a)-Tetrahydrocannabinol-C5 (OTHC), Cannabichromanone D-C5, Cannabicoumaronone-C5 (CBCON-C5), Cannabidiol monomethyl ether-C5 (CBDM), Δ⁹-THCM-C5, (±)-3″-hydroxy-Δ⁴″-cannabichromene-C5, (5aS,6S,9R,9aR)-Cannabielsoin-C5 (CBE), 2-geranyl-5-hydroxy-3-n-pentyl-1,4-benzoquinone-C5, 5-geranyl olivetolic acid, 5-geranyl olivetolate, 8α-Hydroxy-Δ⁹-Tetrahydrocannabinol-C5 (8α-OH-Δ⁹-THC), 8β-Hydroxy-Δ⁹-Tetrahydrocannabinol-C5 (8β-OH-Δ⁹-THC), 10α-Hydroxy-Δ⁸-Tetrahydrocannabinol-C5 (10α-OH-Δ″-THC), 10β-Hydroxy-Δ⁸-Tetrahydrocannabinol-C5 (10β-OH-Δ⁸-THC), 10α-hydroxy-Δ^9,11-hexahydrocannabinol-C5, 9β,10β-Epoxyhexahydrocannabinol-C5, OH-CBD-C5 (OH-CBD), Cannabigerol monomethyl ether-C5 (CBGM), Cannabichromanone-C5, CBT-C4, (±)-6,7-cis-epoxycannabigerol-C5, (±)-6,7-trans-epoxycannabigerol-C5, (−)-7-hydroxycannabichromane-C5, Cannabimovone-C5, (−)-trans-Cannabitriol-C5 ((−)-trans-CBT), (+)-trans-Cannabitriol-C5 ((+)-trans-CBT), (±)-cis-Cannabitriol-C5 ((±)-cis-CBT), (−)-trans-10-Ethoxy-9-hydroxy-Δ^6a(10a))-tetrahydrocannabivarin-C3 [(−)-trans-CBT-OEt], (−)-(6aR,9S,10S,10aR)-9,10-Dihydroxyhexahydrocannabinol-C5 [(−)-Cannabiripsol] (CBR), Cannabichromanone C-C5, (−)-6a,7,10a-Trihydroxy-Δ⁹-tetrahydrocannabinol-C5 [(−)-Cannabitetrol] (CBTT), Cannabichromanone B-C5, 8,9-Dihydroxy-Δ^6a(10a)-tetrahydrocannabinol-C5 (8,9-Di-OHCBT), (±)-4-acetoxycannabichromene-C5, 2-acetoxy-6-geranyl-3-n-pentyl-1,4-benzoquinone-C5, 11-Acetoxy-Δ 9-TetrahydrocannabinolC5 (11-OAc-Δ 9-THC), 5-acetyl-4-hydroxycannabigerol-C5, 4-acetoxy-2-geranyl-5-hydroxy-3-npentylphenol-C5, (−)-trans-10-Ethoxy-9-hydroxy-Δ^6a(10a)-tetrahydrocannabinol-C5 ((−)-trans-CBTOEt), sesquicannabigerol-C5 (SesquiCBG), carmagerol-C5, 4-terpenyl cannabinolate-C5, β-fenchyl-Δ⁹-tetrahydrocannabinolate-C5, α-fenchyl-Δ⁹-tetrahydrocannabinolate-C5, epi-bornyl-Δ⁹-tetrahydrocannabinolate-C5, bornyl-Δ⁹-tetrahydrocannabinolate-C5, α-terpenyl-Δ⁹-tetrahydrocannabinolate-C5, 4-terpenyl-Δ⁹-tetrahydrocannabinolate-C5, 6,6,9-trimethyl-3-pentyl-6H-dibenzo[b,d]pyran-1-ol, 3-(1,1-dimethylheptyl)-6,6a,7,8,10,10a-hexahydro-1-hydroxy-6,6-dimethyl-9H-dibenzo[b,d]pyran-9-one, (−)-(3S,4S)-7-hydroxy-Δ⁶-tetrahydrocannabinol-1,1-dimethylheptyl, (+)-(3S,4S)-7-hydroxy-Δ⁶-tetrahydrocannabinol-1,1-dimethylheptyl, 11-hydroxy-Δ⁹-tetrahydrocannabinol, and Δ⁸-tetrahydrocannabinol-11-oic acid)); certain piperidine analogs (e.g., (−)-(6S,6aR,9R,10aR)-5,6,6a,7,8,9,10,10a-octahydro-6-methyl-3-[(R)-1-methyl-4-phenylbutoxy]-1,9-phenanthridinediol 1-acetate)), certain aminoalkylindole analogs (e g., (R)-(+)-[2,3-dihydro-5-methyl-3-(4-morpholinylmethyl)-pyrrolo[1,2,3-de]-1,4-benzoxazin-6-yl]-1-naphthalenyl-methanone), certain open pyran ring analogs (e.g, 2-[3-methyl-6-(1-methylethenyl)-2-cyclohexen-1-yl]-5-pentyl-1,3-benzenediol and 4-(1,1-dimethylheptyl)-2,3′-dihydroxy-6′alpha-(3-hydroxypropyl)-1′,2′,3′,4′,5′,6′-hexahydrobiphenyl, tetrahydrocannabiphorol (THCP), cannabidiphorol (CBDP), CBGP, CBCP, their acidic forms, salts of the acidic forms, dimers of any combination of the above, trimers of any combination of the above, polymers of any combination of the above, or any combination thereof.

A cannabinoid described in this application can be a rare cannabinoid. For example, in some embodiments, a cannabinoid described in this application corresponds to a cannabinoid that is naturally produced in conventional Cannabis varieties at concentrations of less than 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, 0.9%, 0.8%, 0.7%, 0.6%, 0.5%, 0.25%, or 0.1% by dry weight of the female flower. In some embodiments, rare cannabinoids include CBGA, CBGVA, THCVA, CBDVA, CBCVA, and CBCA. In some embodiments, rare cannabinoids are cannabinoids that are not THCA, THC, CBDA or CBD.

A cannabinoid described in this application can also be a non-rare cannabinoid.

In some embodiments, the cannabinoid is selected from the cannabinoids listed in Table 1.

TABLE 1

Non-limiting examples of cannabinoids according to the present disclosure.

embedded image

Δ⁹-Tetrahydro-

cannabinol

A⁹-THC-C₅

embedded image

Δ⁹-Tetrahydro-

cannabinol-C₄

Δ⁹-THC-C₄

embedded image

Δ⁹-Tetrahydro-

cannabivarin

Δ⁹-THCV-C₃

embedded image

Δ⁹-Tetrahydro-

cannabiorcol

Δ⁹-THCO-C₁

embedded image

(−)-(6aS,10aR)-Δ⁹-

Tetrahydro-

cannabinol

(−)-cis-Δ⁹-THC-C₅

embedded image

Δ⁹-Tetrahydro-

cannabinolic acid A

Δ⁹-THCA-C₅A

embedded image

Δ⁹-Tetrahydro-

cannabinolic acid B

Δ⁹-THCA-C₅B

embedded image

Δ⁹-Tetrahydro-

cannabinolic acid-C₄

A and/or B

Δ⁹-THCA-C₄A and/or B

embedded image

Δ⁹-Tetrahydro-

cannabivarinic acid

A

Δ⁹-THCVA-C₃A

embedded image

Δ⁹-Tetrahydro-

cannabiorcolic acid

A and/or B

Δ⁹-THCOA-C₁A

and/or B

embedded image

(−)-Δ⁸-trans-

(6aR,10aR)-

Δ⁸-Tetrahydro-

cannabinol

Δ⁸-THC-C₅

embedded image

(−)-Δ⁸-trans-

(6aR,10aR)-

Tetrahydro-

cannabinolic

acid A

Δ⁸-THCA-C₅A

embedded image

(−)-Cannabidiol

CBD-C5

embedded image

Cannabidiol

momomethyl ether

CBDM-C5

embedded image

Cannabidiol-C4

CBD-C4

embedded image

Cannabidiolic acid

CBDA-C5

embedded image

Cannabidivarinic acid

CBDVA-C3

embedded image

(−)-Cannabidivarin

CBDV-C3

embedded image

Cannabidiorcol

CBD-C1

embedded image

Cannabigerolic acid

A

(E)-CBGA-C₅A

embedded image

Cannabigerol

(E)-CBG-C₅

embedded image

Cannabigerol

monomethyl ether

(E)-CBGM-C₅A

embedded image

Cannabinerolic acid A

(Z)-CBGA-C₅A

embedded image

Cannabigerovarin

(E)-CBGV-C₃

embedded image

Cannabigerol

(E)-CBG-C₅

embedded image

Cannabigerolic acid

A

(E)-CBGA-C₅A

embedded image

Cannabigerolic acid A

monomethyl ether

(E)-CBGAM-C₅A

embedded image

Cannabigerovarinic

acid A

(E)-CBGVA-C₃A

embedded image

Cannabinolic acid A

CBNA-C5 A

embedded image

Cannabinol methyl

ether

CBNM-C5

embedded image

Cannabinol

CBN-C5

embedded image

Cannabinol-C4

CBN-C4

embedded image

Cannabivarin

CBN-C3

embedded image

Cannabinol-C2

CBN-C2

embedded image

Cannabiorcol

CBN-C1

embedded image

(±)-

Cannabichromene

CBC-C₅

embedded image

(±)-Cannabichromenic

acid A

CBCA-C₅A

embedded image

(±)-

Cannabivarichromene,

(±)-

Cannabichromevarin

CBCV-C₃

embedded image

(±)-Cannabichro-

mevarinic

acid A

CBCVA-C₃A

embedded image

(±)-

Cannabichromene

CBC-C₅

embedded image

(±)-

(1aS,3aR,8bR,8cR)-

Cannabicyclol

CBL-C₅

embedded image

(±)-(1aS,3aR,8bR,8cR)-

Cannabicyclolic acid A

CBLA-C₅A

embedded image

(±)-(1aS,3aR,8bR,8cR)-

Cannabicyclovarin

CBLV-C₃

embedded image

(−)-(9R,10R)-trans-

10-O-Ethyl-

cannabitriol

(−)-trans-CBT-OEt-

C5

embedded image

(±)-

(9R,10R/9S,10S)-

Cannabitriol-C3

(±)-trans-CBT-C3

embedded image

(−)-(9R,10R)-trans-

Cannabitriol

(−)-trans-CBT-C5

embedded image

(+)-(9S,10S)-

Cannabitriol

(+)-trans-CBT-C5

embedded image

(±)-(9R,10S/9S,10R)-

Cannabitriol

(±)-cis-CBT-C5

embedded image

(−)-6a,7,10a-

Trihydroxy-

Δ9-

tetrahydrocannabinol

(−)-Cannabitetrol

embedded image

10-Oxo-Δ6a(10a)-

tetrahydro-

cannabinol

OTHC

embedded image

8,9-Dihydroxy-

Δ6a(10a)-

tetrahydro-

cannabinol

8,9-Di-OH-CBT-C5

embedded image

Cannabidiolic acid A

cannabitriol ester

CBDA-C5 9-OH-CBT-C5

ester

embedded image

(−)-(6aR,9S,10S,10aR)-

9,10-Dihydroxy-

hexahydrocannabinol,

Cannabiripsol

Cannabiripsol-C5

embedded image

(5aS,65,9R,9aR)-

Cannabielsoic acid B

CBEA-C5 B

embedded image

(5aS,6S,9R,9aR)-

C3-Cannabielsoic

acid B

CBEA-C3 B

embedded image

(5aS,6S,9R,9aR)-

Cannabielsoin

CBE-C5

embedded image

(5aS,6S,9R,9aR)-

C3-Cannabielsoin

CBE-C3

embedded image

(5aS,6S,9R,9aR)-

Cannabielsoic acid A

CBEA-C5 A

embedded image

Cannabiglendol-C3

OH-iso-HHCV-C3

embedded image

Dehydro-

cannabifuran

DCBF-C5

embedded image

Cannabifuran

CBF-C5

embedded image

Cannabidiphorol

(CBDP)

embedded image

Tetrahydro-

cannabiphorol

(THCP)

text missing or illegible when filed

Cannabinoids are often classified by “type,” i.e., by the topological arrangement of their prenyl moieties (See, for example, M. A. Elsohly and D. Slade, Life Sci., 2005, 78, 539-548; and L. O. Hanus et al. Nat. Prod. Rep., 2016, 33, 1357). Generally, each “type” of cannabinoid includes the variations possible for ring substitutions of the resorcinol moiety at the position meta to the two hydroxyl moieties. As used in this disclosure, a “CBG-type” cannabinoid is a 3-[(2E)-3,7-dimethylocta-2,6-dienyl]-2,4-dihydroxybenzoic acid optionally substituted at the 6 position of the benzoic acid moiety. As used in this disclosure, “CBC-type” cannabinoids refer to 5-hydroxy-2-methyl-2-(4-methylpent-3-enyl)-chromene-6-carboxylic acid optionally substituted at the 7 position of the chromene moiety. As used in this disclosure, a “THC-type” cannabinoid is a (6aR,10aR)-1-hydroxy-6,6,9-trimethyl-6a,7,8,10a-tetrahydrobenzo[c]chromene-2-carboxylic acid optionally substituted at the 3 position of the benzo[c]chromene moiety. As used in this disclosure, a “CBD-type” cannabinoid is a 2,4-dihydroxy-3-[(1R,6R)-3-methyl-6-prop-1-en-2-ylcyclohex-2-en-1-yl]-benzoic acid optionally substituted at the 6 position of the benzoic acid moiety. In some embodiments, the optional ring substitution for each “type” is an optionally substituted C1-C11 alkyl, an optionally substituted C1-C11 alkenyl, an optionally substituted C1-C11 alkynyl, or an optionally substituted C1-C11 aralkyl.

The terms “varinolic cannabinoid” and “varin cannabinoid” are interchangeable, and mean a cannabinoid that is a derivative of divaric acid or divarinol, a cannabinoid of Formula (X) where R1 is propyl (e.g., n-propyl), a cannabinoid of Formula (X-A), (X-B), (X-C), (11-z), (10-z), where R is propyl (e.g., n-propyl), or any combination of thereof. Exemplary, varinolic cannabinoids and varin cannabinoids include, but are not limited to, CBGV, CBCV (cannabichromevarin), CBDV, CBGVA, THCV, THCVA and/or CBCVA.

Biosynthesis of Cannabinoids and Cannabinoid Precursors

Aspects of the present disclosure provide tools, sequences, and methods for the biosynthetic production of cannabinoids in host cells. In some embodiments, the present disclosure teaches expression of enzymes that are capable of producing cannabinoids by biosynthesis.

As a non-limiting example, one or more of the enzymes depicted in FIG. 2 may be used to produce a cannabinoid or cannabinoid precursor of interest. FIG. 1 shows a cannabinoid biosynthesis pathway for the most abundant phytocannabinoids found in Cannabis. See also, de Meijer et al. I, II, III, and IV (1: 2003, Genetics, 163:335-346; II: 2005. Euphytica, 145:189-198, III: 2009, Euphytica, 165:293-311; and IV: 2009, Euphytica, 168:95-112), and Carvalho et al. “Designing Microorganisms for Heterologous Biosynthesis of Cannabinoids” (2017) FEMS Yeast Research June 1; 17(4), each of which is incorporated by reference in this application in its entirety. FIG. 4 shows a biosynthetic pathway for production of varin cannabinoid compounds.

It should be appreciated that a precursor substrate for use in cannabinoid biosynthesis is generally selected based on the cannabinoid of interest. Non-limiting examples of cannabinoid precursors include compounds of Formulae (1)-(8) in FIG. 2. In some embodiments, polyketides, including compounds of Formula (5), could be prenylated. In certain embodiments, the precursor is a precursor compound shown in FIGS. 1-4. Substrates in which R contains 1-40 carbon atoms are preferred. In some embodiments, substrates in which R contains 3-8 carbon atoms are most preferred.

As used in this application, a cannabinoid or a cannabinoid precursor may comprise an R group. See, e.g., FIG. 2. In some embodiments, R may be a hydrogen. In certain embodiments, R is optionally substituted alkyl. In certain embodiments, R is optionally substituted C1-40 alkyl. In certain embodiments, R is optionally substituted C2-40 alkyl. In certain embodiments, R is optionally substituted C2-40 alkyl, which is straight chain or branched alkyl. In certain embodiments, R is optionally substituted C3-8 alkyl. In certain embodiments, R is optionally substituted C1-C40 alkyl, C1-C20 alkyl, C1-C10 alkyl, C1-C8 alkyl, C1-C5 alkyl, C3-C5 alkyl, C3 alkyl, or C5 alkyl. In certain embodiments, R is optionally substituted C1-C20 alkyl. In certain embodiments, R is optionally substituted C1-C10 alkyl. In certain embodiments. R is optionally substituted C1-C8 alkyl. In certain embodiments, R is optionally substituted C1-C5 alkyl. In certain embodiments, R is optionally substituted C1-C7 alkyl. In certain embodiments. R is optionally substituted C3-C5 alkyl. In certain embodiments, R is optionally substituted C3 alkyl. In certain embodiments, R is unsubstituted C3 alkyl. In certain embodiments, R is n-C3 alkyl. In certain embodiments, R is n-propyl. In certain embodiments, R is n-butyl. In certain embodiments, R is n-pentyl. In certain embodiments, R is n-hexyl. In certain embodiments, R is n-heptyl. In certain embodiments. R is of formula:

embedded image

In certain embodiments, R is optionally substituted C4 alkyl. In certain embodiments, R is unsubstituted C4 alkyl. In certain embodiments, R is optionally substituted C5 alkyl. In certain embodiments, R is unsubstituted C5 alkyl. In certain embodiments, R is optionally substituted C6 alkyl. In certain embodiments, R is unsubstituted C6 alkyl. In certain embodiments. R is optionally substituted C7 alkyl. In certain embodiments, R is unsubstituted C7 alkyl. In certain embodiments R is of formula:

embedded image

In certain embodiments, R is of formula:

embedded image

In certain embodiments, R is of formula:

embedded image

In certain embodiments, R is of formula:

embedded image

In certain embodiments, R is of formula:

embedded image

In certain embodiments, R is optionally substituted n-propyl. In certain embodiments, R is n-propyl optionally substituted with optionally substituted aryl. In certain embodiments, R is n-propyl optionally substituted with optionally substituted phenyl. In certain embodiments. R is n-propyl substituted with unsubstituted phenyl. In certain embodiments, R is optionally substituted butyl. In certain embodiments, R is optionally substituted n-butyl. In certain embodiments, R is n-butyl optionally substituted with optionally substituted aryl. In certain embodiments, R is n-butyl optionally substituted with optionally substituted phenyl. In certain embodiments, R is n-butyl substituted with unsubstituted phenyl. In certain embodiments, R is optionally substituted pentyl. In certain embodiments, R is optionally substituted n-pentyl. In certain embodiments, R is n-pentyl optionally substituted with optionally substituted aryl. In certain embodiments, R is n-pentyl optionally substituted with optionally substituted phenyl. In certain embodiments, R is n-pentyl substituted with unsubstituted phenyl. In certain embodiments, R is optionally substituted hexyl. In certain embodiments, R is optionally substituted n-hexyl. In certain embodiments, R is optionally substituted n-heptyl. In certain embodiments, R is optionally substituted n-octyl. In certain embodiments, R is alkyl optionally substituted with aryl (e.g., phenyl). In certain embodiments, R is optionally substituted acyl (e.g., —C(═O)Me).

In certain embodiments, R is optionally substituted alkenyl (e.g., substituted or unsubstituted C_2-6alkenyl). In certain embodiments, R is substituted or unsubstituted C_2-6alkenyl. In certain embodiments R is substituted or unsubstituted C_2-5alkenyl. In certain embodiments, R is of formula:

embedded image

In certain embodiments, R is optionally substituted alkynyl (e.g., substituted or unsubstituted C_2-6alkynyl). In certain embodiments, R is substituted or unsubstituted C_2-6alkynyl. In certain embodiments, R is of formula:

embedded image

In certain embodiments, R is optionally substituted carbocyclyl. In certain embodiments, R is optionally substituted aryl (e.g., phenyl or napthyl).

The chain length of a precursor substrate can be from C1-C40. Those substrates can have any degree and any kind of branching or saturation or chain structure, including, without limitation, aliphatic, alicyclic, and aromatic. In addition, they may include any functional groups including hydroxy, halogens, carbohydrates, phosphates, methyl-containing or nitrogen-containing functional groups.

For example, FIG. 3 shows a non-exclusive set of putative precursors for the cannabinoid pathway. Aliphatic carboxylic acids including four to eight total carbons (“C4”-“C8” in FIG. 3) and up to 10-12 total carbons with either linear or branched chains may be used as precursors for the heterologous pathway. Non-limiting examples include methanoic acid, butyric acid, pentanoic acid, hexanoic acid, heptanoic acid, isovaleric acid, octanoic acid, and decanoic acid. Additional precursors may include ethanoic acid and propanoic acid. In some embodiments, in addition to acids, the ester, salt, and acid forms may all be used as substrates. Substrates may have any degree and any kind of branching, saturation, and chain structure, including, without limitation, aliphatic, alicyclic, and aromatic. In addition, they may include any functional modifications or combination of modifications including, without limitation, halogenation, hydroxylation, amination, acylation, alkylation, phenylation, and/or installation of pendant carbohydrates, phosphates, sulfates, heterocycles, or lipids, or any other functional groups.

Substrates for any of the enzymes disclosed in this application may be provided exogenously or may be produced endogenously by a host cell. In some embodiments, the cannabinoids are produced from a glucose substrate, so that compounds of Formula 1 shown in FIG. 2 and CoA precursors are synthesized by the cell. In other embodiments, a precursor is fed into the reaction. In some embodiments, a precursor is a compound selected from Formulae 1-8 in FIG. 2.

Cannabinoids produced by methods disclosed in this application include rare cannabinoids. Due to the low concentrations at which cannabinoids, including rare cannabinoids occur in nature, producing industrially significant amounts of isolated or purified cannabinoids from the Cannabis plant may become prohibitive due to, e.g., the large volumes of Cannabis plants, and the large amounts of space, labor, time, and capital requirements to grow, harvest, and/or process the plant materials (see, for example, Crandall, K., 2016. A Chronic Problem: Taming Energy Costs and Impacts from Marijuana Cultivation. EQ Research; Mills, E., 2012. The carbon footprint of indoor Cannabis production. Energy Policy, 46, pp. 58-67: Jourabchi. M. and M. Lahet. 2014. Electrical Load Impacts of Indoor Commercial Cannabis Production. Presented to the Northwest Power and Conservation Council: O'Hare, M., D. Sanchez, and P. Alstone. 2013. Environmental Risks and Opportunities in Cannabis Cultivation. Washington State Liquor and Cannabis Board: 2018. Comparing Cannabis Cultivation Energy Consumption. New Frontier Data; and Madhusoodanan, J., 2019. Can cannabis go green? Nature Outlook: Cannabis; all of which are incorporated by reference in this disclosure). The disclosure provided in this application represents a potentially efficient method for producing high yields of cannabinoids, including rare cannabinoids. The disclosure provided in this application also represents a potential method for addressing concerns related to agricultural practices and water usage associated with traditional methods of cannabinoid production (Dillis et al. “Water storage and irrigation practices for cannabis drive seasonal patterns of water extraction and use in Northern California.” Journal of Environmental Management 272 (2020): 110955, incorporated by reference in this disclosure).

Cannabinoids produced by the disclosed methods also include non-rare cannabinoids. Without being bound by a particular theory, the methods described in this application may be advantageous compared with traditional plant-based methods for producing non-rare cannabinoids. For example, methods provided in this application represent potentially efficient means for producing consistent and high yields of non-rare cannabinoids. With traditional methods of cannabinoid production, in which cannabinoids are harvested from plants, maintaining consistent and uniform conditions, including airflow, nutrients, lighting, temperature, and humidity, can be difficult. For example, with plant-based methods, there can be microclimates created by branching, which can lead to inconsistent yields and by-product formation. In some embodiments, the methods described in this application are more efficient at producing a cannabinoid of interest as compared to harvesting cannabinoids from plants. For example, with plant-based methods, seed-to-harvest can take up to half a year, while cutting-to-harvest usually takes about 4 months. Additional steps including drying, curing, and extraction are also usually needed with plant-based methods. In contrast, in some embodiments, the fermentation-based methods described in this application only take about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 days. In some embodiments, the fermentation-based methods described in this application only take about 3-5 days. In some embodiments, the fermentation-based methods described in this application only take about 5 days. In some embodiments, the methods provided in this application reduce the amount of security needed to comply with regulatory standards. For example, a smaller secured area may be needed to be monitored and secured to practice the methods described in this application as compared to the cultivation of plants. In some embodiments, the methods described in this application are advantageous over plant-sourced cannabinoids.

Prenyltransferase (PT)

A host cell described in this application may comprise a prenyltransferase (PT). As used in this disclosure, a “PT” refers to an enzyme that is capable of transferring prenyl groups to acceptor molecule substrates. Non-limiting examples of prenyltransferases are described in U.S. Pat. No. 7,544,498 and Kumano et al., Bioorg Med Chem. 2008 Sep. 1; 16(17): 8117-8126 (e.g., NphB), PCT Publication No. WO 2018/200888 (e.g., CsPT4), U.S. Pat. No. 8,884,100 (e.g., CsPT1); Canadian Patent No. CA2718469; Valliere et al., Nat Commun. 2019 Feb. 4:10(1):565 (e.g., NphB variants); PCT Publication Nos: WO2019/173770, WO2019/183152, and WO2020/210810 (e.g., NphB variants); Luo et al., Nature 2019 March; 567(7746):123-126 (e.g., CsPT4); WO 2021/034848; U.S. 63/091,292, U.S. 63/188,442 and WO2022/081615 (e.g., CsPT variants and chimeras), which are incorporated by reference in their entireties. In some embodiments, a PT is capable of producing cannabigerolic acid (CBGA), cannabigerophorolic acid (CBGPA), cannabigerovarinic acid (CBGVA), or other cannabinoids or cannabinoid-like substances. In some embodiments, a PT is capable of producing cannabigerol (CBG), cannabigerovarin (CBGV), or other cannabinoids or cannabinoid-like substances. In some embodiments, a PT is cannabigerolic acid synthase (CBGAS). In some embodiments, a PT is cannabigerovarinic acid synthase (CBGVAS).

Example 1 describes the identification of a PT from Phialocephala scopiformis (P. scopiformis; corresponding to UniProt Accession No. A0A132B7II) that can be functionally expressed in host cells such as S. cerevisiae. The protein sequence corresponding to UniProt Accession No. A0A132B7I1 is provided in this disclosure as SEQ ID NO: 34: MKRKSTIEPFSADRLLSDLEHISNSIKAPYSPQAVQEALRVFGENLSNGAIAIRT TNRAGDPLNFWAGEYNRADTISRAVNAGIVSFTHPTVLLLRSWFSMYDNEPE PSTDFDTVYGL AKTWIYFMRLRPVEEVLSAEHVPQSFRDHIDTFKSIGARLVY HVAVNYRSNSVNVYLQIPSEFNPKQATKVVTTLLPDCVPPTAIEMEQMVKCM KPDMPIVFAVTLAYPSGTIERICFYAFMVPKELALSMGIGERLETFLRETPCYD EREVINFGWSFGRTGDRYLKIDTGYCGGFCDILGKLKHN* (SEQ ID NO: 34) A non-limiting example of a nucleic acid sequence encoding SEQ ID NO: 34 is SEQ ID NO: 35.

atgaaacgtaagtctaccatagaaccattttccgccgatagattgctttcggacttagagcacatcagtaatagcattaaggctccttattc accccaggcagtgcaagaagctctaagagtittcggtgaaaacttgtctaacggagctattgctatcaggacaactaatagagccggtg atccactgaacttctgggctggcgaatacaatagagccgacacgatctctcgtgctgtcaacgcaggtattgtttcctttactcatccaac cgtcttgttgttaagatcttggttctccatgtacgataacgagccagaaccttctactgactttgataccgtatatggtttggctaagacctgg atttacttcatgagattaagaccagttgaagaagttttgagtgccgaacacgttccacaatcgtttagagatcatatagacactttcaaatca attggtgctcgtttggtctaccacgtcgctgtgaattacaggtctaactccgttaatgtatatcttcaaatcccatctgagttcaacccaaag caagcaactaaggtcgttacaacgttgctaccagactgcgttcctcctactgctattgaaatggaacaaatggttaaatgtatgaagcca gacatgcctatcgtcttcgccgttacactagcttacccatcaggtaccatcgaaagaatatgtttttatgcttttatggtaccaaaggaatta gccttgtctatgggcattggtgaaagattggaaactttcttgagagaaaccccctgttacgatgagcgtgaagtcattaatttcggttggtc ctttggtagaactggtgatagatatctaaaaatcgacaccggttactgcggtggtttctgtgacatcctgggaaagttaaagcataactaa (SEQ ID NO: 35)

In some embodiments, a PT comprises a sequence (nucleic acid or protein sequence) that is at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or is 100% identical, including all values in between, to SEQ ID NO: 34 or SEQ ID NO: 35. In some embodiments, a PT comprises a conservatively substituted version of SEQ ID NO: 34.

In some embodiments, a PT consists of a sequence corresponding to SEQ ID: 34.

A host cell that expresses a heterologous polynucleotide encoding a PT described in this disclosure may be capable of producing at least 1% (e.g., at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 100%, at least 125%, at least 150%, at least 175%, at least 200%, at least 300%, at least 400%, at least 500%, at least 600%, at least 700%, at least 800%, at least 900%, or at least 1,000%) more CBG than a host cell that expresses a control PT. A host cell that expresses a heterologous polynucleotide encoding a PT described in this disclosure may be capable of producing at least 5, 10, 15, 20 or more than 20 fold more CBG relative to a host cell that expresses a control PT.

In some embodiments, the control PT is a wild-type reference PT. A wild-type reference PT can be full-length or truncated. A wild-type reference PT can be part of a fusion protein.

In some embodiments, a control PT corresponds to NphB from Streptomyces sp. (see, e.g., UniprotKB Accession No. Q4R2T2; see also SEQ ID NO: 2 of U.S. Pat. No. 7,361,483). The protein sequence corresponding to UniprotKB Accession No. Q4R2T2 is provided by SEQ ID NO: 8:

(SEQ ID NO: 8)

MSEAADVERVYAAMEEAAGLLGVACARDKIYPLLSTFQDTLVEGGSVVVF

SMASGRHSTELDFSISVPTSHGDPYATVVEKGLFPATGHPVDDLLADTQK

HLPVSMFAIDGEVTGGFKKTYAFFPTDNMPGVAELSAIPSMPPAVAENAE

LFARYGLDKVQMTSMDYKKRQVNLYFSELSAQTLEAESVLALVRELGLHV

PNELGLKFCKRSFSVYPTLNWETGKIDRLCFAVISNDPTLVPSSDEGDIE

KFHNYATKAPYAYVGEKRTLVYGLTLSPKEEYYKLGAYYHITDVQRGLLK

AFDSLED.

A non-limiting example of a nucleotide

sequence encoding NphB is:

(SEQ ID NO: 9)

atgtcagaagccgcagatgtcgaaagagtttacgccgctatggaagaagc

cgccggtttgttaggtgttgcctgtgccagagataagatctacccattgt

tgtctacttttcaagatacattagttgaaggggttcagttgttgttttct

ctatggcttcaggtagacattctacagaattggatttctctatctcagtt

ccaacatcacatggtgatccatacgctactgttgttgaaaaaggtttatt

tccagcaacaggtcatccagttgatgatttgttggctgatactcaaaagc

atttgccagtttctatgtttgcaattgatggtgaagttactggtggtttc

aagaaaacttacgctttctttccaactgataacatgccaggtgttgcaga

attatctgctattccatcaatgccaccagctgttgcagaaaatgcagaat

tatttgctagatacggtttggataaggttcaaatgacatctatggattac

aagaaaagacaagttaatttgtacttttctgaattatcagcacaaacttt

ggaagctgaatcagttttggcattagttagagaattgggtttacatgttc

caaacgaattgggtttgaagttttgtaaaagatctttctcagtttatcca

actttaaactgggaaacaggcaagatcgatagattatgtttcgcagttat

ctctaacgatccaacattggttccatcttcagatgaaggtgatatcgaaa

agtttcataactacgctactaaagcaccatatgcttacgttggtgaaaag

agaacattagtttatggtttgactttatcaccaaaggaagaatactacaa

gttgggtgcttactaccacattaccgacgtacaaagaggtttattgaaag

cattcgatagtttagaagactaa.

In other embodiments, a control PT corresponds to CsPT1, which is disclosed as SEQ ID NO: 2 in U.S. Pat. No. 8,884,100 (Cannabis sativa; corresponding to SEQ ID NO: 10 in this disclosure):

(SEQ ID NO: 10)

MGLSSVCTFSFQTNYHTLLNPHNNNPKTSLLCYRHPKTPIKYSYNNFPSK

HCSTKSFHLQNKCSESLSIAKNSIRAATTNQTEPPESDNHSVATKILNFG

KACWKLQRPYTIIAFTSCACGLFGKELLHNTNLISWSLMFKAFFFLVAIL

CIASFTTTINQIYDLHIDRINKPDLPLASGEISVNTAWIMSIIVALFGLI

ITIKMKGGPLYIFGYCFGIFGGIVYSVPPFRWKQNPSTAFLLNFLAHIIT

NFTFYYASRAALGLPFELRPSFTFLLAFMKSMGSALALIKDASDVEGDTK

FGISTLASKYGSRNLTLFCSGIVLLSYVAAILAGIIWPQAFNSNVMLLSH

AILAFWLILQTRDFALTNYDPEAGRRFYEFMWKLYYAEYLVYVFI.

In some embodiments, a control PT corresponds to CsPT4, which is disclosed as SEQ ID NO: 110 in WO2018200888, corresponding to SEQ ID NO: 11 in this disclosure:

(SEQ ID NO: 11)

MGLSLVCTFSFQTNYHTLLNPHNKNPKNSLLSYQHPKTPIIKSSYDNFPS

KYCLTKNFHLLGLNSHNRISSQSRSIRAGSDQIEGSPHHESDNSIATKIL

NFGHTCWKLQRPYVVKGMISIACGLFGRELFNNRHLFSWGLMWKAFFALV

PILSFNFFAAIMNQIYDVDIDRINKPDLPLVSGEMSIETAWILSIIVALT

GLIVTIKLKSAPLFVFIYIFGIFAGFAYSVPPIRWKQYPFTNFLITISSH

VGLAFTSYSATTSALGLPFVWRPAFSFIIAFMTVMGMTIAFAKDISDIEG

DAKYGVSTVATKLGARNMTFVVSGVLLLNYLVSISIGIIWPQVFKSNIMI

LSHAILAFCLIFQTRELALANYASAPSRQFFEFIWLLYYAEYFVYVFI.

In some embodiments, a control PT corresponds to a truncated CsPT4, which is provided as SEQ ID NO: 12 in this disclosure:

(SEQ ID NO: 12)

MSAGSDQIEGSPHHESDNSIATKILNFGHTCWKLQRPYVVKGMISIACGL

FGRELFNNRHLFSWGLMWKAFFALVPILSFNFFAAIMNQIYDVDIDRINK

PDLPLVSGEMSIETAWILSIIVALTGLIVTIKLKSAPLFVFIYIFGIFAG

FAYSVPPIRWKQYPFTNFLITISSHVGLAFTSYSATTSALGLPFVWRPAF

SFIIAFMTVMGMTIAFAKDISDIEGDAKYGVSTVATKLGARNMTFVVSGV

LLLNYLVSISIGIIWPQVFKSNIMILSHAILAFCLIFQTRELALANYASA

PSRQFFEFIWLLYYAEYFVYVFI

PTs for use in producing cannabinoids may be selected based on any one or more desired features, such as substrate selectivity, potential products formed, yield/titer of a product of interest, and/or solubility (cytosolic localization) of the enzyme.

a. Substrate Selectivity

Many prenyltransferases are known to have promiscuity in regard to prenyl donors and acceptors, which may result in a broad spectrum of potential products formed using a particular enzyme (Chen et al. Nat. Chem. Biol. (2017): 13(2): 226-234). Without being bound by a particular theory, promiscuous enzymes may be useful in some embodiments because different products may be produced by the enzyme by varying the substrate. In some embodiments, a promiscuous enzyme may be useful in producing different products from a composition of heterogenous substrates.

In other instances, it may be preferable for the prenyltransferase to have high specificity and not be promiscuous. For example, it may be preferable for the prenyltransferase to be specific for a particular substrate, so that the prenyltransferase produces a more homogenous product mix (i.e., greater product purity). Without being bound by a particular theory, an enzyme that has high specificity for a particular substrate may be useful because it may reduce possible by-products due to impurities in the substrate composition. For instance, when an enzyme is used with a host cell, the host cell may have intracellular mechanisms to convert a particular feed substrate into an undesirable substrate. In such instances, an enzyme that is highly specific for the non-converted substrate may be used to produce a product that has a higher purity of a compound of interest. In some instances, a highly specific enzyme may be useful for simplifying downstream processing, e.g., removing the need for further product purification.

As a non-limiting example, the PT from Streptomyces sp., NphB, has been previously shown to prenylate both olivetol and olivetolic acid (Kuzuyama et al. Nature, 2005). Wild-type NphB has also been reported to display a high degree of both substrate and product promiscuity. Similarly, C. sativa CsPT4 has been previously shown to prenylate both olivetol and olivetolic acid (Luo et al. Nature, 2019).

However, at least the Streptomyces sp. aromatic prenyltransferase NphB has been reported to have poor kinetics with respect to prenylation of olivetol (Kumano et al, Bioorg. Med. Chem., 2008). Particularly, NphB has been shown to be incapable of efficiently utilizing olivetol for producing CBG. The consumption of olivetol by NphB may not be sufficient for meaningful production of CBG and for downstream cannabinoid biosynthesis.

Surprisingly, the inventors of the present disclosure identified an aromatic PT from P. scopiformis which has significant activity against olivetol and which is capable of prenylayting olivetol to form CBG. The discovery allows efficient utilization of olivetol, which has been viewed as a “dead-end” metabolite in the cannabinoid biosynthesis pathway.

In some embodiments, as shown in FIG. 5, a PT is capable of catalyzing a compound of Formula 5:

embedded image

or as shown in FIG. 2, is capable of catalyzing a compound of Formula 5a:

embedded image

to produce a compound of Formula 8a-1:

embedded image

In some embodiments, a PT may be capable of consuming a substrate of a compound of Formula 5a in FIG. 5B at a rate that is at least 1% (e.g., at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 100%, at least 125%, at least 150%, at least 175%, at least 200%, at least 300%, at least 400%, at least 500% f, at least 600%, at least 700%, at least 800%, at least 900%, or at least 1,000%) faster or slower relative to a PT control.

In some embodiments, a PT may be capable of consuming olivetol (Formula 5a) at a rate that is at least 1% (e.g., at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 100%, at least 125%, at least 150%, at least 175%, at least 200%, at least 300%, at least 400%, at least 500%, at least 600%, at least 700% t, at least 800%, at least 900%, or at least 1,000%) faster relative to a PT control. In some embodiments, the PT comprises a sequence that is at least 90% identical to SEQ ID NO: 34. In some embodiments, the PT comprises the sequence of SEQ ID NO: 34.

In some embodiments, a PT may be capable of consuming at least 1% (e.g., at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 100%, at least 125%, at least 150%, at least 175%, at least 200%, at least 300%, at least 400%, at least 500%, at least 600%, at least 700%, at least 800%, at least 900%, or at least 1,000%) more olivetol (Formula 5a) relative to a PT control. In some embodiments, the PT comprises a sequence that is at least 90% identical to SEQ ID NO: 34. In some embodiments, the PT comprises a sequence that corresponds to SEQ ID NO: 34.

In some embodiments, a PT may be capable of consuming at least 5,000 μg/L, at least 6,000 μg/L, at least 7,000 μg/L, at least 8,000 μg/L, at least 9.000 μg/L, at least 10,000 μg/L, at least 11,000 μg/L, at least 12,000 μg/L, at least 13,000 μg/L, at least 14,000 μg/L, at least 15,000 μg/L, at least 16,000 μg/L, at least 17,000 μg/L, at least 18,000 μg/L, at least 19,000 μg/L, at least 20,000 μg/L, at least 21,000 μg/L, at least 22,000 μg/L, at least 23,000 μg/L, at least 24.000 μg/L, at least 25,000 μg/L, at least 26,000 μg/L, at least 27,000 μg/L, at least 28,000 μg/L, at least 29,000 μg/L, at least 30,000 μg/L, at least 31,000 μg/L, at least 32,000 μg/L, at least 3300 μg/L, at least 34,000 μg/L, at least 35,000 μg/L, at least 36,000 μg/L, at least 37,000 μg/L, at least 38,000 μg/L, at least 39,000 μg/L, or at least 40,000 μg/L more olivetol (Formula 5a) relative to a PT control. In some embodiments, the PT comprises a sequence that is at least 90% identical to SEQ ID NO: 34. In some embodiments, the PT comprises a sequence that corresponds to SEQ ID NO: 34.

In some embodiments, the control is a wild-type reference PT. A wild-type reference PT can be full-length or truncated. A wild-type reference PT can be part of a fusion protein. In some embodiments, the PT control is NphB (SEQ ID NO: 8). See, e.g., U.S. Pat. No. 7,544,498; and Kumano et al., Bioorg Med Chem. 2008 Sep. 1; 16(17): 8117-8126, which are incorporated by reference in this application in their entireties.

b. Prenylation

In addition to promiscuity in regard to potential substrates utilized, many prenyltransferases are known to also be promiscuous as to the products formed due to the ability to prenylate a prenyl acceptor at different sites, further resulting in a broad spectrum of potential products formed using a particular enzyme (Chen et al. Nat. Chem. Biol. (2017): 13(2): 226-234). When tested for activity using geranyl pyrophosphate (GPP) and olivetolic acid (OA) as substrates, NphB and CsPT4 produce multiple prenylation products (Kumano et al. Bioorganic Medicinal Chemistry, 2008; Luo et al. Nature, 2019). In particular, on OA at carbon positions labeled 3 and 5 and oxygen positions labeled 2 and 4 in Structure 6a (FIG. 5). Zirpel et al. reported the major prenylation product of wild-type NphB to be 2-O-Geranyl Olivetolic Acid (OGOA, Formula (8b) in FIG. 5)), with CBGA produced as the minor product (Formula (8a) in FIG. 1 and FIG. 5, Zirpel et al. Journal of Biotechnology, 2017). Functional expression of NphB and production of CBGA in S. cerevisiae was detected (Zirpel et al. Journal of Biotechnology, 2017).

The carboxyl group of olivetolic acid has been described as “crucial for the [geranyl-olivetolic acid transferase] reaction” in the biosynthesis of cannabinoids in planta (see, Taura et al, 2007. Phytocannabinoids in Cannabis sativa: recent studies on biosynthetic enzymes. Chemistry & Biodiversity, 4(8), pp. 1649-1663 at 1659.). Thus, olivetol has been considered a “dead-end” metabolite, where no downstream products can be produced in a conventional cannabinoid biosynthesis pathway (FIG. 1). Olivetol, therefore, has not been frequently used as a substrate for creating prenylation products in this pathway.

In some instances, it may be preferable to prenylate at a particular position in Formula (6) or Formula (5). For example, it may be preferable to use a prenyltransferase (e.g., in combination with a terminal synthase) to produce phytocannabinoids, which are commonly prenylated at the C3 position of Formula (6).

In some instances, prenylation at a particular position in Formula (6) or Formula (5) may be used to alter the pharmacokinetic profile of cannabinoid products. For example, prenylation at a particular position in Formula (6) or Formula (5) may allow for the development of a cannabinoid product that crosses the blood brain barrier.

In some embodiments, a PT described in this disclosure transfers one or more prenyl groups to any of positions 1, 2, or 3 in a compound of Formula (5), shown below:

embedded image

In some embodiments, a PT described in this disclosure transfers one or more prenyl groups to position 3 in a compound of Formula (5), shown below:

embedded image

In some embodiments, a PT described in this disclosure transfers one or more prenyl groups to any of positions 1, 2, or 3 in a compound of Formula (5), shown below:

embedded image

to form one or more compounds of Formula (8w-1-a), Formula (8x-1), and/or Formula (8′-1):

embedded image

In some embodiments, a PT described in this disclosure transfers a prenyl group to a compound of Formula (5), shown below:

embedded image

to form a compound of Formula (8-1):

embedded image

In some embodiments, as shown in FIG. 5, a PT described in this disclosure transfers a prenyl group to a compound of Formula (5a), shown below:

embedded image

to form a compound of Formula (8a-1):

embedded image

In some embodiments, provided is a method for producing a prenylated product of a compound of Formula (5a):

embedded image

comprising contacting:

(a) a compound of Formula (5a):

embedded image

in the presence of (b) a PT comprising a sequence that is at least 90% identical to the sequence of SEQ ID NO: 34. In some embodiments, the PT comprises the sequence of SEQ ID NO: 34.

In some embodiments, a PT described in this disclosure transfers one or more prenyl groups to any of positions 1, 2, 3, 4, or 5 in a compound of Formula (6), shown below:

embedded image

In some embodiments, the PT transfers a prenyl group to any of positions 1, 2, 3, 4, or 5 in a compound of Formula (6), shown below:

embedded image

to form a compound of one or more of Formula (8w), Formula (8x), Formula (8′), Formula (8y), Formula (8z):

embedded image

or a pharmaceutically acceptable salt, solvate, hydrate, polymorph, co-crystal, tautomer, stereoisomer, isotopically labeled derivative, or prodrug thereof, wherein a is 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10.

In some embodiments, the PT transfers a prenyl group to any of positions 1, 2, 3, 4, or 5 in a compound of Formula (6), shown below:

embedded image

to form a compound of one or more of Formula (8w), Formula (8x), Formula (8′), Formula (8y), Formula (8z), wherein a is 1, 2, 3, 4, or 5. In some embodiments, the PT transfers a prenyl group to any of positions 1, 2, 3, 4, or 5 in a compound of Formula (6), shown below:

embedded image

to form a compound of one or more of Formula (8w), Formula (8x), Formula (8′), Formula (8y), Formula (8z), or a pharmaceutically acceptable salt thereof, wherein a is 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10.

In some embodiments, a PT described in this application transfers one or more prenyl groups to any of positions 1, 2, or 3 in a compound of Formula (5), shown below:

embedded image

In some embodiments, the PT transfers a prenyl group to any of positions 1, 2, or 3 in a compound of Formula (5), shown below:

embedded image

to form one or more compounds of Formula (8w-1-a), Formula (8x-1), and/or Formula (8′-1):

embedded image

to form a compound of Formula (8-1):

embedded image

or a pharmaceutically acceptable salt, solvate, hydrate, polymorph, co-crystal, tautomer, stereoisomer, isotopically labeled derivative, or prodrug thereof.

In some embodiments, the PT catalyzes the synthesis of (e.g., by transferring a prenyl group to result in the synthesis of) a compound of Formula (8-1):

embedded image

or a pharmaceutically acceptable salt, solvate, hydrate, polymorph, co-crystal, tautomer, stereoisomer, isotopically labeled derivative, or prodrug thereof.

In some embodiments, the PT transfers a prenyl group to a compound of Formula (5a), shown below:

embedded image

to form a compound of Formula (8a-1):

embedded image

or a pharmaceutically acceptable salt, solvate, hydrate, polymorph, co-crystal, tautomer, stereoisomer, isotopically labeled derivative, or prodrug thereof.