Showing posts with label chemical space. Show all posts
Showing posts with label chemical space. Show all posts

Monday, 20 May 2024

A time and place for Nature in drug discovery?

I’ll be reviewing Y2022 (The Time and Place for Nature in Drug Discovery) in this post and stating my position on natural products in modern drug discovery is a good place to start. I certainly see value in screening natural products and natural product-like compounds (especially in phenotypic assays) and there is currently a great deal of interest in chemical probes (I’ll point you toward an article on the Target 2035 initiative and a link to the Chemical Probes Portal).  In general, a natural product or natural product-like active identified by screening would either need to exhibit novel phenotypic effects or be significantly more potent than other known actives for me to enthusiastic about following it up. I would certainly consider screening fragments that are only present in natural product structures although these would need to still need comply with the criteria (typically defined in terms of properties such as molecular size, molecular complexity and lipophilicity) used to select fragments. I see significant benefits coming from the increased use of  biocatalysis, both in drug discovery and for manufacturing drugs, but I don’t see these benefits as being restricted to synthesis of natural products or natural product-like compounds. 

This will be a very long post (for which I make no apology) and it's a good point to say something about how the review is presented. I've used section headings (in bold text) used in Y2022 for my commentary and quoted text has been indented (my comments on the quoted text enclosed with square brackets and italicized in red). I'd like to raise four general points before starting my review: 

  1. Proprietary data cannot accurately be described as “facts” or “evidence” and it’s not valid to claim that you’ve proven or demonstrated something on the basis of analysis of proprietary data.  
  2. If continuous data such as oral bioavailability measurements have been made categorical (e.g., high | medium | low) prior to analysis then it’s generally a safe assumption that any trends "revealed" by the analysis are weak.  
  3. If basing claims on analysis of locations or distributions within a particular chemical space it is necessary to demonstrate the chemical space is actually relevant to the claims being made. One way to do this is to build usefully predictive models of relevant quantities such as aqueous solubility or permeability using only the dimensions of the chemical space as descriptors.  
  4. There are generally many ways to partition a region of chemical space into subregions with different average values for a measured quantity. Although the  boundaries resulting from these analyses typically appear to be well-defined (for example, as a line or curve in a 2-dimensional chemical space) it is a serious error to automatically interpret such boundaries as meaningful from a physicochemical perspective.    

I have a number of concerns about the Y2022 article and I’ll focus on the more serious of these in this post. I’ll also be commenting on the Rule of 5 (Ro5; see L1997), logP/logD differences, and the drug discovery “sweet spot” reported in the HK2012 article.  My view is that a number of the assertions and recommendations made by the authors of Y2022 are not supported by the analyses or the data that they’ve presented. Specifically, the authors present results of analyses that had been performed using proprietary and undocumented models and, in my view, they have grossly over-interpreted the predictions made using the models.  At times, the authors appear to be treating natural products as if these occupy a distinct and contiguous region of chemical space (this is a pitfall into which drug-likeness advocates also frequently stumble).  The authors of Y2022 discuss physicochemical properties at considerable length without making any convincing connection between this discussion and natural products. Reading the Y2022 article, I did detect a subliminal message that natural products might be infused with vital force and wouldn’t have been surprised to see Gwyneth Paltrow as a co-author.

I’ll make some general observations before examining Y2022 in detail. If you’re going to base decisions on trends in data then you need to now how strong the trends are because this tells you how much weight to give to the trends when making your decisions. In what I’ll call the ‘compound quality’ field you’ll often encounter data presentations that make it extremely difficult to see how strong (or weak) the trends in the data actually are (see KM2013: Inflation of correlation in the pursuit of drug-likeness). Since Ro5 was introduced in 1997 (see L1997) there has been a free flow of advice from self-appointed compound quality gurus as to how compounds can be made better, more developable and more beautiful (introduction of the term “Ro5 envy” in KM2013 appeared to cause some to spit feathers). This advice frequently comes in the form of dire warnings that exceeding a threshold value of a property, such as molecular weight or predicted octanol/water partition coefficient, will increase the probability of something bad happening. It’s actually very difficult to set thresholds like these objectively and you have to consider the possibility that some of these statements of probability are merely expressions of belief (to some “there is a high probability that God exists” will sound rather more convincing than “I believe in God”).

The graphical abstract is a good place to start my review of Y2022. I don’t know whether biotransformations exist that would convert the Core Scaffold into compounds that would match the Bios Collection generalized structure but a 1,3-diene in conjugation with a tertiary nitrogen is not the sort of substructure that I would want to see in a screening active that I had been charged with optimizing.  

Abstract

The authors of Y2022 state:  

The declining natural product-likeness of licensed drugs and the consequent physicochemical implications of this trend in the context of current practices are noted. [The authors do not make a convincing connection between natural product-likeness and physicochemical properties.]  To arrest these trends, the logic of seeking new bioactive agents with enhanced natural mimicry is considered; notably that molecules constructed by proteins (enzymes) are more likely to interact with other proteins (e.g., targets and transporters), a notion validated by natural products. [I consider this claim to be extravagant and it does need to be supported by evidence. The authors’ use of “validated” reminded me of the extravagant claim made in a Future Medicinal Chemistry editorial that “ligand efficiency validated fragment-based design”. Taking the statement literally, the authors appear to be suggesting that a compound would be more likely to interact with proteins if it had been isolated from natural sources than if it had been synthesized in a laboratory (I was reminded of the "water memory" explanation for why homeopathy works). If “molecules constructed by proteins” really are more likely to interact with other proteins then they’re also more likely to interact with anti-targets like hERG and CYPs. I’m guessing that the response of medicinal chemistry teams tackling CNS targets to suggestions that they should make their compounds more like natural products so as increase the likelihood of recognition by transporters might be to ask which natural products those offering the advice had been smoking.]

Introduction

The authors show time-dependence for the values of a number of parameters calculated for drugs in Figure 1. I see analyses like these as exercises in philately and, when I first encountered examples about two decades ago, I formed a view that some senior medicinal chemists had a bit too much time on their hands. The observation of significant time-dependency for a parameter calculated for drugs can mean one of three things. First, the parameter is irrelevant to drug discovery (however, the absence of a time-dependence shouldn't be taken as evidence that the parameter is relevant to drug discovery). Second, the old ways were best and the medicinal chemists of today have lost their way (I’m guessing this might be Jacob Rees Mogg’s interpretation if he were a medicinal chemist). Third, the old ways no longer work so well and the medicinal chemists of today have learned new ways.

I have a number of concerns about what is shown in Figure 1 (quite aside from these concerns I would question why 1b or 1c were even included in the study). The data values that have been plotted are actually mean values and, as we observed in KM2013, the presentation of mean value (or median) values without showing measures of the spread in the data, such as standard deviation or inter-quartile range, makes trends look stronger than they actually are (others use the term “voodoo correlations”).  This way of presenting data is specifically verboten by J Med Chem and Author Guidelines (viewed 18-May-2024) for that journal specifically state:

If average values are reported from computational analysis, their variance must be documented. This can be accomplished by providing the number of times calculations have been repeated, mean values, and standard deviations (or standard errors). Alternatively, median values and percentile ranges can be provided. Data might also be summarized in scatter plots or box plots.

However, the hidden variation in the response variables is not the only issue that I have with Figure 1. Let’s take a look at Figure 1a which shows “a temporal comparison of natural product likeness of approved drugs assessed by the Natural Product Scout algorithm (12) versus the year of the first disclosure of the drug” although it the caption for Figure 1a is “Natural product class probability. (8)”. I think that the authors do need to explain exactly what they mean by natural product class probability because the true probability that a compound is a natural product is either 1 (it’s a natural product) or 0 (it’s not a natural product). Put another way there are differences between natural products and Prof.  Schrödinger’s unfortunate feline companion. The measure of lipophilicity shown in Figure 1c is XLogP3 although no justification is given for the selection of this particular method for lipophilicity prediction nor is any reference provided.

Before continuing with my review of Y2022 I also need to examine Ro5 and discuss the difference between logP and logD (the reasons for these digressions will hopefully become clear later). Ro5 which was based on physicochemical property distributions for compounds that had been taken into phase 2 of clinical development before 1997 (the year that L1997 was published). My view is that Ro5 certainly raised awareness of the problems associated with excessive lipophilicity and molecular size (A Good Thing) but I’ve never considered Ro5 to be useful in design. Although Ro5 is accepted by many (most?) drug discovery scientists as an article of faith, some are prepared to ask awkward questions and I’ll mention the S2019 study. Let’s take a look at how Ro5 was specified in the L1997 article (the graphic is slide #17 from a presentation that I gave late last year):


Ro5 is stated in terms of likelihood of poor absorption or permeation although no measured oral absorption or permeability data are given in the L1997 study and Ro5 should therefore be regarded as a statement of belief. I realise that to make such an assertion runs the risk of an appointment with the auto-da-fé and I stress that had Ro5 been stated in terms of physicochemical and molecular property distributions I would not have made the assertion.

Medieval cartographers annotated the unknown regions of their maps with “here be dragons” and Ro5’s dragons are poor absorption and poor permeation. However, there's another issue which I touched on in HBD3:

It is significant that attempts to build global models for permeability and solubility, using only the dimensions of the chemical space in which the Ro5 is specified as descriptors, do not appear to have been successful.

What I was getting at in HBD3 is that the chemical space in which Ro5 is specified was not demonstrated to be relevant to permeability or solubility (this relates to the third of the four points that I raised at the start of the post). It must be stressed that I'm definitely not denying that relationships exist between descriptors, such as logP, used to specify Ro5 and properties such as aqueous solubility and permeability that are more directly relevant to getting drugs to where they need to. It’s just that these relationships are weak (see TY2020) and, while we don’t exactly know exactly how weak the relationships are, we do know that they are weak because continuous data have been binned to display them (see also KM2013 and specifically the comments on HY2010). I would generally anticipate that these relationships will be stronger within structural series but in these cases you’ll generally observe different relationships for different structural series. In practical terms this means that a logP of 5 might be manageable in one structural series while in another structural series compounds with logP greater than 3.5 prove to be inadequately soluble. As I advised in NoLE:

Drug designers should not automatically assume that conclusions drawn from analysis of large, structurally-diverse data sets are necessarily relevant to the specific drug design projects on which they are working.

I also need to discuss the distinction between logP and logD since this is a source of confusion for medicinal chemists and compound quality 'experts' alike. Here’s a graphic (it’s slide #18) from the presentation that I did at SancaMedChem in 2019 (if the piranhas did venture into the non-polar phase they'd probably end up swimming backstroke):


The partition coefficient (P) is simply the ratio of the concentration of the neutral form of the compound in the organic phase (usually octanol) to the concentration of the compound in water when both phases are in equilibrium. The distribution coefficient (D) is defined analogously as the ratio of the sum of concentrations of all forms of the compound in the organic phase to the sum of concentrations of all forms of the compound in water. Values of P and D are usually quoted as their logarithms logP and logD. When interpreting logD values it is commonly assumed that that is that only neutral forms of compounds partition into organic phases and if we make this assumption the relationship between logD and logP is given by Eqn 1 (see B2017):

When we perform experiments to quantify lipophilicity it is actually logD that is measured. Values of logP and logD are identical when ionization can be neglected and logP values for ionizable compounds can be obtained by examination of measured logD-pH profiles although this is rarely done. It’s usually a safe assumption that logP values used by drug discovery scientists (and quoted in medicinal chemistry publications) have been predicted and these values vary with the method used for prediction of logP. For example, L1997 states that the upper logP limit for Ro5 is 5 when logP is calculated using the ClogP method (see L1993) but 4.15 when logP is calculated using the method of Moriguchi et al. (see M1992). Values of logD that you encounter in the literature may have been calculated or measured (you might need to dig around to see if you’re dealing with real data) and it’s also important to remember that logD depends on pH. I would argue that logD is less appropriate than logP for defining compound quality metrics because excessive lipophilicity can be countered simply by increasing the extent to which compounds are are ionized (I hope you can see why that would be A Bad Thing). Another way to think about this is to consider an amine with a pKa value of 8 bound to hERG at a pH of 7. Now suppose that you can change the pKa of the amine to 11 without changing anything else in the molecular structure. What effects would you expect this pKa change to have on affinity, on logD and on logP?

I’ll now get back to reviewing Y2022 and let’s take a look at Figure 2 which shows an adapted version of the "drug discovery sweet spot” proposed in the HK2012 study. As with Figure 1b and 1c, I would question why Figure 2 was included in the Y2022 study since the connection with natural products is tenuous. In my view the authors of the HK2012 study made a number of serious errors in their definition of the “sweet spot” and these errors have been reproduced in the Y2022 study. The authors of HK2012 claimed to have identified a “drug discovery sweet spot” in a chemical space defined by “Log P” and “Molecular mass” but they didn’t actually demonstrate that this chemical space is actually relevant to drug discovery (one way to demonstrate relevance is to build convincing global models for prediction of properties like permeability and aqueous solubility using only the dimensions of the chemical space as descriptors).

If claiming to have identified a drug discovery “sweet spot” it’s important that each dimension of the chemical space in which the “sweet spot” corresponds to a single entity. While “Molecular mass” is unambiguous the term “Log P” does not refer to the same entity for each of the data sets from which the “sweet spot” has been derived. As noted previously ClogP (see L1993) was used to specify Ro5 while the Gleeson upper Log P limit (see G2008) and the “μM potency Log P” (see G2011) were specified respectively by values of clogP (calculated logP from ACD) and AlogP (no reference provided). In contrast the Pfizer Golden Triangle (see J2009) is specified using elogD (proprietary logD prediction method for which details were ot provided).  The Waring low and high logP/logD values stated in W2010 are at least partly based on analysis of AZlogD7.4 values (proprietary logD prediction method; details not provided) reported in the WJ2007 and W2009 studies. The W2010 study states that “the optimal range of lipophilicity lies between ~ 1 and 3” but the these are not the values that are depicted in Figure 3 (or indeed in the original HK2012 study). The Gleeson upper limits for Log P and Molecular Mass stated in G2008 reflect the arbitrary schemes used to bin the data and should not be regarded as objectively-determined limits for these quantities. The authors of Y2022 have superimposed ellipses for "SHMs", "Antibiotic Space?" and "bRo5 /  AbbVie MPS space for higher MW" on the HK2012 "sweet spot" in the creation of Figure 2 although it is not clear how these ellipses were constructed.

The Physicochemical Characteristics of Drugs

The authors assert:

A principle advocated by Hansch that drug molecules should be made as hydrophilic as possible without loss of efficacy (47) is commonly expressed and utilized as Lipophilic Ligand Efficiency (LLE). (48) [If actually using this principle advocated by Hansch you would optimize leads by varying hydrophilicity and observing efficacy. While LLE is one way to express Hansch’s principle it is by no means the only way and (pIC50 – 0.5 ´ logP) would be equally acceptable as a lipophilic ligand efficiency metric from the perspective of the Hansch’s principle.] This metric, widely accepted and exploited in drug discovery as a key metric in optimization, is expressed on a log scale as activity (e.g., −log10[XC50]) [The logarithm function is not defined for dimensioned quantities such as XC50 (see M2011) and, while it may appear to nitpicking to point it out, this is the source of the invalidity of the ligand efficiency metric as was discussed at length in NoLE.] minus a lipophilicity term (typically the Partition coefficient or log10 P or sometimes log D7.4). (49) [Although it is common to see LLE values quoted in the drug discovery literature it’s much less clear how (or even whether) the metric was actually used to make project decisions. In many studies, however, the focus is on plots of pIC50 against logP (or logD) rather than values of the metric itself. In lead optimization, medicinal chemists typically need to balance activity against properties such as permeability, aqueous solubility, metabolic stability and off-target activity. In these situations, experienced medicinal chemists typically give much more weight to structure-activity relationships (SARs) and structure-property relationships (SPRs) that they've observed within the structural series that they're optimising than to crude metrics of questionable relevance and predictivity. It is noteworthy that the authors of ref 49 use logD rather than logP to define LLE (which they call LiPE) and if you do this then you can make compounds more efficient simply by increasing the extent to which they are ionized.] The impact of lipophilicity on efficacy needs to be considered in the context that reducing lipophilicity (equating to increasing hydrophilicity) will generally increase the solubility, reduce the metabolism, and reduce the promiscuity of a given compound in a series. (50) [The relationships between these properties and lipophilicity shown in ref 50 are for structurally diverse data sets rather than for individual series. I consider the activity criterion (pIC50 > 5) used to quantify promiscuity in ref 50 to be at least an order of magnitude too permissive to be pharmaceutically relevant.]

Let’s take a look at Figure 3 in which values of “Calc Chrom Log D7.4” are plotted against “CMR”. This is what the authors of say about Figure 3 in the text of Y2022:

The distribution of marketed oral drugs in terms of their lipophilicity and size, shows a remarkably similar distribution to the set of compounds designed by Kell as a representative set of natural products to investigate carrier mechanisms (Figure 3). (64) [To state “shows a remarkably similar distribution” is arm-waving given that there are methods for assessing the similarity of two distributions in an objective manner.]

As is the case for Figure 1a, what is written in the text about Figure 3 differs significantly from the caption for this figure:

Figure 3. Natural products are found across most size lipophilicity combinations, as exemplified in a representative set designed and compiled by O’Hagan and Kell (64) superimposed on the Chrom log D7.4 vs cmr training set of compounds with >30% bioavailability. (51) [It is unclear why this training set was restricted to compounds with >30% bioavailability.  The LDF is shown in this figure with “Limits of confidence” but the level of confidence to which these limits correspond is not given.]

The first criticism that I’ll make is that the authors of Y2022 have not actually demonstrated the relevance of chemical space specified by the axes of Figure 3 (this is the essence of the third of the four points that I raised at the start of the post and the same criticism can be made of Figure 4 and Figure 5). The authors note, with some arm-waving, that cmr “largely correlates with MW” which does rather beg the question of why they consider this particular measure of molecular size to be superior to MW for this type of analysis. The authors claim that “the GSK model based on log D7.4 vs calculated molar refraction” (it is actually molar refractivity as opposed to molar refraction that was calculated) is a useful guide to predict oral exposure. I consider this claim to be extravagant because one would need to have access to the proprietary model for calculation of Chrom Log D7.4 in order to use the model. The proprietary nature of the GSK model means that predictions made using this model cannot credibly be presented as “evidence”.

Details of the models for calculating Chrom Log D7.4 and for prediction of oral exposure are sketchy and I regard each of these proprietary models as undocumented. A linear discriminant function (LDF) model was reportedly used for prediction of oral exposure but it is unclear how the model was trained (or if it was even validated). An LDF is a classification model and it is not clear what how the classes were defined for prediction of oral exposure. I’m assuming that the oral absorption classes used in GSK oral exposure model have been defined by categorization of continuous data (I’m happy to be corrected on this point but, given the sketchiness of details, I can be forgiven for speculation) and setting thresholds like these is difficult to achieve in an objective manner. If this was indeed the case I'd assume that the threshold value used to categorize the continuous data was arbitrary (you’ll get a different LDF model if you use a different threshold to define the classes). My view is that that an LDF is an inappropriate way to model this type of data because the categorization of the data discards a huge amount of information.

Here's the caption for Figure 4:

Figure 4. Proposed regions of size/lipophilicity space for an oral drug set, (51) using the effectual combination of Chrom Log D7.4 vs calculated molar refraction (cmr) as a description of chemical space. [It’s actually molar refractivity as opposed to molar refractivity that was calculated. It is unclear what the authors mean by "bRo5 principles".] The highlighted regions suggest likely absorption mechanisms, based on ref (65) with compounds colored by binned NPScout probability scores. [The authors of Y2022 appear to be using a proprietary and undocumented LDF model of unknown predictivity to infer absorption mechanisms (this is what I was getting at in the fourth of the four points points that I raised at the start of the post). The depiction of data shown in Figure 4 would be much more informative had compounds known (as opposed to believed) by to be orally absorbed by one of these mechanisms been plotted in this chemical space.] Below the LDF line, then mean NPScout score is 0.45, (median 0.33) and above it (indicative of likely oral exposure) the mean is 0.31 and median 0.17 (p < 0.01) [It is unclear what (p < 0.01) refers to.]

Here's the caption for Figure 5: 

Figure 5. Illustration of antibiotic drug space, expressed as Calculated Chrom Log D7.4 vs cmr adapted from data in ref (65) colored by antibiotics (circles) and TB drugs (diamonds) which are sized by NP class probabilities and colored by prediction of likelihood of oral exposure (either side of the diagonal “linear discriminant function line” so to be oral, transporters a likely mechanism for the red colored compounds, which mostly have a high NPScout score). [As is the case for Figure 4, the authors of Y2022 appear to be using a proprietary and undocumented LDF model of unknown predictivity to infer absorption mechanisms. Stating that "mostly have a high NPScout score" is arm-waving.]  Vertical (cmr < 8) and horizontal lines (Chrom Log D7.4 < 2.5) together represent likely boundaries for paracellular absorption. [The basis (measured data or belief) for this assertion is unclear. The depiction of data shown in Figure 5 would have been more convincing had compounds known to be and known not to be absorbed by the paracellular route been plotted in this chemical space. While the problems of achieving good oral absorption for antibiotics should not be underestimated, I see getting compounds into cells as the bigger issue and in some cases the transporters cause active efflux (see R2021). The depiction of data shown in Figure 5 would have been much more informative had compounds known (as opposed to believed) to exhibit active influx and active efflux been plotted in this chemical space. Although Figure 5 is presented as a description of antibiotic drug space, the study (ref 65) on which Figure 5 is based is actually focused on antitubercular drug space (one of the challenges to discovery of antitubercular drugs is that Mycobacterium tuberculosis is an intracellular pathogen; see WL2012). One article that I recommend to all drug discovery scientists, especially those working on infectious diseases, is the SM2019 review on intracellular drug concentration.]

The authors suggest:

A logical extension of this hypothesis would be to consider recognition processes with natural molecules, which are likely to have discrete interactions with carrier proteins and therapeutic targets. [The authors do need to articulate what they mean by "discrete interactions" and why "natural molecules" are likely to have "discrete interactions" with carrier proteins and therapeutic targets.] Small molecule drugs are noted to be relatively promiscuous, so making interactions with several proteins is a likely event. (76) [This assertion is not supported by ref 76 which is actually a study of nuisance compounds, PAINS filters, and dark chemical matter in a proprietary compound collection. Promiscuity of a compound is typically defined by a count of the number of targets against which activity exceeds a specific threshold and promiscuity generally increases with the permissiveness of the activity threshold (it’s therefore meaningless to describe a compound as “promiscuous” without also stating the activity threshold). The activity threshold for the analysis reported in ref 76 is ³ 50% inhibition at a concentration of 10 µM which is appropriate if you’re worried about assay interference but, in my view, is at least an order of magnitude too permissive if considering the possibility of off-target activity for a drug in vivo.]  It similarly is logical to consider that a molecule made by a recognition process in a catalytic enzyme may also interact with another protein in a similar manner. (77) [This is not quite as logical as the authors would have us believe since enzymes catalyze reactions by stabilizing transition states. A high binding affinity of an enzyme for its reaction product would generally be expected to result in inhibition of the enzyme by the reaction product.]

Natural Product Fragments in Fragment-Based Drug Discovery

The authors note:

Fragment-based drug discovery (FBDD) can be employed to rapidly explore large areas of chemical space for starting points of molecular design. (91 | 92 | 93) However, most FBDD libraries are composed of privileged substructures of known synthetic drugs and drug candidates and populate already well-explored areas of chemical space, (94 | 95 | 96[I do not consider refs 94-96 to support this assertion (none of these three articles has a fragment screening library design focus and the most recent one was published in 2007).] often through the use of fragments with high sp2-character. (97)  Underexplored areas of chemical space can be rapidly explored by employing fragments derived from NPs that are already biologically prevalidated by evolution. [The authors appear to be suggesting that the physiological effects of natural products are more due to the fragments from which they have been constructed than of the way in which the fragments have been combined.] 

Molecular recognition

The authors state:

That the embedded recognition of natural products for proteins correlates with recognition of the biosynthetic enzyme is an increasingly validated concept. (118 | 119 | 120) [I have no idea what “embedded recognition” means and I’m guessing that the authors might be in a similar position.] The biosynthetic imprint translates to recognition of other proteins using similar interactions. [As I’ve already noted, high binding affinity of a natural product for the enzyme that catalysed its formation would lead to inhibition of the enzyme.] For example, the analysis of protein structures of 38 biosynthetic enzymes gave 64 potential targets for 25 natural products. (121) [Concepts are usually validated with measured data and not by making predictions.]

Conclusions and Prospects for Future Development

The authors assert:

More natural molecules will increase quality through their inherently improved permeability and solubility; [At the risk of appearing pedantic, permeability and solubility are properties of compounds as opposed to molecules. That said, the authors appear to be treating “natural molecules” as occupying a distinct and contiguous region of chemical space by making this claim and it is unclear what the improvements will be relative to. The authors do not present any measured data for permeability or solubility to support their claim.] this is a case of investing time and effort in the early stages of drug discovery to reap rewards with improvements in the later stages through more predictability in trials (and thus a greater chance of success, where quality rather than speed demonstrably impacts (170)) [Many, including me, do indeed believe that investing time and effort in the early stages of drug discovery increases the chances of success in the later stages. However, I would challenge the assertion by the authors of Y2022 that ref 170 actually demonstrates this.] and more sustainable manufacturing methods driven by the transformative power of biocatalysis. (171)

So that concludes my review of Y2022 and thanks for staying with me. I'll leave you with a selfie here in Trinidad's Maraval Valley with my faithful canine companions BB and Coco providing much-needed leadership (a few minutes earlier I had patiently explained to them why ligand efficiency is complete bollocks).

Wednesday, 27 September 2023

Five days in Vermont

A couple of months ago I enjoyed a visit to the US (my first for eight years) on which I caught up with old friends before and after a few days in Vermont (where a trip to the golf course can rapidly become a National Geographic Moment). One highlight of the trip was randomly meeting my friend and fellow blogger Ash Jogalekar for the first time in real life (we’ve actually known each other for about fifteen years) on the Boston T Red Line.  Following a couple of nights in green and leafy Belmont, I headed for the Flatlands with an old friend from my days in Minnesota for a Larry Miller group reunion outside Chicago before delivering a short harangue on polarity at Ripon College in Wisconsin. After the harangues, we enjoyed a number of most excellent Spotted Cattle (Only in Wisconsin) in Ripon. I discovered later that one of my Instagram friends is originally from nearby Green Lake and had taken classes at Ripon College while in high school. It is indeed a small world.

The five days spent discussing computer-aided drug design (CADD) in Vermont are what I’ll be covering in this post and I think it’s worth saying something about what drugs need to do in order to function safely.  First, drugs need to have significant effects on therapeutic targets without having significant effects on anti-targets such as hERG or CYPs and, given the interest in new modalities, I’ll be say “effects” rather than “affinity”, although Paul Ehrlich would have reminded us that drugs need to bind in order to exert effects. Second, drugs need to get to their targets at sufficiently high concentrations for their effects to be therapeutically significant (drug discovery scientists use the term ‘exposure’ when discussing drug concentration). Although it is sometimes believed that successful drugs simply reduce the numbers of patients suffering from symptoms it has been known from the days of Paracelsus that it is actually the dose that differentiates a drug from a poison.

Drug design is often said to be multi-objective in nature although the objectives are perhaps not as numerous as many believe (this point is discussed in the introduction section of NoLE, an article that I'd recommend to insomniacs everywhere). The first objective of drug design can be stated in terms of minimization of the concentration at which a therapeutically useful effect on the target is observed (this is typically the easiest objective to define since drug design is typically directed at specific targets). The second objective of drug design can be stated in analogous terms as maximization of the concentration at which toxic effects on the anti-targets are observed (this is a more difficult objective to define because we generally know less about the anti-targets than about the targets). The third objective of drug design is to achieve controllability of exposure (this is typically the most difficult objective to define because drug concentration is a dose-dependent, spaciotemporal quantity and intracellular concentration cannot generally be measured for drugs in vivo). Drug discovery scientists, especially those with backgrounds in computational chemistry and cheminformatics, don’t always appreciate the importance of controlling exposure and the uncertainty in intracellular concentration always makes for a good stock question for speakers and panels of experts.

I posted previously on  artificial intelligence (AI) in drug design and I think it’s worth highlighting a couple of common misconceptions. The first misconception is that we just need to collect enough data and the drugs will magically condense out of the data cloud that has been generated (this belief appears to have a number of adherents in Silicon Valley).  The second misconception is that drug design is merely an exercise in prediction when it should really be seen in a Design of Experiments framework. It’s also worth noting that genuinely categorical data are rare in drug design and my view is that many (most?) "global" machine learning (ML) models are actually ensembles of local models (this heretical view was expressed in a 2009 article and we were making the point that what appears to be an interpolation may actually be an extrapolation). Increasingly, ML is becoming seen as a panacea and it’s worth asking why quantitative structure activity relationship (QSAR) approaches never really made much of a splash in drug discovery.

I enjoyed catching up with old friends [ D | K | S | R/J | P/M ] as well as making some new ones [ G | B/R | L ]. However, I was disappointed that my beloved Onkel Hugo was not in attendance (I continue to be inspired by Onkel’s laser-like focus on the hydrogen bonding of the ester) and I hope that Onkel has finally forgiven me for asking (in 2008) if Austria was in Bavaria. There were many young people at the gathering in Vermont and their enthusiasm made me greatly optimistic for the future of CADD (I’m getting to the age at which it’s a relief not to be greeted with: "How nice to see you, I thought you were dead!"). Lots of energy at the posters (I learned from one that Voronoi was Ukrainian) although, if we’d been in Moscow, I’d have declined the refreshments and asked for a room on the ground floor (left photo below).  Nevertheless, the bed that folded into the wall (centre and right photos below) provided plenty of potential for hotel room misadventure without the ‘helping hands’ of NKVD personnel.

It'd been four years since CADD had been discussed at this level in Vermont so it was no surprise to see COVID-19 on the agenda. The COVID-19 pandemic led to some very interesting developments including the Covid Moonshot (a very different way of doing drug discovery and one I was happy to contribute to during my 19 month sojourn in Trinidad) and, more tangibly, Nirmatrelvir (an antiviral medicine that has been used to treat COVID-19 infections since early 2022). Looking at the molecular structure of Nirmatrelvir you might have mistaken trifluoroacetyl for a protecting group but it’s actually a important feature (it appears to be beneficial from the permeability perspective). My view is that the alkane/water logP (alkane is a better model than octanol for the hydrocarbon core of a lipid bilayer) for a trifluoroacetamide is likely to be a couple of log units greater than for the corresponding acetamide.



I’ll take you through how the alkane/water logP difference between a trifluoroacetamide and corresponding acetamide can be estimated in some detail because I think this has some relevance to using AI in drug discovery (I tend to approach pKa prediction in an analogous manner). Rather than trying to build an ML model for making the prediction, I’ve simply made connections between measurements for three different physicochemical properties (alkane/water logP, hydrogen bond basicity and hydrogen bond acidity) which is something that could easily be accommodated within an AI framework. I should stress that this approach can only be used because it is a difference in alkane/water logP (as opposed to absolute values) that is being predicted and these physicochemical properties can plausibly be linked to substructures.

Let’s take a look at the triptych below which I admit that is not quite up to the standards of Hieronymus Bosch (although I hope that you find it to be a little less disturbing). The first panel shows values of polarity (q) for some hydrogen bond acceptors and donors (you can find these in Tables 2 and 3 in K2022) that have been derived from alkane/water logP measurements. You could, for example, use these polarity values to predict that reducing the polarity of an amide carbonyl oxygen to the extent that it looks like a ketone will lead to a 2.2 log unit increase in alkane/water logP.  The second panel shows measured hydrogen bond basicity values for three hydrogen bond acceptors (you can find these in this freely available dataset) and the values indicate that a trifluoroacetamide is an even weaker hydrogen bond acceptor than a ketone. Assuming a linear relationship between polarity and hydrogen bond basicity, we can estimate that the trifluoroacetamide carbonyl oxygen is 2.4 log units less polar than the corresponding acetamide. The final panel shows measured hydrogen bond acidity values (you can find these in Table 1 of K2022) that suggest that an imide NH (q = 1.3; 0.5 log units more polar than typical amide NH) will be slightly more polar than the trifluoroacetamide NH of Nirmatrelvir. So to estimate he difference in alkane/water logP values you just need to subtract the additional polarity of trifluoroacetamide NH (0.5 log units) from the lower polarity of the trifluoroacetamide carbonyl oxygen (2.4) to get 1.9 log units.


Chemical space is a recurring theme in drug design and its vastness, which defies human comprehension, has inspired much navel-gazing over the years (it’s actually tangible chemical space that’s relevant to drug design). In drug discovery we need to be able to navigate chemical space (ideally without having to ingest huge quantities of Spice) and, given that Ukrainian chemists have revolutionized the world's idea of tangible chemical space (and have also made it a whole lot larger), it is most appropriate to have a Ukrainian guide who is most ably assisted by a trusty Transylvanian sidekick. I see benefits from considering molecular complexity more explicitly when mapping chemical space. 
   
AI (as its evangelists keep telling us) is quite simply awesome at generating novel molecular structures although, as noted in a previous post, there’s a little bit more to drug design than simply generating novel molecular structures. Once you’ve generated a novel molecular structure you need to decide whether or not to synthesize the compound and, in AI-based drug design, molecular structures are often assessed using ML models for biological activity as well as absorption, distribution, metabolism and excretion (ADME) behaviour. It’s well-known that you need a lot of data for training these ML models but you also need to check that the compounds for which you’re making predictions lie within the chemical space occupied by the training set (one way to do this is to ensure that close structural analogs of these compounds exist in the training set) because you can’t be sure that the big data necessarily cover the regions of chemical space of interest to drug designers using the models. A panel discusses the pressing requirement for more data although ML modellers do need to be aware that there’s a huge difference between assembling data sets for benchmarking and covering chemical space at sufficiently high resolution to enable accurate prediction for arbitrary compounds.  

There are other ways to think about chemical space. For example, differences in biological activity and ADME-related properties can also be seen in terms of structural relationships between compounds. These structural relationships can be defined in terms of molecular similarity (Tanimoto coefficient for the molecular fingerprints of X and Y is 0.9) or substructure (X is the 3-chloro analog of Y). Many medicinal chemists think about structure-activity relationships (SARs) and structure-property relationships (SPRs) in terms of matched molecular pairs (MMPs: pairs of molecular structures that are linked by specific substructural relationships) and free energy perturbation (FEP) can also be seen in this framework. Strong nonadditivity and activity cliffs (large differences in activity observed for close structural analogs) are of considerable interest as SAR features in their own right and because prediction is so challenging (and therefore very useful for testing ML and physics-based models for biological activity). One reason that drug designers need to be aware of activity cliffs and nonadditivity in their project data is that these SAR features can potentially be exploited for selectivity.
        
Cheminformatic approaches can also help you to decide how to synthesize the compounds that you (or your AI Overlords) have designed and automated synthetic route planning is a prerequisite for doing drug discovery in ‘self-driving’ laboratories. The key to success in cheminformatics is getting your data properly organized before starting analysis and the Open Reaction Database (ORD), an open-access schema and infrastructure for structuring and sharing organic reaction data, facilitates training of models. One area that I find very exciting is the use of high-throughput experimentation in the search for new synthetic reactions which can led to better coverage of unexplored chemical space. It’s well known in industry that the process chemists typically synthesize compounds by routes that differ from those used by the medicinal chemists and data-driven multi-objective optimization of catalysts can lead to more efficient manufacturing processes (a higher conversion to the desired product also makes for a cleaner crude product). 

It’s now time to wrap up what’s been a long post. Some of what is referred to as AI appears to already be useful in drug discovery (especially in the early stages) although non-AI computational inputs will continue to be significant for the foreseeable future. I see a need for cheminformatic thinking in drug discovery to shift from big data (global ML models) to focused data (generate project specific data efficiently for building local ML models) and also see advantages in using atom-based descriptors that are clearly linked to molecular interactions. One issue for data-driven approaches to prediction of biological activity such as ML and QSAR modelling is that the need for predictive capability is greatest when there's not much relevant data and this is a scenario under which physics-based approaches have an advantage. In my view, validation of ML models is not a solved problem since clustering in chemical space can cause validation procedures to make optimistic assessments of model quality. I continue to have significant concerns about how relationships (which are not necessarily linear) between descriptors are handled in ML modelling and remain generally skeptical of claims for interpretability of ML models (as noted in NoLE, the contribution of a protein–ligand contact to affinity is not, in general, an experimental observable).

Many thanks for staying with me to the end and hope to see many of you at EuroQSAR in Barcelona next year. I'll leave you with a memory from the early days of chemical space navigation.



Thursday, 17 January 2019

Thoughts on AI in Drug Discovery - A Practical View From the Trenches


I’ll be taking a look at machine learning (ML) in this post which was prompted by AI in Drug Discovery - A Practical View From the Trenches by Pat Walters in Practical Cheminformatics. Pat’s post appears to be triggered by Artificial Intelligence in Drug Design - The Storm Before the Calm? by Allan Jordan that was published as a viewpoint in ACS Medicinal Chemistry Letters. Some of what I said in the Nature of QSAR is relevant to what I’ll be saying in the current post and I'll also direct readers to Will CADD ever become relevant to drug discovery? by Ash at Curious Wavefunction.  Pat notes that Allan “fails to highlight specific problems or to define what he means by AI” and goes on to say that he prefers “to focus on machine learning (ML), a relatively well-defined subfield of AI”. Given that drug discovery scientists have been modeling activity and properties of compounds for decades now, some clarity would be welcome as to which of the methods used in the earlier work would fall under the ML umbrella.

While not denying the potential of AI and ML in drug design, I note that both are associated with a lot of hype and it would be an error to confuse skepticism about the hype with criticism of AI and ML. Nevertheless, there are some aspects of cheminformatic ML, such as chemical space coverage, that don't seem to get discussed quite as much as much as I think they should be and these are what the current post is focused on. I also have a suspicion that some of the ML approaches touted for drug design may be better suited for dealing with responses that are categorical (e.g. pIC50 > 6 ) rather than continuous (e.g. pIC50 = 6.6). When discussing ML in drug design, it can be useful to draw a distinction between 'direct applications' of ML (e.g. prediction of behavior of compounds) and 'indirect applications' of ML (e.g. synthesis planning; image analysis). This post is primarily concerned with direct applications of ML.

As has become customary, I’ve included some photos to break up the text a bit. These are all feature albatrosses and I took them on a 2009 visit to the South Island of New Zealand. Here's a live stream of nest at the Royal Albatross Centre on the Otago Peninsula.  

Spotted on Kaikoura whale watch

My comment on Pat’s post has just appeared so I’ll say pretty much what I said in that comment here. I would challenge the characterization of ML as “a relatively well-defined subfield of AI”. Typically, ML in cheminformatics focuses on (a) finding regions in descriptor space associated with particular chemical behaviors or (b) relating measures of chemical behavior to values of descriptors.  I would not automatically regard either of these activities as subfields of AI any more than I would regard Hansch QSAR, CoMFA, Free-Wilson Analysis, Matched Molecular Pair Analysis, Rule of 5 or PAINS filters as subfields of AI. I’m sure that there will be some cheminformatic ML models that can accurately be described as a subfield of AI but to tout each and every ML method as AI would be a form of hype.

At Royal Albatross Centre, Otago Peninsula.   

Pat states “In essence, machine learning can be thought of as ‘using patterns in data to label things’” and this could be taken as implying that ML models can only handle categorical responses. In drug design, the responses that we would like to predict using ML are typically continuous (e.g. IC50; aqueous solubility; permeability; fraction unbound; clearance; volume of distribution) and genuinely categorical data are rarely encountered in drug discovery projects. Nevertheless, it is common in drug discovery for continuous data to be made categorical (sometimes we say that the data has been binned). There are a number of reasons why this might not be such a great idea. First, binning continuous data throws away huge amounts of information. Second, binning continuous data distorts relationships between objects (e.g. a pIC50 activity threshold of 6 makes pIC50 = 6.1 appear to be more similar to pIC50 = 9 than to pIC50 = 5.9). Third, categorical analysis does not typically account for ordering (e.g. high | medium | low) of the categories. Fourth, one needs to show that the conclusions of analysis do not depend on how the continuous data has been categorized. The third and fourth issues are specifically addressed by the critique of Generation of a Set of Simple,Interpretable ADMET Rules of Thumb that was presented in Inflation of Correlation in the Pursuit of Drug-likeness.

Royal Albatross Centre, Otago Peninsula. 

Overfitting is always a concern when modelling multivariate data and the fit to the training data generally gets better when you use more parameters. One of my concerns with cheminformatic ML is that it is not always clear how many parameters have been used to build the models (I’m guessing that, sometimes, even the modelers don’t know) and one does need to account for numbers of parameters if claiming that one model has outperformed another. When building models from multivariate data, one also needs to account for relationships between the molecular descriptors that define the region(s) of chemical space occupied by the training set. In ‘traditional’ multivariate data analysis, it is assumed that relationships between descriptors are linear and modelers use principal component analysis (PCA) to determine the dimensions of the relevant regions of space. If relationships between descriptors are non-linear then life gets a lot more difficult. Another of my concerns with ML models is that it is not always clear how (or if) relationships between descriptors have been accounted for.

At Royal Albatross Centre, Otago Peninsula. 

Although an ML method may be generic and applicable to data from diverse sources, it is still useful to consider the characteristics of cheminformatic data that distinguish them from other types of data. As noted in Structure Modification in Chemical Databases, the molecular connection table (also known as the 2D molecular structure) is the defining data structure of cheminformatics. One characteristic of cheminformatic data is that is possible to make meaningful (and predictively useful) comparisons between structurally-related compounds and this provides a motivation for studying molecular similarity. In cheminformatic terms we can say that differences in chemical behavior can be perceived and modeled in terms of structural relationships between compounds. This can also be seen as a distance-geometric view of chemical space. Although this may sound a bit abstract, it’s actually how medicinal chemists tend to relate molecular structure to activity and properties (e.g. the bromo-substitution led to practically no improvement in potency but now it sticks like shit to the proverbial blanket in the plasma protein binding assay). This is also a useful framework for analysis of output from high-throughput screening (HTS) and design of screening libraries. It is typically difficult to perceive structural relationships between compounds using models based on generic molecular descriptors.

At Royal Albatross Centre, Otago Peninsula

I have been sufficiently uncouth as to suggest that many ‘global’ cheminformatic models may simply be ensembles of local models and this reflects a belief that training set compounds are often distributed unevenly in chemical space. As we move away from traditional Hansch QSAR to ML models, the molecular descriptors become more numerous (and less physical). When compounds are unevenly distributed in chemical space and molecular descriptors are numerous, it becomes unclear whether the descriptors are capturing the relevant physical chemistry or just organizing the compounds into groups of structurally related analogs. This is an important distinction and the following graphic (which does not feature an albatross) shows why. The graphic shows a simple plot of Y versus X and we want to use this to predict Y for X = 3.  If X is logP and Y is aqueous solubility then it would be reasonable to assume that X captures (at least some of) the physical chemistry and we would regard the prediction as an interpolation because X = 3 is pretty much at the center of this very simple chemical space. If X is simply partitioning the six compounds into two groups of structurally related analogs then making a prediction for X = 3 would represent an extrapolation. While this is clearly a very simple example, it does illustrate an issue that the cheminformatics community needs to take a bit more notice of.


Chemical space coverage is a key consideration for anyone using ML to predict activity and properties of for a series of structurally-related compounds. The term "Big Data" does tend to get over-used but being globally "big" is no guarantee that local regions of chemical space (e.g. the structural series that a medicinal chemistry team may be working on) are adequately covered. The difficulty for the chemists is that is they don't know whether their structural series is in a cluster in the training set space or in a hole. In cheminformatic terms, it is unclear whether or not the series that the medicinal chemistry team is working on lies within the applicability domain of the model.

Validation can lead to an optimistic view of model quality when training (and validation) sets are unevenly distributed in chemical space and I’ll ask you to have another look at Figure 1 and to think about what would happen if we did leave one out (LOO) cross validation. If we leave out any one of the data points from either group of in Figure 1, the two remaining data points ensure that the model is minimally affected. Similar problems can be encountered even when an external test set is used. My view is that training and test sets need to be selected to cover chemical space as evenly as possible in order to get a realistic assessment of model quality from the validation.  Put another way, ML modelers need to view the selection of training and test sets as a design problem in its own right.

At Royal Albatross Centre, Otago Peninsula

Given that Pat's post is billed as a practical view from the trenches, it may be worth saying something about some of the challenges of achieving genuine impact with ML models in real life drug design projects. Drug discovery is incremental in nature and a big part of the process is obtaining the data needed to make decisions as efficiently as possible. In order to have maximum impact on drug discovery, cheminformaticians will need to be involved how the data is obtained as well as analyzing the data.

Using an ML model is a data-hungry way to predict biological activity and, at the start of a project, the team is not usually awash with data. Molecular similarity searching, molecular shape matching and pharmacophore matching can deliver useful results using much less data than you would need for building a typical ML model while docking can be used even when there are no known ligands.

ML models that simply predict whether or not a compound will be "active" are unlikely to be of any value in lead optimization. Put another way, if you suggest to lead optimization chemists that they should make compound X rather than compound Y because it is more likely to have better than micromolar activity, they may think that you'd just stepped off the shuttle from the Planet Tharg. To be useful in lead optimization, a model for prediction of biological activity needs to predict pIC50 values (rather than whether or not pIC50 will exceed a threshold) and should be specific to the region of chemical space of interest to the lead optimization team. A model satisfying these requirements may well be more like the boring old QSAR that has been around for decades than the modern ML model. One difficulty that QSAR modelers have always faced when working on real life drug discovery projects is that key decisions have already been made by the time there is enough data with which to build a reliable model.

While I do not think that ML models are likely to have significant impact for prediction of activity against primary targets in drug discovery projects, they do have more potential for prediction of physicochemical properties and off-target activity (for which measured data are likely to be available for a wider range of chemotypes than is the case for the primary project targets). Furthermore, predictions for physicochemical properties and off-target activity don't usually need to be as accurate as predictions for activity against the primary target. Nevertheless, there will always be concerns about how effectively a model covers  relevant chemical space (e.g. structural series being optimized) and it may be safer to just get some measurements done. My advice to lead optimization chemists concerned about solubility would generally be to get measurements for three or four compounds spanning the lipophilicity range in the series and examine the response of aqueous solubility to lipophilicity.

I do have some thoughts on how cheminformatic models can be made more intelligent but this post is already too long so I'll need to discuss these in a future post. It's "até mais" from me (and the Royal Albatrosses of the South Island).