Showing posts with label lipophilic efficiency. Show all posts
Showing posts with label lipophilic efficiency. Show all posts

Tuesday, 30 July 2024

A Nobel for property-based drug design?

[This post was updated on 10-Aug-2025 to mention my review of CNM2025 (Return to Flatland) which  critically examines (35) (Escape from Flatland: Increasing Saturation as an Approach to Improving Clinical Success)] 
 
This post was updated on 04-Aug-2024. I thank Tim Ritchie (see RM2009 | RM2014) for bringing YG2003 (Prediction of Aqueous Solubility of Organic Compounds by Topological Descriptors) to my attention.]

"The problems of ADME are precisely those that determine success or failure of a drug in vivo. In vitro data can give a clearer picture of the receptor characteristics, but knowledge and control of ADME are also vital. A common trap in binding studies is that binding generally increases with lipophilicity, so that one may obtain extremely potent binding that is totally unattainable in vivo."

SH Unger (1987) Computer-Aided Drug Design in the Year 2000. 
Drug Information Journal 21:267-275 DOI
******************************************

In this post I’ll be reviewing an Editorial (Property-Based Drug Design Merits a Nobel Prize) that was recently published in J Med Chem. For me, the Editorial raises questions about the critical thinking skills of its authors and of the judgement of the J Med Chem Editors (I’m guessing that some of the courteous and cultured members of the Nobel Prize committee might regard it to be somewhat pushy, and possibly even uncouth, for journals to be publishing nominations for Nobel Prizes as editorials). My advice to anybody nominating individuals for a Nobel Prize is to be aware of an observation, usually attributed to Jocelyn Bell Burnett, that it’s better that people ask why you didn’t win a Nobel Prize than why you did. Where applicable, I've used the the same reference numbers that were used in the Editorial and I’ll start by reproducing the Nobel Prize proposal (as is usual in posts at Molecular Design, I’ve inserted some comments, italicized in red and enclosed in square brackets, into the quoted text):
We propose that a Nobel Prize in Physiology or Medicine should be awarded for property-based drug design, with Christopher A. Lipinski, Paul D. Leeson, and Frank Lovering as the proposed recipients for their development of “important principles for drug design” [I would describe what the proposed Nobel laureates have introduced as a rule, a metric and a molecular descriptor rather than principles.], principles that have contributed to the development of numerous approved drugs. [The authors do need to provide convincing evidence to support what appear to be some wildly extravagant claims. Specifically, the authors need to demonstrate that the rule, metric and molecular descriptor (which they describe as “principles”) were actually critical to the decision-making in projects that led to the development of numerous drugs.] While drug design previously focused primarily on optimizing potency, they introduced a more holistic approach based on the consideration of how fundamental molecular and physicochemical properties affect pharmaceutical, pharmacodynamic, pharmacokinetic, and safety properties. [My view is that none of proposed Nobel laureates even demonstrated a single convincing link between molecular and physicochemical properties, and pharmaceutical, pharmacodynamic, pharmacokinetic, and safety properties.] The development of the Rof5 by Christopher A. Lipinski in 1997 introduced a new principle for how molecular and physicochemical properties affect oral bioavailability. The development of LipE by Paul D. Leeson in 2007 introduced a new principle for how physicochemical properties impact potency, selectivity, and safety. Finally, the development of Fsp3 by Frank Lovering in 2009 introduced a new principle for how molecular shape affects pharmaceutical properties and developability.

Before examining the contributions of the three nominated individuals it's worth saying something about the objectives of drug design. First, a drug needs to be highly active against its target(s). Second, activity against anti-targets should be very low (ideally too low to even be measured). Third, as I note in 34, the exposure (concentration at the site of action) of the drug needs to be controllable (one challenge in drug design is that intracellular drug concentration can’t generally be measured in vivo and I recommend that all drug discovery scientists read SR2019). I see controlling exposure as the primary focus of property-based design and one fundamental challenge is that structural modifications that lead to increased engagement potential for the therapeutic target(s) frequently result in reduced controllability of exposure as well as increased engagement potential for anti-targets. I’ve tried to capture these points in the graphic shown below.


It's generally accepted that excessive lipophilicity and molecular size are risk factors in drug design and the “compound quality” (CQ) literature abounds with fire-and-brimstone sermons on the evils of "molecular obesity" (see H2011). Nevertheless, the relationships between these descriptors and properties such as binding affinity for anti-targets, permeability, aqueous solubility and metabolic lability are generally not quite as strong as is commonly believed (or claimed). When using trends in data to inform design it’s really important to know how strong the trends are because this tells you how much weight to give to the trends when making decisions. It’s not unknown in CQ studies for trends in data to be made to appear to be stronger than they actually are which endows the CQ field with what I’ll politely call a “whiff of the pasture” (the term “correlation inflation” has been used; see KM2013). Transformation of continuous data (IC50 values) to categorical data (high | medium | low) prior to analysis should trigger a deafening cacophony of alarm bells as should any averaging of groups of continuous data values without showing the spread in the data values. Some examples of studies in which I consider the strengths of trends to have been exaggerated include 29, 35, HMO2016 and HY2010.

I think that one thing that everybody who actually works (or has worked) on drug discovery projects agrees on is that drug discovery is really difficult. My view is that, by focusing on Rof5, LipE and Fsp3, the Editorial actually trivializes the challenges faced by drug discovery scientists. Most drug design (as opposed to ligand design) takes place during lead optimization and lead optimization teams are typically addressing specific problems (for example, structural changes that result in increased potency also result in reduced aqueous solubility).  Lead optimization teams typically work with a lot of measured data (a significant component of drug design is efficient generation of data to enable decision-making) and a weak correlation between logP and aqueous solubility reported in the literature would be of no practical relevance when the lead optimization team is using aqueous solubility measurements for compounds in the structural series that they’re optimizing. It is common (see M2001 | G2008) for the simplicity of rules, guidelines and metrics to be touted and we noted in KM2013 that:   

Given that drug discovery would appear to be anything but simple, the simplicity of a drug-likeness model could actually be taken as evidence for its irrelevance to drug discovery.

Guidelines for successful drug discovery are often presented in terms of something good (or bad) being more likely to happen when the value of a calculated property such as Fsp3 exceeds a threshold. When using guidelines like these be aware that it’s actually very difficult to set these threshold values objectively and that the guidelines would have been stated in an identical manner had different threshold values been chosen to specify them. One difficulty with using guidelines like these is that the creators of the guidelines don’t usually say what they mean by “more likely” (millions of people book flights knowing that one is “more likely” to die in a plane crash if one takes a flight than if one doesn’t take a flight). A number of published guidelines (some of which have been referenced in the Editorial) claim that compounds that comply with the guidelines are more likely to be developable. However, giving weight to these claims would require that developability be defined in an objective manner that enables compounds with arbitrary molecular structures and differing biological activity to be meaningfully compared.   

I’ll examine the contributions of the three proposed laureates for the Nobel Prize in Physiology or Medicine following the order in the Editorial. Let's start with the first:
  
The development of the Rof5 by Christopher A. Lipinski in 1997 introduced a new principle for how molecular and physicochemical properties affect oral bioavailability. [As a reviewer of the manuscript I would have pressed the authors to explicitly state the new principle that their first nominee for the Nobel Prize for Physiology or Medicine had introduced 1997.]

My view is that the publication of the Rof5 (22) has certainly proven to be highly influential in that it made many drug discovery scientists aware of the need to take account of physicochemical properties, in particular lipophilicity, in drug design. What is less well-known, but possibly more important in my view, is that publication of the Rof5 sent a clear message to Pharma/Biotech management that high-throughput screening wasn’t going to be the panacea that many believed that it would be. However, I don't see the Rof5 as quite the epiphany that the authors of the Editorial would have us believe it to be. The quote with which I started this post was taken from an article that had been published ten years before 22 and the inverse nature of the relationship between aqueous solubility and lipophilicity was being discussed in the scientific literature (see YV1980) more than forty years ago. The NC1996 study is also worthy of mention because it was published more than a year before 22 and it makes the important point that optimal logP values are likely to vary with chemotype ("each congeneric series for a drug backbone usually demonstrates its own optimal log P").       

Questions can be raised about the data analysis presented in support of the Rof5 and readers may find it helpful to take a look at the S2019 study as well as my comments on the Rof5 in HBD3 and in this post. I would argue that the Rof5 does not have any practical value as a drug design tool and I would challenge the assertion made in the Editorial that the publication of 22 demonstrated how “molecular and physicochemical properties affect oral bioavailability”. One aspect of the analysis presented (22) in support of the Rof5 that isn't always fully appreciated is that the compounds for which the descriptors are calculated were all treated as having equivalent oral bioavailability (compounds were selected for the analysis on the basis of having been taken into phase 2 clinical trials at some point before the Rof5 had been published in 1997). This is one reason that it’s not credible to assert that the analysis demonstrates that these molecular and physicochemical properties are linked to bioavailability (it must be stressed that, like many, I do actually believe that excessive lipophilicity and molecular size are risk factors in drug design). I make the following point in a blog post (I’ve modified the original text very slightly for consistency with the Editorial):

The Rof5 is stated in terms of likelihood of poor absorption or permeation although no measured oral absorption or permeability data are given in 22 and the Rof5 should therefore be regarded as a statement of belief. I realise that to make such an assertion runs the risk of an appointment with the auto-da-fé and I stress that had the Rof5 been stated in terms of physicochemical and molecular property distributions I would not have made the assertion.

To see what I was getting at let’s take a look at how the Rof5 was stated in 22 (“The ‘rule of 5’ states that: poor absorption or permeation are more likely when…”). However, the analysis presented in support of the Rof5 was of the distribution of compounds in chemical space defined by molecular weight, logP and numbers of hydrogen bond donors and acceptors with no account being taken of variation in either absorption or permeation for the compounds. Analysis like this can be informative but you need to demonstrate that the chemical space is actually relevant to the phenomena of interest. One way that you can demonstrate that a chemical space is relevant is to build predictive models for the phenomena of interest using only the dimensions of the chemical space as descriptors. Alternatively you might observe meaningful differences between the distributions in the chemical space for compounds that have respectively passed and failed at at a particular stage in clinical development.  

So that’s all that I’ll be saying about Rof5 and it’s time to take a look at the contributions of the second proposed Nobel Laureate:

The development of LipE by Paul D. Leeson in 2007 introduced a new principle for how physicochemical properties impact potency, selectivity, and safety. [As a reviewer of the manuscript I would have pressed the authors to explicitly state the new principle that their second nominee for the Nobel Prize for Physiology or Medicine had introduced 2007.]

I'll start by saying that LipE is a simple mathematical formula and I suggest that one shouldn't be confusing simple mathematical formulae with principles when nominating people for Nobel Prizes. There are, however, other errors and these are not the kind of errors that you can afford to make when nominating people for Nobel Prizes. First, the term used in 29 is actually “ligand-lipophilicity efficiency” (LLE) although this appears to have mutated to “lipophilic ligand efficiency” (also LLE) by 2014 (see H2014). The term “LipE” was actually introduced by Pfizer scientists (see R2009) and it is significant that the more recent J2018 article defines LipE in terms of logD rather than logP (doing so means that you can make compounds more efficient simply by increasing extent of ionization and, as a drug design tactic, this is likely to end about as well as things did for the Sixth Army at Stalingrad).

The second (and more serious from the perspective of a Nobel nomination) error is that the metric had already been discussed, although not named, in the literature when 29 was published (I’m guessing that a suggestion that naming a metric merits a Nobel Prize for Physiology or Medicine might cause some members of the Nobel Prize committee to choke on their surströmming).  The L2006 book chapter, published fifteen months before 29, states:

Thus, to achieve compounds with a not too high log P while still retaining potency, the difference between the log potency and the log D can be utilised.

From the A2007 perspective which was published three months before 29

Lipophilicity is thought to be a driving force for binding to anti-targets such as the hERG ion channel and cytochrome p450 enzymes and potency can be scaled by lipophilicity by subtracting measured or calculated 1-octanol water partition coefficients from pIC50.

It might be helpful to say something about efficiency metrics since LiPE (or LLE if you prefer) is an example of an efficiency metric. The idea behind efficiency metrics is to “normalize” a compound’s activity (typically quantified by potency or affinity) by the value of a risk factor such as lipophilicity or molecular size (for the masochists among you there’s an entire section in 34 on normalization of binding affinity). Ligand efficiency (LE) was introduced in 2004 (see H2004) and is generally regarded as the original efficiency metric although its creators do acknowledge the influence of the K1999 study. I’ve argued at length in 34 (Table 1 and Figure 1 in the article capture the essence of the argument) that LE is physically meaningless because perception of efficiency changes if you use a different concentration to define the standard state (by convention ΔGbinding values correspond to an arbitrary 1 M standard concentration) and there is no way to objectively select any particular value of the standard concentration for calculation of LE.  The problem doesn’t go away if you try to define ligand efficiency in terms of logarithmically expressed values of IC50, Ki or Kd instead of ΔGbinding because these quantities still have to be divided by an arbitrary concentration value in order to be expressed as logarithms (see M2011).  My view is that LE shouldn't even be described as a metric and I sometimes appropriate a quote ("it's not even wrong") that is usually attributed to Pauli because those who advocate the use of LE in drug design are unable (or unwilling) to say what it measures.

The meaninglessness of LE stems from it being defined by scaling ΔGbinding by the design risk factor (molecular size). In contrast, LipE is defined by offsetting pIC50 by the risk factor (logP) and can be interpreted (see 34) as the energetic cost of moving the ligand from octanol to its target binding site (this interpretation is only valid when the ligand binds in its neutral form and is predominantly neutral in the aqueous phase).  When considering lipophilicity in property-based design it is important to be aware that octanol is an arbitrary choice of solvent for measurement of partition coefficients and that the logP (or logD) calculated for a compound may differ significantly depending on the algorithm used for the calculations. That said, the hydrogen bond donors/acceptors and ionizable groups tend to be relatively conserved within structural series which means that the details of exactly how lipophilicity is quantified are likely to be less critical in lead optimization than for structurally-diverse sets of compounds.

When we use LipE we’re actually assuming that logP (or logD) is predictive of properties such as aqueous solubility, affinity for anti-targets and metabolic lability. That is why it’s not accurate to state that the introduction of LipE showed how “physicochemical properties impact potency, selectivity, and safety”.  In some published studies the focus is less on the LipE metric and more on what might be called the "lipophilic efficiency concept" (aim for top left corner of a plot of potency against lipophilicity). It is common to show reference lines of constant LipE to plots of potency against lipophilicity in this type of analysis and if you're doing this you really should be citing R2009 rather than 29

I'll finish the commentary on LipE (or LLE if you prefer) with this statement made in the Editorial:

Emerging from an analysis of approved drugs, this rubric predicts a compound is more likely to be clinically developable when LipE > 5. [I don’t know what the authors of the Editorial mean by “rubric” (I'm not even sure that they do) but as a reviewer of the manuscript I would have pressed them to justify their claim. Specifically I would have been looking for a literature reference (for me, the choice of the word “emerging” does rather conjure up an image of hot gases and stoned priestesses at Delphi) and a coherent explanation for why a value of 5 yields a better rubric than values of 4 or 6.]

That’s all that I’ll be saying about LipE (or LLE if you prefer) and it’s time to take a look at the contributions of the third nominee for the Nobel Prize in Physiology or Medicine:

Finally, the development of Fsp3 by Frank Lovering in 2009 introduced a new principle for how molecular shape affects pharmaceutical properties and developability. [As a reviewer of the manuscript I would have pressed the authors to explicitly state the new principle that their third nominee for the Nobel Prize for Physiology or Medicine had introduced in 2009. My view is that Fsp3 is a thoroughly unconvincing descriptor of molecular shape and I suggest readers consider the suggestion that cyclohexane (Fsp3 = 1) would have a better shape match with benzene (Fsp3 = 0) than with either methane (Fsp3 = 1) or adamantane (Fsp3 = 1).]

[10-Aug-2025 update: The authors of the CNM2025 study claim "we repeated an analysis similar to that of Lovering et al. to assess Fsp3 in drugs approved post-2009 and those in active clinical development as of mid-2024" and conclude "there appeared to be no clear relationship between highest phase reached and Fsp3, suggesting the key conclusion noted by Lovering et al. has not persisted". My view expressed in this 05-Aug-2025 post is that the analysis is not sufficiently similar to support this conclusion.] 
 
[04-Aug-2024 update: The Fsp3 descriptor had actually been used as i_ali in the YG2003 study (Prediction of Aqueous Solubility of Organic Compounds by Topological Descriptors) six years before the publication of 35:

The aliphatic indicator of a molecule (i_ali) is equal to the number of sp3 carbons divided by the total number of carbon atoms in the molecule.

The YG2003 study discussed prediction of aqueous solubility using i_ali (renamed as Fsp3 in 35) in conjunction with other topological descriptors. In contrast with the claims made in 35 for Fsp3 the YG2003 study made no suggestion that i_ali was a highly effective predictor of aqueous solvation when used by itself.]   

Before discussing the contributions of the third nominee for the Nobel Prize for Physiology or Medicine I should stress that I certainly consider gratuitous use of aromatic rings to be a very bad thing in drug design (it was the data analysis in 35 that was criticized in KM2013 but not the eminently sensible suggestion that drug designers should look beyond what the authors referred to as ‘Flatland’). Having sp3 carbon atoms in a scaffold provides drug designers with a wider range of options for placement of substituents than would be the case for a fully aromatic scaffold and we stated in KM2013 that:   

One limitation of aromatic rings as components of drug molecules is that some regions above and below the plane defined by the atomic nuclear positions are not directly accessible to substituents. Molecular recognition considerations suggest a focus on achieving axial substitution in saturated rings with minimal steric footprint, for example by exploiting the anomeric effect or by substituting N-acylated cyclic amines at C2. 

My view is that deleterious effects of aromatic rings on aqueous solubility would be more plausibly explained by molecular interactions stabilizing the solid state than in terms of molecular shape (this point is discussed in more detail in HBD3). I also see saturated ring systems such as bicyclo[1.1.1]pentane and cubane as potentially more resistant to metabolism than benzene. 

There’s one point that I need to make before discussing 35 from the data analysis perspective which is that molecular structures with basic nitrogen atoms tend to have higher Fsp3 values than molecular structures that lack basic nitrogen atoms (see L2013). This means that you can’t tell whether the benefits of higher Fsp3 values are actually caused by the higher Fsp3 values or by the presence of basic nitrogen atoms.

The Editorial states:

Stemming from an analysis of discovery compounds, investigational drugs, and approved drugs, Fsp3 predicts a discovery compound is more likely to become a drug when Fsp3 > 0.40. [Figure 3 in 35 does not actually depict a significant difference between mean Fsp3 values for for discovery compounds and marketed drugs (the significant difference between mean Fsp3 values is for discovery and Phase 2 compounds).]  

It’s not clear (at least to me) where the figure of 0.40 comes from and I would argue that that compound X (IC50 against therapeutic target = 50 μM; Fsp3 = 0.80) would actually be less likely to become a drug than compound Y (IC50 against therapeutic target = 10 nM; Fsp3 = 0.20). I’m assuming that what the Editorial refers to as “analysis of discovery compounds, investigational drugs, and approved drugs” is what is shown by Figure 3 in 35. Presenting data in this manner hides the variation in Fsp3 for the compounds at each stage of development and makes the trends look much stronger than they actually are (this is verboten according to current J Med Chem author guidelines which state "If average values are reported from computational analysis, their variance must be documented.") I would challenge the suggestion that what is shown in Figure 3 in 35 can be used to calculate the probability that an arbitrary compound will become a drug (my view is that it’s not feasible to even define the probability that a compound will become a drug in a meaningful manner). Analyses of success in clinical development are generally more convincing when comparisons are made between compounds that pass or fail in individual phases of clinical development than between compounds in different phases of clinical development. 

The Editorial continues:        

This observation was ascribed to increased Fsp3 leading to increased aqueous solubility, a critical physiochemical property for successful drug discovery.

I’m assuming that what the Editorial refers to as “increased Fsp3 leading to increased aqueous solubility” is the trend shown by Figure 5 of 35 (this featured prominently in the KM2013 correlation inflation article) which claims to show the relationship between Fsp3 and log S (aqueous solubility expressed as a logarithm).  This claim is not accurate because the log S values have been binned and the relationship is actually between centre point of bin and mean log S value for bin. The authors of 35 used public domain aqueous solubility data for their analysis and we showed (KM2013; see Figure 5) that the Pearson correlation coefficient for the relationship between log S and Fsp3 is only 0.25 (the corresponding value for the binned data is 0.97).  I consider the suggestion that such a weak correlation could have any relevance whatsoever to the the likelihood of success in clinical trials to be wild and uninformed conjecture.      

I'll finish my commentary on Fsp3 by reproducing this claim made in the Editorial:

Much like the Rof5 and LipE, Fsp3 has proven to be enduringly useful for the design of compounds with improved chances of clinical success. (37) [My view is there is insufficient evidence to justify this claim and I'm perplexed by the citation of 37. In any case, members of the Nobel committee are likely to focus more on whether or not Fsp3 is usefully predictive than on the endurance of this molecular descriptor.]  

It’s now time to summarise what has been a long and at times pedantic blog post, and I thank all readers who’ve stayed with me. I don’t consider any of the three studies (22 | 29 | 35) that form the basis of the Nobel Prize nomination to have reported significant scientific discoveries and I would also challenge the claim made in the Editorial that these studies introduced new principles. I’m aware that 22 is heavily cited and I certainly agree that it is common to see values of LipE and Fsp3 quoted in the drug discovery literature. Nevertheless, I would argue that that the Editorial failed to provide even a single convincing example of the Rof5, LipE or Fsp3 making a critical contribution to the discovery of a marketed drug (this should be quite sufficient to rule out the award of a share in the Nobel Prize for Physiology or Medicine to any of these nominees). Furthermore, the Editorial doesn’t provide any convincing evidence that the Rof5, LipE or Fsp3 are usefully predictive in drug discovery projects.

Aside from the failure of the Editorial to demonstrate significant impact for the Rof5, LipE and Fsp3, I do have some scientific concerns about this Nobel Prize nomination. First, the Rof5 is not actually supported by data in the form that it is stated. Second, LipE had already been discussed, although not named, in the drug discovery literature when 29 was published. Third, Fsp3 had been already been introduced (as i_ali) for aqueous solubility prediction and the data analysis in 35 would fail to comply with current J Med Chem author guidelines.

Saturday, 11 May 2019

Efficient trajectories


I'll examine an article entitled ‘Mapping the Efficiency and Physicochemical Trajectories of Successful Optimizations’ (YL2018) in this post and I should note that the article title reminded me that abseiling has been described as the second fastest way down the mountain. The orchids in Blanchisseuse have been particularly good this year and I’ll include some photos of them to break the text up a bit.


It’s been almost 22 years since the rule of 5 (Ro5) was published. While the Ro5 article highlighted molecular size and lipophilicity as pharmaceutical risk factors, the rule itself is actually of limited utility as a drug design tool. Some of the problems associated with excessive lipophilicity had actually been recognized (see Yalkowsky | Hansch) over a decade before the publication of Ro5 in 1997 and there’s also this article that had been published in the previous year. However, it was the emergence of high-throughput screening that can be regarded as the trigger for Ro5 which, in turn, dramatically raised awareness of the importance of physicochemical properties in drug design. The heavy citation and wide acceptance of Ro5 provided incentives for researchers to publish their own respective analyses of large (usually proprietary) data sets and this has been expressed more succinctly as “Ro5 envy”.



So let's take a look at YL2018 and the trajectories. I have to concede that ‘trajectory’ makes it all seem so physical and scientifically rigorous even though ‘path’ would be more appropriate (and easier to say after a few beers). As noted in ‘The nature of ligand efficiency’ (NoLE), I certainly believe that it is a good idea for medicinal chemistry teams to both plot potency (e.g. pIC50) against risk factors such as molecular size or lipophilicity for their project compounds and to analyze the relationships between potency and these quantities. However, it is far from clear that a medicinal chemistry team optimizing a specific structural series against a particular target would necessarily find the plots corresponding to optimization of other structural series against other targets to be especially relevant to their own project.

YL2018 claims that “the wider employment of efficiency metrics and lipophilicity control is evident in contemporary practice and the impact on quality demonstrable”. While I would agree that efficiency metrics are integral to the philatelic aspects of modern drug discovery, I don’t believe that YL2018 actually presents a single convincing example of efficiency metrics being used for decision making in a specific drug design project. I should also point out that each of the authors of YL2018 provided cannon fodder (LS2007 | HY2010 ) for the correlation inflation article and you might want to keep that in mind when you read the words “evident” and “demonstrable”. They also published 'Molecular Property Design: Does Everyone Get It?' back in 2015 and you may find this review of that seminal contribution to the drug design literature to be informative.

I reckon that it would actually be a lot more difficult to demonstrate that efficiency metrics were used meaningfully (i.e. for decision making rather than presentation at dog and pony shows) in projects than it would be to demonstrate that they were predictive of pharmaceutically relevant behavior of compounds. In NoLE, I stated:

"However, a depiction [6] of an optimization path for a project that has achieved a satisfactory endpoint is not direct evidence that consideration of molecular size or lipophilicity made a significant contribution toward achieving that endpoint. Furthermore, explicit consideration of lipophilicity and molecular size in design does not mean that efficiency metrics were actually used for this purpose. Design decisions in lead optimization are typically supported by assays for a range of properties such as solubility, permeability, metabolic stability and off-target activity as well as pharmacokinetic studies. This makes it difficult to assess the extent to which efficiency metrics have actually been used to make decisions in specific projects, especially given the proprietary nature of much project-related data."



YL2018 states, “Trajectory mapping, based on principles rather than rules, is useful in assessing quality and progress in optimizations while benchmarking against competitors and assessing property-dependent risks.” and, as a general point, you need to show you're on top of the physical chemistry if you're going write articles like this.

Ligand efficiency represents something of a liability for anybody claiming expertise in physical chemistry. The reason for this is that perception of efficiency depends on the unit that you use to express affinity and this is a serious issue (in the "not even wrong" category) that was highlighted in 2009 and 2014 before NoLE was published. While YL2018 acknowledges that criticisms of ligand efficiency have been made, you really need to say exactly why this dependence of perception is not a problem if you're going lecture about principles to readers of Journal of Medicinal Chemistry.

Ligand lipophilic efficiency (LLE) which is also known as ligand lipophilicity efficiency (LLE) and lipophilic efficiency (LipE) can be described as offset efficiency metric (lipophilicity is subtracted from potency). As such, perception of efficiency does not change when you use a different unit to express potency and, provided that ionization of ligand is insignificant, efficiency can be seen as a measure of the ease of transfer of ligand from octanol to its binding site. Here's a graphic that illustrates this:

LLE (LipE) measures ease of transfer of ligand from octanol to binding site

I'm not entirely convinced that the authors of YL2018 properly understood the difference between logP and logD. Even if they did, they needed to articulate the implications for drug design a lot more clearly than they have done. Here's an equation that expresses logD as a function of logP and the fraction of ligand in the neutral form at the experimental pH (assuming that only neutral forms of ligands partition into the octanol).


The equation highlights the problems that result from using logD (rather than logP) to define "compound quality". In essence the difficulty stems from the composite nature of logD which means that logD can be also be reduced by increasing the extent of ionization. While this is likely to result in increased aqueous solubility, it is much less likely that problems associated with binding to anti-targets will be addressed. Increasing the extent of ionization may also compromise permeability.    


YL2018 is clearly a long article and I'm going to focus on two of the ways in which the authors present values of efficiency metrics. The first of these is the "% better" statistic which is used to reference specific compounds (e.g. optimization endpoints) to sets of compounds (e.g. everything synthesized by project chemists). The statistic is calculated as the fraction of compounds in the set for which both LE and LLE values are greater than the corresponding values for the compound of interest. The smallest values of the "% better" statistic are considered to correspond to the most optimal compounds. The use of the "% better" statistic could be taken as indicating that absolute thresholds for LE and LLE are not useful for analyzing optimization trajectories..

The fundamental problem with analyzing data in this manner is that LE has a nontrivial dependence on the concentration unit in which affinity is expressed (this is shown in Table 1 and Fig. 1 in NoLE). One consequence of this nontrivial dependence is that both perception of efficiency and the "% better" statistic vary with the concentration unit used to express efficiency.

The second way that the authors of YL2018 present values of efficiency metrics is to plot LE against LLE and, as has already been noted, this is a particularly bone-headed way to analyze data. One problem is that the plot changes in a nontrivial manner if you express affinity in a different unit. This makes it difficult to explain to medicinal chemists why they need to convert the micromolar potencies from their project database to molar units in order for The Truth to be revealed. Another problem is that LE and LLE are both linear functions of pIC50 (or pKi) and that means that the appearance of the plot is heavily influenced by the (trivial) correlation of potency with itself.

A much better way to present the data is to plot LLE against number of non-hydrogen atoms (or any other measure of molecular size that you might prefer). In such a plot, expressing potency (or affinity) in a different unit simply shifts all points 'up' or 'down' to the same extent which means that you no longer have the problem that the appearance of plot changes when you change units. The other advantage of plotting the data in this manner is that there is no explicit correlation between the quantities being plotted. I have used a variant of this plot in NoLE (see Fig. 2b) to compare some fragment to lead optimizations that had been analyzed previously.

I think this is a good point to wrap things up. Even if you have found the post to be tedious, I hope that you have at least enjoyed the orchids. As we would say in Brazil, até mais!

    

Friday, 3 June 2016

Yet more on ligand efficiency metrics

In this post, I'll be responding to a couple of articles in the literature that cited our gentle critique of ligand efficiency metrics (LEMs). The critique has also been distilled into a harangue and  readers may find that a bit more digestible that the article. As we all know, Ligand Efficiency (LE), the original LEM was introduced to normalize affinity with respect to molecular size. Before getting started, I'd like to ask you, the reader, to ask yourself exactly what you take to mean by the term 'normalize'.

The first article which I'll call L2016 states:

Optimisation frequently increases molecular size, and on average there is a trade-off between potency and size gains, leading to little or no gain in LE [42,52] but an increase in SILE [52]. This, and the nonlinear dependence of LE on heavy atom count, together with thermodynamic considerations, has led some authors to question the validity of LE [76,77], while others support its use [52,78,79].

This statement is misleading because the "thermodynamic considerations" are that our perception of efficiency changes when we change the concentration units in which affinity and potency are expressed.  As such, LE is a physicochemically meaningless quantity and, in any case, references 52 and 78 precede our challenge to the thermodynamic validity of LE (although not an equivalent challenge in 2009). Reference 78 uses a mathematically invalid formula for LE when claiming to have shown that LE is mathematically valid and reference 79 creates much noise while evading the challenge. I have responded to reference 79 (aka the 'sound and fury article') in two blog posts ( 1 | 2 ).

This is a good place for a graphic to break up the text a bit and I'll use the table (pulled from an earlier post) that shows how our perception of ligand efficiency changes with the concentration units used to define affinity. I've used base 10 logarithms and dispensed with energy units (which are often discarded) to redefine LE as generalized LE (GLE) so that we can explore the effect of changing the concentration unit (which I've called a reference concentration). Please take special of note how a change in concentration unit can change your perception of efficiency for the three compounds. Do you think it makes sense to try to 'correct' LE for the effects of molecular size? 


Another article also cites our LEM critique.  Let's take a look at how the study, which I'll call M2016, responds to our criticism of LE (reference 69 in this study):

The appeal of LE and GE is in the convenience and rapidity with which these factors can be assessed during lead optimization, but the simplistic nature of these metrics requires an understanding of, and appreciation for, their inherent limitations when interpreting data.[67,68,69,70The relevance of LE as a metric has been challenged based on the lack of direct proportionality to molecular size and an inconsistency of the magnitude of effect between homologous series, both attributed to a fundamental invalidity underlying its mathematical derivation.[65,67] These criticisms have stimulated considerable discussion and provoked discourse that attempts to moderate the perspective and provide guidance on how to use LE and GE as rule-of-thumb metrics in lead optimization.[68,69,70]

To be blunt, I don't think that the M2016 study does actually respond to our criticism of LE as a metric which is that our perception of efficiency changes when we change the concentration unit with which we specify affinity or potency. This is an alarming characteristic for something that is presented as a tool for decision making and, if it were a navigational instrument, we'd be talking about fundamental design flaws rather than "limitations". The choice of 1 M is entirely arbitrary and selecting a particular concentration unit for calculation of LE places the burden of proof on those making the selection to demonstrate that this particular concentration unit is indeed the one that is most fit for purpose. 


The other class of LEM that is commonly encountered is exemplified by what is probably best termed lipophilic efficiency (LipE).  Although the term LLE is more often used, there appears to be some confusion as to whether this should be taken to mean ligand-lipophilicity efficiency or lipophilic ligand efficiency so it's probably safest to use LipE. Let's see what the M2016 study has to say about LipE:

LLE is an offsetting metric that reflects the difference in the affinity of a drug for its target versus water compared to the distribution of the drug between octanol and water, which is a measure of nonspecific lipophilic association.[69,12]

If I knew very little about LEMs, I would find this sentence a bit confusing although I think that it is essentially correct. We used (and possibly even introduced) the term 'offset' in the LEM context to describe metrics that are defined by subtracting risk factor from affinity (or potency). This is in contrast to LE and its variations which are defined by dividing affinity (or potency) by molecular size and can be described as scaled. There is still an arbitrary aspect to LipE in that we could ask whether (pIC50 - 0.5 ´ logP) might not be a better metric than  (pIC50 - logP).  Unlike LE, however, LipE is a quantity that actually has some physicochemical meaning, provided that the compound in question binds to its target in an uncharged form. Specifically, LipE can be considered to quantify the ease (or difficulty) of moving the compound from octanol to its binding site in the target as shown in the figure below:


Let's see what M2016 study has to say:

However, care needs to be exercised in applying this metric since it is dependent on the ionization state of a molecule, and either Log P or Log D should be used when appropriate.

This statement fails to acknowledge a third option which is that there may be situations in which  neither logP nor logD is appropriate for defining LipE. One such situation is when the compound binds to its target in a charged form. When this is the case, neither logP nor logD quantifies the ease (or difficulty) of moving the bound form of compound from octanol to water. As an aside, using logD to quantify compound quality suggests that increasing the extent of ionization will lead to better compounds and I hope that readers will see that this is a strategy that is likely to end in tears.

Let's take a look at LEMs from the perspective of folk who are working in lead optimization projects or doing hit-to-lead work. Merely questioning the value of LEMs is likely incur the wrath of Mothers Against Molecular Obesity (MAMO) so I'll stress that I'm not denying that excessive lipophilicity and  molecular size are undesirable. We even called them "risk factors" in our LEM critique. That said, in the compound quality and drug-likeness literature, it is much more common to read that X and Y are correlated, associated or linked than to actually be shown how strong the correlation, association or linkage is. When you do get shown the relationship between X and Y, it's usually all smoke and mirrors (e.g. graphics colored in lurid, traffic light hues). When reading M2016 you might be asking why can't we see the relationship between PFI and aqueous solubility presented more directly (or even why iPFI is preferred over PFI for hERG and promiscuity). A plot of one against the other perhaps even a correlation coefficient? Is it really too much to ask?

The reason for the smoke and mirrors is that the correlations are probably weak. Does this mean that we don't need to worry about risk factors like molecular size and lipophilicity? No, it most definitely does not! "You speak in more riddles than a Lean Six Sigma belt", I hear you say, "and you tell us that the correlations with the risk factors have been smoked and mirrored and yet we still need to worry about the risk factors".  Patience, dear reader, because the apparent paradox can be resolved once you realize some much stronger local correlations may be lurking beneath an anemic global correlation. What this means is that potencies of compounds in different projects (and different chemical series in the same project) may respond differently to risk factors like lipophilicity and molecular size. You need to start thinking of each LO project as special (although 'unique' might be a better term because 'special projects' were what used to happen to senior managers at ICI before they were put out to pasture).  

Another view of LEMs is that they represent reference lines. For example, we can plot potency against molecular size and draw a line with positive slope from a point corresponding to a 1 M IC50 on the potency axis and say that all points on the line correspond to the same LE. Analogously, we can draw a line of unit slope on a plot of pIC50 against logP and say that all points on the line correspond to the same LipE.  You might be thinking that these reference lines are a bit arbitrary and you'd be thinking along the right lines. The intercept on the potency axis is entirely arbitrary and that was the basis of our criticism of LE. A stronger case can be made for considering  a line of unit slope on a plot of pIC50 against logP to represent constant LipE but only if the compounds bind in uncharged forms to their target.

Let's get back to that project you're working on and let's suppose that you want to manage risk factors like lipophilicity and molecular size. Before you calculate all those LEMs for your project compounds, I'd like you to plot pIC50 against molecular size (it actually doesn't matter too much what measure of molecular size you use). What you now have in front of you is the response of potency to molecular size.  Do you see any correlation between pIC50 and molecular size? Why not try fitting a straight line to your data to get an idea of the strength of the correlation? The points that lie above the line of fit beat the trend in the data and the points that lie below the line are beaten by the trend. The residual for a point is simply the distance above the line for that point and its value tells you how much the activity that it represents beats the trend in the data. Are there structural features that might explain why some points are relatively distant from the line that you've fit? In case you hadn't realized it, you've just normalized your data. Vorsprung durch technik! Here's a graphic to give you an idea how this might work. 



The relationship between affinity and molecular size shown in the plot above is likely to be a lot tighter than what you'll see for a typical project. In the early stages of a project, the range in activity for the project compounds will often be too narrow for the response of activity to risk factor to be discerned. You can make assumptions about the response of affinity (or potency) to risk factor (e.g. that LipE will remain constant during optimization) in order to forecast outcome but it's really important to continually monitor the response of activity to risk factor to check that your assumptions still hold. If affinity (or potency) is strongly correlated with risk factor then you want the response to risk factor to be as steep as possible. Could this be something to think about when trying to prioritize between series?  

So it's been a long post and there are only so many metrics that one can take in a day. If you want to base your decisions on metrics that cause your perception to change with units then as consenting adults you are free to do so (just as you are free to use astrological charts or to seek the face of a deity in clouds). A former prime minister of India drank a glass of his own urine every day and lived to 98. Who would have predicted that? Our LEM critique was entitled 'Ligand efficiency metrics considered harmful' and I now I need to say why. When doing property-based design, it is vital to get as full an understanding as possible of the response of affinity (or potency) to each of the properties in which you're interested. If exploring the relationship between X and Y, it is generally best to analyse the data as directly as possible and to keep X and Y separate (as opposed to looking at the response of a function of Y and X to X). When you use LEMs you're also making assumptions about the response of Y to X and you need to ask yourself whether that's a sensible way to explore the response of Y to X. If you want to normalize potency by risk factor, would you prefer to use the trend that you've actually observed in your data or an arbitrary trend that 'experts' recommend on the basis that it's "simple"?

Next week, PAINS...

Saturday, 13 September 2014

Saving Ligand Efficiency?

<< Previous |

I’ll be concluding the series of posts on ligand efficiency metrics (LEMs) here so it’s a good point at which to remind you that the open access for the article on which these posts are based is likely to stop on September 14 (tomorrow). At the risk of sounding like one of those tedious twitter bores who thinks that you’re actually interested in their page load statistics, download now to avoid disappointment later. In this post, I’ll also be saying something about the implications of the LEM critique for FBDD and I’ve not said anything specific about FBDD in this blog for a long time.

LEMs have become so ingrained in the FBDD orthodoxy that to criticize them could be seen as fundamentally ‘anti-fragment’ and even heretical.  I still certainly believe that fragment-based approaches represent an effective (and efficient although not in the LEM sense) way to do drug discovery. At the same time, I think that it will get more difficult, even risky, to attempt to tout fragment-based approaches primarily on a basis that fragment hits are typically more ligand-efficient than higher molecular weight starting points for synthesis.  I do hope that the critique will at least reassure drug discovery scientists that it really is OK to ask questions (even of publications with eye-wateringly large numbers of citations) and show that the ‘experts’ are not always right (sometimes they’re not even wrong).

My view is that the LEM framework gets used as a crutch in FBDD so perhaps this is a good time for us to cast that crutch aside for a moment and remind ourselves why we were using fragment-based approaches before LEMs arrived on the scene.  Fragment-based approaches allow us to probe chemical space efficiently and it can be helpful to think in terms of information gained per compound assayed or per unit of synthetic effort consumed.  Molecular interactions lie at the core of pharmaceutical molecular design and small, structurally-prototypical probes allow these to be explored quantitatively while minimizing the confounding effects of multiple protein-ligand contacts.  Using the language of screening library design, we can say that fragments cover chemical space more effectively than larger species and can even conjecture that fragments allow that chemical space to be sampled at a more controllable resolution.  Something I’d like you to think about is the idea of minimal steric footprint which was mentioned in the fragment context in both LEM (fragment linking) and Correlation Inflation (achieving axial substitution) critiques.  This is also a good point to remind readers that not all design in drug discovery is about prediction.  For example, hypothesis-driven molecular design and statistical molecular design can be seen as frameworks for establishing structure-activity relationships (SARs) as efficiently as possible.

Despite criticizing the use of LEMs, I believe that we do need to manage risk factors such as molecular size and lipophilicity when optimizing lead series. It’s just I don’t think that the currently used LEMs provide a generally valid framework for doing this.  We often draw straight lines on plots of activity against risk factors in order to manage the latter.  For example, we might hypothesize that 10 nM potency will be sufficient for in vivo efficacy and so we could draw the pIC50 = 8 line to identify the lowest molecular weight compounds above this line.   Alternatively we might try to construct line a line that represents the ‘leading edge’ (most potent compound for particular value of risk factor) of a plot of pIC50 against ClogP.  When we draw these lines, we often make the implicit assumption that every point on the line is in some way equivalent. For example we might conjecture that points on the line represent compounds of equal ‘quality’.  We do this when we use LEMs and assume that compounds above the line are better than those below it.

Let’s take a look at ligand efficiency (LE) in this context and I’m going to define LE in terms of pIC50 for the purposes of this discussion. We can think of LE in terms of a continuum of lines that intersect the activity axis at zero (where pIC50  = 1 M).  At this point, I should stress that if the activities of a selection of compounds just happen to lie on any one of these lines then I’d be happy to treat those compounds as equivalent.  Now let’s suppose that you’ve drawn a line that intersects the activity axis at zero.  Now imagine that I draw a line that intersects the activity axis at 3 zero (where IC50  = 1 mM) and, just to make things interesting, I’m going to make sure that my line has a different slope to your line.  There will be one point where we appear to agree (just like −40° on the Celsius and Fahrenheit temperature scales) but everywhere else we disagree. Who is right? Almost certainly neither of us is right (plenty of lines to choose from in that continuum) but in any case we simply don’t know because we’ve both made completely arbitrary decisions with respect to the points on the activity axis where we’ve chosen to anchor our respective lines.  Here’s a figure that will give you a bit more of an idea what I'm talking about.


However, there may just be a way out this sorry mess and the first thing that has to go is that arbitrary assumption that all IC50 values tend to 1 M in the limit of zero molecular size.  Arbitrary assumptions beget arbitrary decisions regardless of how many grinning LeanSixSigma Master Black Belts your organization employs.  One way out of the mire is link efficiency with the response (slope) and agree that points lying on any straight line (of finite non-zero slope) when activity is plotted against risk factor represent compounds of equal efficiency.  Right now this is only the case when that line just happens to intersect the activity axis at a point where IC50  = 1 M.  This is essentially what Mike Schultz was getting at in his series of (  1  | 2  |  3  ) critiques of LE even if he did get into a bit of a tangle by launching his blitzkrieg on a mathematical validity front. 
The next problem that we’ll need to deal with if we want to rescue LE is deciding where we make our lines intersect the activity axis.  When we calculate LE for a compound, we connect the point on the activity versus molecular size representing the compound to a point on the activity axis with a straight line.  Then we calculate the slope of this line to get LE but we still need to find an objective way to select the point on the activity axis if we are to save LE.  One way to do this is to use the available activity data to locate the point for you by fitting a straight line to the data although this won’t work if the correlation between risk factor and activity is very weak.  If you can’t use the data to locate the intercept then how confident do you feel about selecting an appropriate intercept yourself?  As an exercise, you might like to take a look at Figure 1 in this article (which was reviewed earlier this year at Practical Fragments) and ask if the data set would have allowed you to locate the intercept used in the analysis.

If you buy into the idea of using the data to locate the intercept then you’ll need to be thinking about whether it is valid to mix results from different assays.   It may be the same intercept is appropriate to all potency and affinity assays but the validity of this assumption needs to be tested by analyzing real data before you go basing decisions on it.  If you get essentially the same intercept when you fit activity to molecular size for results from a number of different assays then you can justify aggregating the data. However, it is important to remember (as is the case with any data analysis procedure) that the burden of proof is on the person aggregating the data to show that it is actually valid to do so. 

By now hopefully you’ve seen the connection between this approach to repairing LE and the idea of using the residuals to measure the extent to which the activity of a compound beats the trend in the data.  In each case, we start by fitting a straight line to the data without constraining either the slope or the intercept.  In one case we first subtract the value of this intercept from pIC50 before calculating LE from the difference in the normal manner and the resulting metric can be regarded as measuring the extent to which activity beats the trend in the data.  The residuals come directly from the process and there is no need to create new variables, such as the difference between from pIC50 and the intercept, prior to analysis. The residuals have sign which means that you can easily see whether or not the activity of compound beats the trend in the data.  With residuals there is no scaling of uncertainty in assay measurement by molecular size (as is the case with LE).  Finally, residuals can still be used if the trend in the data is non-linear (you just need to fit a curve instead of a straight line).  I have argued the case for using residuals to quantify extent to which activity beats the trend in the data but you can also generate a modified LEM from your fit of the activity to molecular size (or whatever property by which you think activity should be scaled by).

It’s been a longer-winded post than I’d intended to write and this is a good point at which to wrap up.  I could write the analogous post about LipE by substituting ‘slope’ for ‘intercept’ but I won’t because that would be tedious for all of us.  I have argued that if you honestly want to normalize activity by risk factor then you should be using trends actually observed in the data rather than making assumptions that self-appointed 'experts' tell you to make.   This means thinking (hopefully that's not too much to ask although sometimes I fear that it is) about what questions you'd like to ask and analyzing the data in a manner that is relevant to those questions. 

That said, I rest my case.