Showing posts with label philately. Show all posts
Showing posts with label philately. Show all posts

Sunday, 30 November 2025

Covalent ligand efficiency

It appears that whoever first described Economics as the ‘Dismal Science’ had never encountered a ligand efficiency metric. I’ll be taking a look at the FK2025 study (Covalent ligand efficiency) in this post and the study has already been reviewed by Dan. Something that I’ve observed repeatedly over the years is that authors of ligand efficiency studies exhibit a lack of understanding of units and dimensions associated with physicochemical quantities that would shame a first year undergraduate studying introductory physical chemistry (this is somewhat ironic given that creators of ligand efficiency metrics frequently tout their creations in physicochemical terms). I consider covalent ligand efficiency (CLE) as defined in the FK2025 study to have no value whatsoever for design of drugs that bind irreversibly to their targets through covalent bond formation given that the metric is time-dependent and based on an invalid measure of bioactivity. The formidable Lady Bracknell is clearly unimpressed and I should mention that the photo is from the wikipedia page for English actress Rose Leclercq (1843-1899). Given the serious deficiencies in the FK2025 study this is going to be a long and tedious post (even more so than usual 😁😁😁) so please ensure that you have strong coffee close to hand. As is usually the case in posts here I've used the same reference numbers as were used in FK2025 and quoted text is indented with my comments in red italics. I've organized some of the mathematical material into three tables and references to tables in the post are to these (and not to any of the tables in the FK2025 study).    

Before starting the review of FK2025 it’s worth examining irreversible covalent inhibition from a molecular design perspective and I’ll direct readers to the informative S2016 and McW2021 reviews, and the recent L2025 study which presents COOKIE-Pro for covalent inhibitor binding kinetics profiling on the proteome scale. Covalent bond formation between RNA and ligands can also be exploited (S2025 | K2025 | L2015)  and I generally use 'target' rather than 'protein' in blog posts and journal articles. An irreversible covalent inhibitor acts by first binding non-covalently to its target in the first step with the covalent bond forming in the second step between an electrophilic ligand atom (the term warhead is commonly used) and a nucleophilic target atom such as the sulfur atom of a cysteine. A commonly used measure of activity for irreversible covalent inhibitors is the kinact/KI ratio which can be thought of as the product of affinity (1/KI) and reactivity (kinact). In design of irreversible covalent inhibitors we try to place the electrophilic atom of the warhead within reacting distance of the nucleophilic atom of the target (this is relatively easy if you have a reliable structure of a complex of the target with a relevant ligand that lacks the electrophilic warhead). The non-covalent complex between target and ligand is stabilised by the non-covalent contacts between the target and ligand (the term ‘molecular interactions’ is also used although I prefer to think in terms of ‘non-covalent contacts’ since the latter can be observed experimentally). However, non-covalent contacts also determine reactivity of the non-covalently bound complex by stabilising the transition state (I consider it more correct to think in terms of reactivity of the complex than in terms of reactivity of either the electrophilic warhead or the target nucleophile). In the design context, this means attempting to tune non-covalent contacts to stabilise the transition state to a greater extent than the non-covalent complex.

The LE and CLE metrics share a very serious deficiency in that your perception of efficiency can be altered if you change the value of an arbitrary term in the formula for the metric and I'll start the review of FK2025 by critically examining LE. The meaningless of LE stems from a fundamental misunderstanding of how logarithms work and I'll by point you toward M2011 (Can one take the logarithm or the sine of a dimensioned quantity or a unit? Dimensional analysis involving transcendental functions) that was published in the Journal of Chemical Education. In drug discovery we frequently need to calculate logarithms for quantities and you need to be aware you can’t calculate the logarithm for a dimensioned quantity. Let’s take pIC50 as an example and this quantity is commonly defined as the negative logarithm of the IC50 in mole per litre (M). However, what you actually do when you calculate pIC50 is that you take the negative logarithm of the numerical value of the IC50 when expressed in mole per litre (this is a bit of a mouthful and it can be written more compactly as equation 1 below). While not denying that it is useful to have a convention such as this for expressing potency values logarithmically it should be remembered that the choice of mol per litre (M) is entirely arbitrary and it would be equally correct to use other valid concentration units such as μM or nM. One consequence of choosing mole per litre (M) for expressing IC50 values is that pIC50 values (or at least measured pIC50 values) will generally be positive because of the extreme difficulty of measuring meaningful IC50 values that are greater than 1 M. 


Let’s take a look at the binding free energy ΔG° and you’ll notice that I’ve written it with a degree symbol which indicates that this quantity corresponds to a standard state defined by a concentration value C° (the standard concentration). Equation 2 shows how the binding free energy is defined as the difference in chemical potential between the associating species (target + ligand) and the target-ligand complex with each species at the standard concentration (the degree symbol indicates that that both the binding free energy and chemical potential depend on the value of C° and I’ve also shown this explicitly in the equation although this is not actually necessary). Equation 3 shows the dependence of chemical potential on the concentration C of the species and the standard concentration C°. Taken together, Equation 2 and Equation 3 should clarify the origins of the dependence of binding free energy on the standard concentration comes from (there are two associating species but only one complex). We can’t actually measure binding free energy directly but we can calculate it from the dissociation constant KD using Equation 4 (which can be derived from Equation 2 and Equation 3). It’s important to be aware that if you use Equation 4 to convert ΔG° values between different values of the standard concentration C° you’ll be making the assumption that solutions are dilute (ΔH is independent of concentration) and this is indicated in Equation 5.

The standard concentration is a source of much confusion in the ligand efficiency field and I’ll direct readers to ‘The Nature of Ligand Efficiency’ (p9). While the standard concentration is integral to a valid thermodynamic treatment of target-ligand binding the value of C° is entirely arbitrary (to suggest otherwise would mean that you’ve abandoned thermodynamics). It is conventional in drug discovery (and biochemical) literature to use a C° value of 1 M when reporting ΔG° values. While this convention is certainly beneficial, it is no more (and no less) valid to use a value of 1M than it is to use a value of 1 μM for this purpose. Furthermore a standard state defined by a C° value of 1 M is not biophysically realistic (consider the difficulty of accommodating a mole of target in a volume of 1 litre and the likelihood of a ligand exhibiting aqueous solubility of 1 M). I assume that most biochemists and biophysicists would agree that it is generally not feasible to measure KD values of greater of 1 M (I would be happy to be proven wrong on this point) and this means that drug discovery scientists tend to assume that ΔG° values are necessarily negative.

Let’s now a take a look at ligand efficiency (LE) and you can see from the photo above that some heretics regard the metric as physically nonsensical (if you're interested in how I came to be chatting with fellow blogger Ash then take a look at this post). The LE metric which is regarded as an article of  faith in the fragment-based design community was introduced in the (p5) study with the symbol Δg (see Equation 6 below) and the authors of that study did not actually state that it had to be calculated using a C° value of 1 M (I consider it unlikely that any of the authors were even aware of the dependence of ΔG° on C°). In The Nature of Ligand Efficiency (p9) I defined the quantity ηbind (see Equation 7) by dividing Δg (LE) by RT (when LE values are quoted the molar energy units are usually discarded and T often does not correspond to the temperature at which the assay was run) and by the factor (2.303) used to convert between natural logarithms and base 10 logarithms. The quantity ηbind is directly proportional to Δg (LE) and using it makes it much easier to see how using a different standard concentration can alter your perception of efficiency.  Take a look at Table 1 in (p9) and you’ll see that the three compounds (a fragment, a lead and a clinical candidate) bind with equal efficiency when C° is 1 M. Change C° to 0.1 M and the clinical candidate is binds more efficiently than the fragment but when C° to 10 M the fragment becomes more ligand-efficient than the clinical candidate. As noted in (p9) “In thermodynamic analysis, a change in perception resulting from a change in a standard state definition would generally be regarded as a serious error rather than a penetrating insight.” 

Here's what the authors of FK2025 say about LE: 

LE depends on the choice of the standard concentration (normally 1 M) (p8) (p9) and its maximal available value is size dependent.(p10) (p11) [It's true that LE depends on C° but it’s also true that ΔG° depends on C° and the difference in the two dependencies is that is that LE “depends upon the choice of standard concentration in a nontrivial fashion” (p8). The issue is not the so much that LE depends on C° but that using a different unit to express KD changes how we perceive efficiency. The ΔΔG values that determine perception of affinity don’t change when you use a different value of C° (equivalent to using a different unit to express KD). However, if you use a different value of C° for calculating LE you can see from Table 1 in (p9) that even the ordering of LE values between two ligands can change. I consider the molecular size dependencies of LE observed by the authors of (p10) and (p11) to be artefactual and I’ll point you toward Fig. 1 in (p9) which shows that using a different value of C° can change how we perceive the molecular size dependency of LE.] Nevertheless, LE is an established tool to normalize potency and facilitate the comparison of ligands with a range of potencies and sizes. [It is not uncommon for adherents of religions to consider their beliefs to be established facts.] The usefulness of LE and other efficiency metrics in drug discovery has been extensively analyzed and reviewed elsewhere. (p6) (p12) (p13) (p14) (p15) (p16) (p17) [My view is that nobody has actually demonstrated the usefulness of LE and I’m unconvinced that it would even be possible to do so meaningfully in an objective manner (consider the feasibility of comparing success rates between a group of individuals using LE in discovery projects and a control group of individuals not using LE in discovery projects). Usefulness means that using something provides demonstrable benefits and ‘widely-used’ is not equivalent to ‘useful’ (I’m guessing that more people use homeopathic ‘medicines’ than use ligand efficiency metrics). One piece of advice that I’ll offer to anybody advocating the use of LE in drug design is to ensure that you fully understand the implications of changes in perception resulting from using different units to express quantities not least because you might find yourself lecturing to people who do understand.]

After a lengthy preamble it’s now time to take a look at how the FK2025 study addresses ligand efficiency in the context of irreversible covalent inhibition. One of the challenges in design of drugs that engage their targets irreversibly is that it’s not possible to meaningfully quantify activity with a single parameter. This is particularly relevant to definition of efficiency metrics which are typically derived by either scaling or offsetting a measured activity value by a risk factor such as molecular size or lipophilicity. While you can certainly measure an IC50 value for an irreversible covalent inhibitor the value that you measure will be time-dependent and it’s not generally meaningful to compare two IC50 values that have been measured using different incubation times. While the kinact/KI ratio is time-independent using it as a measure of activity necessarily entails a degree of information loss.

The authors of FK2025 state:

Our starting point is the LE introduced for noncovalent ligands as a useful metric for lead selection. (p5) [LE  was claimed to be useful when it was introduced although no evidence was presented in support of the claim.]

Let’s take a look at Table 2 which shows two equations from FK2025. The first equation, which appears in the text of the article, illustrates two common errors in the efficiency metric field (taking logarithms of dimensioned quantities and discarding units). It should ring alarm bells for the reader when authors make either error (especially if the authors interpret values of the efficiency metrics).


The authors of FK2025 assert that “LE can be decomposed into contributions from the noncovalent recognition and the covalent reaction (Box 2, Equation III)” and this is reproduced in Table 2 as Equation 2. The first term is a commonly-used mathematical formula for LE when inhibition is reversible and it is important to be aware that KI has been divided by an arbitrary concentration value (1 M) in order that it can be expressed as a logarithm (see Equation 1 in Table 1). The argument of the logarithm in the second term is dimensionless although its magnitude does vary with t. Each term in Equation III (Box 2) has a nontrivial dependence on the value of an arbitrary quantity (the 1 M concentration in the first term and t, in the second term). This means that your perception of efficiency when calculated according to Equation III (Box 2) will be altered if you use either a different concentration unit or a different value of t. You can see this effect in Figure 2 (effect of varying of t) and the appearance of Figure 1 will be altered if you use a value of t other than 1 h or a concentration unit other than M for the calculation of LE.

It's now time to examine CLE (defined as Equation II in Box 3 of the FK2025 study) and I’ll direct you to Table 3 below in which I’ve made some comments. Using CLE requires that the IC50 values for the inhibitors of interest all correspond to the same time point (t) and it is not clear whether the authors are suggesting that that the IC50 values should all be measured using the same incubation time or need to be calculated from measured KI and kinact values using Equation III in Box 2. A quantity t is also explicitly present in the argument of the logarithm in Equation II in Box 3 and this is necessary for the argument of the logarithm to be dimensionless (see M2011). The argument of the logarithm in Equation II (Box 3) is clearly time-dependent and this means that your perception of efficiency will be altered if you use a different value of t when calculating CLE (just as your perception of efficiency will be altered if you use a different concentration unit to express IC50 when you calculate LE for reversible inhibitors). It also means the molecular size dependency of CLE will vary with time just as the molecular size dependency of LE varies with the concentration unit used to express affinity as can be seen in Fig. 1 of (p9).

However, there is another difficulty which is that the argument of the logarithm in Equation II (Box 3) is not a valid measure of activity (the same criticism can also be made of the xLE metric introduced in the Z2025 study that Dan has already reviewed).  This problem is a bit more subtle and it’s important to remember that knowing the IC50 value for a reversible inhibitor enables you to generate a concentration response for inhibition.  When you express an IC50 value as a logarithm you need to scale it by a concentration value to ensure that the argument of the logarithm function is dimensionless (see M2011) but it’s important to remember that the concentration unit is still there even though it’s not shown (see Equation 1 in Table 1).

This is a good point at which to wrap up and I’ve argued that CLE has two deficiencies. First, perception of efficiency and its dependency on molecular size both vary with an arbitrary quantity (t) in the argument of the logarithm (this is analogous to the problems caused by the arbitrary nature of the concentration unit used for scaling affinity/potency in the definition of LE for reversible binders). Second, the argument of the logarithm is not a valid measure of activity because it cannot be used to generate a concentration response. Furthermore, I would question the value of aggregating results from multiple assays for analysis even for a valid metric without these deficiencies and I offered the following advice in (p9):   

Drug designers should not automatically assume that conclusions drawn from analysis of large, structurally-diverse data sets are necessarily relevant to the specific drug design projects on which they are working.

I’ve criticized the FK2025 study at length and saying how I might use data like this in drug design projects is a good way to conclude the post.  A general criticism that I have made of drug design efficiency metrics is that they are based on assumptions of relationships between activity and risk factors such as molecular size.  I argued in (p9) that one should use the trend that is actually observed in the data to normalize activity with respect to risk factors and I’ll point you to the relevant section (Alternatives to ligand efficiency for normalization of affinity) in that article. I would start by attempting to model the relationship between kinact and reactivity with glutathione. The objective of this exercise is to identify inhibitors that best exploit their intrinsic reactivity when forming covalent bonds with the target residue (you can quantify this by how far the point for an inhibitor lies above the trend line and the most interesting compounds have the largest positive residuals). I might  also examine the relationship between kinact and KI for inhibitors with the same intrinsic reactivity (e.g., incorporating the same warhead) with a view to identifying the inhibitors for which non-covalent interactions with the target most effectively stabilise the transition state relative to the non-covalent complex. I should stress that there is no suggestion that these analyses would necessarily yield useful insight.

It's been a long post so thanks for staying with me. This will be the the last post until after Christmas and, as I extend best wishes to all for a happy and peaceful festive season, I'm keenly aware that Christmas will be neither happy nor peaceful for many of our fellow human beings.  

Tuesday, 5 August 2025

Return to Flatland

Whoever first referred to Economics as ‘The Dismal Science’ had clearly never read an article on ‘3Dness’ in drug discovery.  My own experience reading articles on this topic is a sensation of having my life force slowly sucked out (I even suggested that reviewing the '3Dness' literature might be considered as an appropriate penance when I recently confessed my sins at St Gallen Cathedral) and the subject of Confession reminds me of a song that the late great Tom Lehrer sang about the Second Vatican Council.

In this post I review the CNM2025 study (Return to Flatland) which examines the heavily-cited LBH2009 study (Escape from Flatland: Increasing Saturation as an Approach to Improving Clinical Success).This also is a good point to mention a Journal of Medicinal Chemistry Editorial (Property-Based Drug Design Merits a Nobel Prize) that I reviewed in a 30-Jul-2024 post. The CNM2025 study, which has already been reviewed by Dan and Ash, opens with: 

The year is 2009, Barack Obama has just been inaugurated and both Lady Gaga and The Black Eyed Peas are at the height of their popularity

This couldn’t help but remind me of the “WORLD WAR 2 BOMBER FOUND ON MOON” headline that appeared on the front page of the Sunday Sport twenty-one years before the publication of LBH2009 (it was was accompanied by a photo of a B-17 in a lunar crater). A few weeks later the headline was “WORLD WAR 2 BOMBER FOUND ON MOON VANISHES” (this time accompanied by a photo of the now empty lunar crater).

I’ll start my review of CNM2025 by quoting from it and, as is usual for posts here at Molecular Design, quoted text is indented with any comments by me italicized in red and enclosed in square brackets. 

The hypothesis was attractive, and the data clearly showed the relationship between Fsp3 and clinical progression with pairwise significance P < 0.001. [This statement is inaccurate and Figure 3 of the LBH2009 study shows statistically significant differences at this level between (a) discovery and phase 2 compounds (b) phase 1 and phase 3 compounds (c) phase 2 compounds & drugs.  The authors of LBH2009 state: “The change in average Fsp3 was statistically significant between adjacent stages in only one case (phase 1 to phase 2)” but they neither show this in Figure 3 of their article nor do they report a P-value for the statistical significance of the mean difference in Fsp3 between phase 1 and phase 2 compounds.] The statistics seemed compelling, though the effect size was modest — an increase in average Fsp3 of 0.09 between sets of phase I and approved drugs equates to a difference of around two additional sp3 carbons per drug molecule only. [The authors of LBH2009 did not actually report this difference to be statistically significant so it is unclear why the authors of CNM2025 have stated that the “statistics seemed compelling”.]

The LBH2009 study is effectively a call to think beyond aromatic rings in drug design and my view is that there are considerable benefits in doing so even though I consider the data analysis in the study to be shaky. Almost three decades ago I included a quinuclidine in the Zeneca fragment library for NMR screening and later at AstraZeneca I would actively search (with minimal success) for amides and heteroaryls derived from bicyclic amines. I see the advantages in looking beyond aromatic rings as stemming primarily from increased molecular diversity and a more controllable coverage of chemical space, and in KM2013 we wrote:

Molecular recognition considerations suggest a focus on achieving axial substitution in saturated rings with minimal steric footprint, for example by exploiting the anomeric effect or by substituting N-acylated cyclic amines at C2.

Although data analyses (for example, see HY2010) presented in support of the belief that aromatic rings adversely affect aqueous solubility are typically underwhelming I consider the suggestion to be plausible and suggested in K2022 that deleterious effects of aromatic rings are more likely to be due to their potential for making molecular interactions than to their planarity. That said, I should also point out that the analysis of the relationship between aqueous solubility and Fsp3 presented in Figure 5 of LBH2009 is a textbook example of correlation inflation (see Fig. 5 in KM2013) and I suspect that if a team had submitted this analysis at Statistiques Sans Frontières the judges would have either awarded “nul points” or come to the conclusion that the team had played its joker. Given the Lady Gaga reference in CNM2025 I couldn't resist linking this Peter Gabriel song which includes the lyrics "Adolf builds a bonfire, Enrico plays with it" even though I have absolutely no idea what the the lyrics actually mean.

While the analysis of the relationship between aqueous solubility presented in Figure 5 of LBH2009 does endow the study with what I’ll politely call a whiff of the pasture it’s not directly related to the analysis of clinical progression presented in the study. Let’s take a look at Figure 3 in LBH2009 which shows mean Fsp3 values for compounds in discovery, at the three phases of clinical development, and approved drugs. As an aside this analysis would fall foul of current Journal of Medicinal Chemistry author guidelines (see link; accessed 05-Aug-2025) which clearly mandate that “If average values are reported from computational analysis, their variance must be documented”.  As mentioned earlier in this post Figure 3 in LBH2009 shows statistically significant (P value < 0.001) differences between (a) discovery and phase 2 compounds (b) phase 1 and phase 3 compounds (c) phase 2 compounds & drugs. It’s also worth stressing that Figure 3 in LBH2009 does not show statistically significant differences in Fsp3 for any of the clinical development transitions (phase 1 to phase 2; phase 2 to phase 3; phase 3 to approved drug). Figure 3 of in LBH2009 shows 591 phase 2 compounds but only 376 phase 1 compounds, raising questions about the numbers of compounds that have been in clinical development without being recorded in the database.

I think that there are some problems with how the authors of the LBH2009 study have analysed the relationship between Fsp3 and progression through the stages of clinical development.  If charged with analysing this data I would focus on the three clinical development transitions (phase 1 to phase 2; phase 2 to phase 3; phase 3 to approved drug) and wouldn’t waste time on comparisons between discovery compounds and clinical compounds. If analysing the relationship between Fsp3 and the progression from phase 1 to phase 2, I would partition the set of phase 1 compounds into a ‘YES’ subset of compounds that had progressed to phase 2 and a ‘NO’ subset of compounds that had not progressed to phase 2. I would certainly be taking a close  look at distributions of Fsp3 values (some approaches to assessing statistical significance are based on the assumption of Normally-distributed data values) and I’d also be thinking about assessing effect size in addition to statistical significance. However, the problems with the LBH2009 analysis are more fundamental than non-Normal distributions of Fsp3 values.

The authors of LBH2009 assess the progression from phase 1 to phase 2 by comparing the mean Fsp3 value for the phase 1 compounds with the mean Fsp3 value for phase 2 compounds. The problem is that the Fsp3 values for the YES compounds (that have progressed from phase 1 to phase 2) are present in both the data sets for which comparisons are being made. This means that the observed differences in mean Fsp3 values will reflect both the difference between YES and NO compounds (relevant to relationship between Fsp3 and progression from phase 1 to phase 2) and the relative numbers of YES and NO compounds in the phase 1 data (not relevant to relationship between Fsp3 and progression from phase 1 to phase 2). Analysing the data in the way that the authors of LBH2009 have done effectively adds noise to the signal and it’s possible that they would have observed more statistically significant differences in mean Fsp3 values had they analysed the data in a more appropriate manner.

This is an appropriate point at which to discuss correlation in the context of studies such as LBH2009 and CNM2025. It’s actually well known (see L2013) that that Fsp3 values for chemical structures tend to be greater when amine nitrogen atoms are present (this does not invalidate the observed trends in the data but has big implications for how you interpret these trends). There is, however, a much bigger issue which is that correlation does not imply causation. Let’s suppose that you’ve just joined a drug discovery team as they are preparing to select a clinical candidate (I concede that this is most improbable scenario but it does illustrate a point). The team have an excellent understanding of the structure-activity relationship (SAR) and have successfully addressed a number of issues during the lead optimization process (the chemical structures of the compounds have been quite literally shaped by the problems that the team members have solved). Now consider the likely reaction of the team members to a suggestion that probability of success in the clinic would increase if the chemical structure of the best compound were modified so as to increase its Fsp3 value. My view is that the team might think that the person making such a suggestion had just stepped off the shuttle from Planet Tharg (an alien from this planet used to make occasional Sunday Sport  appearances). I see the trends in data observed by the authors of LBH2009 as effects rather than causes (the vanishing B-17 was never there in the first place).

Let’s return to the CNM2025 study and its authors state:

Using data from the Cortellis Drug Discovery Intelligence database, we repeated an analysis similar to that of Lovering et al. to assess Fsp3 in drugs approved post-2009 and those in active clinical development as of mid-2024 (Fig. 1). [I would challenge the claim that the analysis presented in CNM2025 is similar to that presented in LBH2009. The supplementary material for CNM2025 indicates that the data summarised in Fig. 1b correspond to the period 2012 through 2024 (it is not clear whether the database has been updated to account for compounds that have fallen out of active development during this period. As is the case for Figure 3 in LBH2009, Fig.1b in CNM2025 shows more phase 2 compounds (816) than phase 1 compounds (421), raising similar questions about the numbers of compounds that have been in clinical development without being recorded in the database. I thank fellow blogger  Dan Erlanson for suggesting that I examine the supplemental information for CNM2025.] Although our methods used contemporary data sources different to Lovering et al., we obtained comparable Fsp3 data for approved drugs prior to 2009. More recently however, the picture appears to have changed with approvals shifting to lower Fsp3 drugs (Fig. 1a). Similarly, when looking at drugs currently in clinical development (Fig. 1b), there appeared to be no clear relationship between highest phase reached and Fsp3, suggesting the key conclusion noted by Lovering et al. has not persisted. In all data sets, exemplars with Fsp3 = 0 as well as Fsp3 = 1 are extensively seen. [It is necessary to account for the number of hypotheses have been tested for statistical significance when quoting P-values (see R2016 and VM2018).]

Fig. 1a in CNM2025 shows the time-dependence of Fsp3 distributions for approved drugs according to approval date and I remain unconvinced of the value of analysis like this (on first encountering analysis of time-dependence of drug properties a quarter of a century ago I recall being left with the distinct impression that some senior medicinal chemists where I worked had a bit too much time on their hands). However, it is immaterial whether or not you are as underwhelmed as I am by time-dependence of drug properties because no such analysis is actually reported in LBH2009 and this is one reason that I challenge the claim by made by the authors of CNM2025 that they “repeated an analysis similar to that of Lovering et al. to assess Fsp3 in drugs approved post-2009 and those in active clinical development as of mid-2024”.  

Now let’s take a look at Fig. 1b in CNM2025 and this should be compared with Figure 3 in LBH2009. In some ways the former is an improvement on the latter since the violin plots show the distributions of Fsp3 values for each group of compounds and, as mentioned earlier in the post, I don’t think that it makes any sense to include discovery compounds in analysis like this (as the authors of LBH2009 did). Although these two figures look superficially similar they are actually very different and, given that the authors of CNM2025 only included "compounds in clinical trials as of mid-2024" in their study, I would argue that their study does not properly examine the link between Fsp3 and clinical progression. I agree that the difference between mean Fsp3 values for drugs approved up to 2009 and for drugs approved after 2009 is statistically significant. What is not clear from the analysis summarized in Fig. 1b in CNM2025 is whether the lower Fsp3 values of drugs that were approved after 2009 reflect smaller increases in Fsp3 over the course of clinical development (the B-17 has disappeared from the lunar crater) or lower Fsp3 values for compounds entering clinical development (the B-17 is still in the lunar crater). I think it's possible to address this question but you would need to analyse the data a lot more carefully than the authors of CNM2025 appear to have done. For example, you might examine the time-dependencies of mean Fsp3 values for compounds evaluated in phase 1 and the corresponding mean Fsp3 values for compounds that progressed or failed to progress to phase 2. While I consider more careful analysis of progression to be feasible I see little or no value from the perspective of real world drug discovery in actually performing the analysis more carefully.

This is a good point at which to wrap up and, unless the the trends in the data can shown to reflect causation, the debate can be described as bald men fighting over a comb (as one who is follicly challenged I always find it painful to use this phrase). I see variation in drug properties with time as an effect rather than a cause and Forrest Gump would have been well aware of this fifteen years before the publication of LBH2019 when he famously observed that "shit happens". One point on which the CNM2025 authors and I do appear to agree is that there is not currently a B-17 in a lunar crater. Where we appear to differ is that they seem to be suggesting this was because it has vanished while I never believed that it was ever there in the first place. I’ll let the late great Dave Allen have the last word.

Tuesday, 31 December 2024

Natural Intelligence?

**********************************************************
My pulse will be quickenin'
With each drop of strychnine
We feed to a pigeon
It just takes a smidgin
To poison a pigeon in the park

Tom Lehrer, Poisoning Pigeons in the Park | video
*********************************************************

I’ll be reviewing the H2024 study (Occurrence of “Natural Selection in Successful Small Molecule Drug Discovery) in this post. Derek has already posted on the H2024 study which has been included in the BL2024 Virtual Special Issue on natural products (NPs) in medicinal chemistry. I'll also mention reviews here at Molecular Design of the the related studies (4) (see post) and (24) (see post). As is usual for Molecular Design reviews of literature I have used the same reference numbers that were used in H2024 and quoted text is indented with any comments by me in square brackets and italicised in red. Given the serious concerns I have about H2024 this is going to be a long post and there are a couple of disclaimers that I need to make before starting the review:

  1. I regard identification and biological characterisation of NPs as vital scientific activities that should be generously funded and Derek puts it very well in his recent post ("When you see specific and complex small molecules that living creatures are going to the metabolic trouble to prepare, there are surely survival-linked functions behind them."). In particular, I see it as important that NPs be screened in diverse phenotypic assays and here’s a link to the Chemical Probes Portal. While my criticisms of H2024 are certainly serious it would be grossly inaccurate to take these criticisms as indicative of an anti-NP position.
  2. Automation of workflows (N2017) and generation of datasets from databases such as ChEMBL are far from trivial and (33), which highlights some of the challenges faced by researchers in this area, was the subject of a recent post at Molecular Design. I consider method development in this area to be an important cheminformatic activity that should be adequately supported. It must also be stressed that the design, building and updating of databases such as ChEMBL (G2012 | B2014 | P2015 | G2017 | 23) are vital scientific activities that should be generously funded (had it not been of the vision and foresight of the creators of the PDB over half a century ago it is improbable that the 2024 Chemistry Nobel Prize would have been awarded for “computational protein design” and “protein structure prediction”). While my criticisms of H2024 are certainly serious it would be grossly inaccurate to take these as criticisms of the automated dataset generation described in the study (and recently published in H2024b) or of the contributions by a number of individuals that have made ChEMBL an invaluable resource for drug discovery scientists and chemical biologists.

Hampi, November 2013

Having made the disclaimers, I’ll open my review of H2024 with some general observations. First, I do not consider that H2024 presents any insights of practical value to medicinal chemists nor do I consider the analyses presented in the study to support the assertion that “there is untapped potential awaiting exploitation, by applying nature’s building blocks─’natural intelligence’─to drug design” (in my view the use of the term “natural intelligence” does rather endow the study with what I’ll politely refer to as a distinctly pastoral odour). Second, the results of the analyses presented in H2024 do not demonstrate any tangible benefits from the drug design perspective of incorporating structural features that have been anointed as 'natural' by the authors (my view is that it would be extremely difficult to design data analyses to address the relevant questions in an objective manner). Third, the authors of H2024 present a ‘scaffold-centric’ view of NPs in which the naturalness of NPs is due to cyclic substructures present within their chemical (2D) structures (it is almost as if these 'natural' substructures are considered to be infused with 'vital force') and I would question whether this is a realistic view from the molecular recognition and physicochemical perspectives.  Fourth, the meaning of what the authors of H2024 are calling 'enrichment' of pseudo-NPs (PNPs) in clinical compounds is unclear and, in any case, the 'enrichment' values do seem rather low (never more than twofold) when you consider the numbers of compounds that successful discovery project teams typically have to synthesize in order to deliver a drug that gets to market.

It's not clear (at least to me) what the authors of H2024 mean by ‘natural selection’ and at times their view of natural selection appears to be closer to Lysenkoism than Darwinism. For example, they assert in the conclusions section of H2024 that “NP structural motifs are provided predesigned by nature, constructed for biological purposes as a result of 4 billion years of evolution.” Design actually has no place in natural selection and perhaps the authors are thinking of 'Intelligent Design' which is a doctrine with many adherents in the Creationist community.  While I don’t dispute that the chemical structures of many clinical compounds contain substructures that are also found in the chemical structures of NPs, I think that it would be extremely difficult to objectively compare different explanations for the observations (it's worth remembering that correlation does not imply causation). The explanation favoured by the authors of H2024 is that compounds assembled from Nature’s building blocks are ‘better’ and a stated aim of the study is “to seek further support for the existence of ‘natural selection’ in drug discovery” (this video will give readers an idea of what the late great Dave Allen might have made of this). In my view the data analyses presented in H2024 are not actually based on statistics and are therefore unfit for the purpose of testing hypotheses. Put another way, if you're going to use data analysis to look for something then it would be a good idea to use methods capable of telling you that you that haven't found what you were looking for.    
 
The data analyses in H2024 are largely based on quantities (PNP_Status | Frag_coverage_Murcko | NP-likeness) that are calculated from the chemical (2D) structures of compounds.  However, the authors do not state which software was used to perform the calculations and, had I been a reviewer, I would have drawn their attention to the following directive in the Data Requirements section in the J Med Chem Author Guidelines (accessed 27-Dec-2024):

9. Software. Software used as a part of computer-aided drug design should be readily available from reliable sources, and the authors should specify where the software can be obtained.

As was the case for my review of (24) I see much of the analysis in H2024 as relatively harmless “stamp collecting” (in contrast, as discussed in KM2013, I consider presentations and analyses of data that exaggerate trend strength, such as those used in the HMO2006LS2007, LBH2009, HY2010 and TY2020 studies to be anything but harmless). The analyses that I’ll be examining in this post are of comparisons between clinical compounds and reference compounds although I'll comment in general terms on the analyses of time-dependencies of characteristics of clinical compounds. My general criticism of H2024 is not that the analyses presented by its authors are necessarily invalid but that they fail to provide any useful insight and I’ll share an insightful observation by Manfred Eigen (1927-2019):

A theory has only the alternative of being right or wrong. A model has a third possibility: it may be right, but irrelevant.

I first encountered analyses of time-dependencies of drug properties about two decades ago and rapidly came to the conclusion that some senior medicinal chemists where I worked had a bit too much time on their hands.  The fundamental flaw in the interpretation of these analyses is that time-dependencies of the properties of drugs and other clinical compounds are presented as causes rather than effects and it has never been clear how medicinal chemists working on drug discovery projects in the real world should use the results from such analyses. The authors claim that “changes to drug properties over time are significant” and I would challenge them to present even a single example of such analysis being used to meaningfully inform decision-making in a drug discovery project. It must be stressed that my criticism of analyses of time-dependency of the properties of drugs and other clinical compounds is simply that they don't provide useful insights and not that the analyses are necessarily invalid. That said, I do have general concerns about how time-dependencies are compared when some of the properties are expressed as logarithms and some are not. As reviewer I would have recommended that the vertical axis of the plot in the graphical abstract be drawn from 0% to 100% rather than from 30% to ~67%.

As is the case for analyses of time-dependency, my criticism of analyses of the differences between clinical compounds and reference compounds is that they don’t provide useful insight and there is no suggestion that the analyses are necessarily invalid. Before looking at the analyses presented in H2024 I’ll quote from the abstract of (24) because this will give you an idea of what I mean by analyses not providing useful insight:

Drugs are differentiated from target comparators by higher potency, ligand efficiency (LE), lipophilic ligand efficiency (LLE), and lower carboaromaticity.

As I noted in this post (this focused principally on the invalidity of the LE metric as discussed in NoLE) reporting that an analysis has shown drugs to be differentiated by potency from target comparators does seem to be stating the obvious and, given how LE and LLE are defined, it is perhaps not the most penetrating of insights to observe that values of these efficiency metrics tend to be greater for drugs than for comparator compounds. While the observation of lower carboaromaticity of drugs relative to comparator compounds is non-obvious, it does not constitute information that can be used for medicinal chemistry decision-making in specific discovery projects (as we noted in KM2013 carboaromaticity and lipophilicity can both be reduced simply by replacing a benzene ring with benzoquinone).

Let’s take a look at how this type of analysis is used in H2024. The authors of H2024 note that “comparing Figure 3a,b shows a clear ‘enrichment’ of PNPs in clinical compounds versus reference compounds in the post-2008 period” and two of these authors, writing in (17), assert that “PNPs have increasingly been explored in recent drug discovery programs, and are strongly enriched in clinical compounds”.  What the authors of H2024 are calling 'enrichment' is rather different to the enrichment in structural features that results from high-throughput screening (HTS) and it’s important to understand the difference. Let’s suppose that we’ve screened a library of compounds of which 1% are pyrimidines and 1% are pyrazines and we find that 10% of the hits are pyrimidines and 0.1% are pyrazines (to simplify things you can assume there is no compound in the library with a pyrimidine and a pyrazine in its chemical structure). In this case we would conclude that the process of screening has resulted in a tenfold enrichment for pyrimidines and a tenfold impoverishment for pyrazines. Now let's create a 'selected azines' category by combining the pyrimidines and pyrazines which as a structural class comprise 2% of the screening library compounds but 10.1% of the hits. What I'm getting at here is that enrichment of an more inclusive structural class such as 'selected azines' (or PNPs) does not imply that each and every one of the structural classes covered by the inclusive structural class definition will also be enriched.

Now let’s take a look at how the 'enrichment' of PNPs in clinical compounds is assessed in H2024. First, a set of reference compounds is generated for each clinical compound (this is discussed in detail in H2024b) and the sets of reference compounds are combined. 'Enrichment' is then assessed by comparing the fraction of clinical compounds that are PNPs with the fraction of compounds in the combined reference sets that are PNPs. When we assess enrichment of chemotypes in HTS the hits are all selected (by the screening process) from the same reference pool of compounds. In contrast, each clinical compound in the H2024 analysis is associated with a different reference set of compounds (from the perspective of data analysis combining reference sets defined in this manner gratuitously throws information away). As a reviewer I would have pressed the authors to enlighten readers as to how they should interpret the proportions of PNPs in the reference sets for individual compounds.

It's worth thinking about what the reference compound set might look like for a clinical compound that is a PNP. The proportion of PNPs in the reference set will generally be influenced by factors such as availability of data, the ‘rarity’ of the structural features of the drug and the ‘tightness’ of the structure-activity relationship (SAR).  A more permissive definition of ‘activity’ would generally be expected to make SAR appear to be less ‘tight’ (or ‘looser’ if you prefer). Compounds were defined as ‘active’ for the analysis on the basis of a recorded pChEMBL value against one of the clinical compound’s targets (as a reviewer I’d have suggested that the authors define the term ‘pChEMBL’) which means that a compound might have been selected for inclusion in a reference set on the basis of an IC50 value of 100 μM.

Let’s define 'enrichment' by dividing the fraction of the clinical compounds that are PNPs by the fraction of reference compounds that are PNPs. When we select a reference set for a clinical compound that is a PNP then it’s extremely unlikely that every single compound in the reference set will also be a PNP (especially if we’re accepting compounds with IC50 values 100 μM as ‘active’) and it’s even less likely that every single compound in the combined reference sets will be a PNP. This means that we should generally expect the clinical compounds that are PNPs to be ‘enriched’ in PNPs when compared with their combined reference sets. We can apply exactly the same logic to conclude that we should expect that the combined reference sets for the clinical compounds that are not PNPs  (under this scenario we would conclude that the set of clinical compounds that are not PNPs are infinitely impoverished in PNPs when compared with their combined reference sets). This means that we should expect that the 'enrichment' of PNPs in the clinical compound set in comparison with their combined reference sets will increase with the fraction of clinical compounds that are PNPs.

Let’s take another look at the plot in the graphical abstract which shows the fractions of clinical compounds and reference compounds that are PNPs as a function of time. Notice how the lines tend to be furthest apart when the fraction of clinical compounds that are PNPs is relatively high. As a reviewer, I would have required that the authors examine the correlation between the logarithm of the fraction of clinical compounds and the logarithm of the enrichment (a relatively strong correlation would indicate that the information added by the combined reference sets is minimal). The 'enrichments' calculated from the plot in the graphical abstract are underwhelming (the highest degree of enrichment is the 2014 value of just over 1.5-fold and this value seems very low when you consider the numbers of compounds that successful discovery project teams typically need to synthesize in order to get drugs approved).  From 2011 the fraction of clinical compounds that are PNPs exceeds 50% but I wouldn't consider it accurate to use the term "strongly enriched" (17) because the fraction of reference compounds that are PNPs is 40% or greater for this time period (plotting the vertical axis in the graphical abstract from 30% to ~67%  creates the illusion that the 'enrichment' is greater than it actually is).

I do have a number of other gripes about the data analysis in H2024 but I do also need to take a look at PNPs and the following assertion by the authors is an appropriate point at which to start this discussion:

The PNP concept has been validated by its appearance in the literature (16,17) and by the design of several new classes of biologically active compounds. (18,19) [As a reviewer I would have pressed the authors to clearly articulate the “PNP concept” (just as I would have pressed the authors of this Editorial to clearly articulate the new principles that their nominees for the Nobel Prize in Physiology or Medicine had introduced).  My view is that it is verging on megalomania to claim that a concept “has been validated by its appearance in the literature” and I don’t consider (18) to support the claim for “design of several new classes of biologically active compounds”. To support such a claim, one would ideally need to demonstrate that screening of libraries of compounds designed as PNPs resulted in the discovery of viable lead series against a range of therapeutic targets. At absolute minimum, one would need to show that libraries of compounds designed as PNPs exhibited exploitable activity across a range of target-related assays (although interesting, the results from the “cell painting assay” would not by themselves support a claim for “design of several new classes of biologically active compounds”). I should also mention that some in the compound quality field (see B2023 and my review of that article) interpret activity against multiple targets for a set of compounds based on a particular scaffold as evidence for pan-assay interference even when the individual compounds don’t themselves exhibit frequent-hitter behaviour. I don't have access to (19) and am therefore unable to assess the degree to which that article supports the authors claim for “design of several new classes of biologically active compounds”.]

The PNP status of a compound is determined by how “NP library fragments” (these are cyclic substructures extracted from the chemical structures of compounds in an NP-focussed screening library that had been generated over a decade ago for fragment-based drug discovery) are combined in its chemical structure.
 
PNP_Status. Compounds were assigned to one of four categories according to their NP fragment combination graphs. (16,17) The NP library fragments used for this purpose are Murcko scaffolds (26) [It would be actually more appropriate to refer to these as ‘Bemis scaffolds’ in order to properly recognize the corresponding author of this article.] (the core structures containing all rings without substituents except for double bonds, n = 1673) derived (16) from a representative set of 2000 NP fragment clusters. (15) [I see this approach as unlikely to capture all the relevant cyclic substructures present in NPs.  My view is that it would have been better to first extract the relevant cyclic substructures from the chemical structures of all NPs for which this information is available, and then do the selection and filtering in one or more subsequent steps. The other advantage of doing things this way is that you’ll get a better assessment of the frequencies with which the different cyclic substructures occur in the chemical structures of NPs.]  Because of their ubiquitous appearances in NPs, the phenyl ring and glucose moieties were specifically excluded as fragments. (16) [I would expect exclusion of the benzene ring (I consider ‘benzene ring’ more correct than ‘phenyl ring’ in this context) as a fragment to result is a significant reduction in number of the number of compounds that are considered to be PNPs (and, by implication, the ‘enrichment’ associated with membership of the PNP class).  Even though the benzene ring has been excluded for the purpose of assigning PNP status it should still be considered to be one of Nature’s building blocks.]

As I mentioned earlier in the post, the view of NPs presented in H2024 is ‘scaffold-centric’ and I would question how realistic this view is given that non-scaffold atoms at the periphery of a molecular structure will generally be more exposed to targets (and anti-targets) than scaffold atoms at the core of the molecular structure. What I’m getting at here is that it is far from clear how much of a compound’s pharmacological activity can be attributed to the presence of individual substructural features in the chemical structure of the compound (modifying a point made in NoLE, I would argue that the contribution of a structural feature to the binding affinity of a compound is not actually an experimental observable). This is one reason that unless matched molecular pairs are available it would not generally be possible to demonstrate the superiority of one structural feature over another in an objective manner.

Something that you need to pay very close attention to when extracting substructures from chemical structures of compounds is the ‘environment’ of the substructure (I prefer to use the term ‘substructural context’). For example, two piperidine rings linked through nitrogen look very different from the perspective of a therapeutic target protein depending on whether the link is a carbonyl carbon or a tetrahedral carbon (most medicinal chemists will be aware that the protonation states differ but there are also subtle, although still significant, differences in the shape of the piperidine ring in the two substructures). You also need to be aware that fusing rings can have profound effects on physicochemical characteristics and I would consider it a bad idea to extract monocyclic substructures from fused or bicyclic ring systems.

There are some things that don't look quite right and I would have flagged these up if I’d been reviewing the manuscript. Let’s take a look at the first entry (Sotorasib) in Table 1 and you can see that the oxygen of the 2-pyrimidone substructure is coloured lilac indicating that this substructure can be found in the chemical structures of one or more NPs (I would still challenge the view that the result of fusing 2-pyrimidone with pyridine should be considered 'natural' on the basis that the heterocycles from which it is derived from are both found in chemical structures of NPs). Now take a look the second entry (Dolutegravir) in Table 1 and you'll notice that the oxygen in the 4-pyridone substructure is not coloured green. This implies that 4-pyridone does not occur in the chemical structure of any NP and, in the absence of  information, I can only assume that it has been anointed as 'natural' because of its structural analogy with pyridine (while there is a nitrogen atom and five trigonal carbon atoms in each substructure the molecular recognition characteristics of the two substructures differ far too much for them to be regarded as equivalent from the perspective of assigning PNP status). Six of the substructures in Figure 5 appear to be in unstable tautomeric forms (first, fifth, ninth, twelfth entries in line 2 | seventh entry in line 3 | first entry in line 5).   

I'll conclude my review of  H2024 by commenting on claims made by the authors:

This is further evidence that the three NP metrics can be considered as independent measures of clinical compound quality. [I would consider the claim that any of these “NP metrics” can be considered as a measure of“clinical compound quality” to be wildly extravagant (the authors haven't even stated how "clinical compound quality" is defined yet they claim to be able to measure it). I would argue that compound quality cannot be meaningfully compared for clinical compounds that have been developed for different diseases or disorders. Describing a compound as 'clinical' implies that a large body of measured data has actually been generated for it and the authors of H2024 might find it instructive to ask themselves why they think a simple metric calculated from the chemical structure of the compound would be of interest to a project team with access to this large of body of measured data One criticism that I make of drug discovery metrics is that they trivialize drug discovery and we noted in KM2013: “Given that drug discovery would appear to be anything but simple, the simplicity of a drug-likeness model could actually be taken as evidence for its irrelevance to drug discovery.” ]

The overall results are supportive of the occurrence of “natural selection” being associated with many successful drug discovery campaigns. [My view is that the authors of H2024 have not clearly articulated what they mean by“natural selection” in the context of this study.]  It has been proposed that NP-likeness assists drug distribution by membrane transporters, (21) [The author of (20c) asserts "Over the years, my colleagues and I have come to realise that the likelihood of pharmaceutical drugs being able to diffuse through whatever unhindered phospholipid bilayer may exist in intact biological membranes in vivo is vanishingly low" and, by implication, that entry of the vast majority of drugs into cells is transporter mediated. I keep an open mind on this issue although I note that what is touted by some as a universal phenomenon does seem to have been remarkably difficult to observe directly by experiment. The difficulties caused by active efflux are widely recognized by drug discovery scientists and it may be instructive for the authors of H2024 to consider how an experienced medicinal chemist working in the CNS area might view a suggestion that compounds should be made more like NPs to increase the likelihood of being transporter substrates.] and we further speculate that employing NP fragments may result in less attrition due to toxicity, a major cause of preclinical failure. (55[This does seem to be grasping at straws. The focus of the cited article is actually clinical failure and not preclinical failure.]

There is untapped potential for further exploitation of currently used and unused NP fragments, especially in fragment combinations and the design of PNPs, without the need to resort to chemically diverse ring systems and scaffolds. [This exemplifies what can be called the ‘Ro5 mentality’ (‘experts’ advising medicinal chemists to not explore but to focus on regions of chemical space that have been blessed by the ‘experts’). As I note in this blog post Ro5 (as it is stated) is not actually supported by data and in NoLE, I advise drug designers not to “automatically assume that conclusions drawn from analysis of large, structurally-diverse data sets are necessarily relevant to the specific drug design projects on which they are working.” An equally plausible 'explanation' for the observation that a high fraction of clinical compounds are PNPs is simply that medicinal chemists are working with what they're most familiar with (in this case the advice would be to look beyond Nature's building blocks for inspiration).] To exploit these opportunities, “NP awareness” needs to be added to the repertoire of medicinal chemists. [My view is that it would be more important for critical thinking to be added to the repertoire of medicinal chemists so they are better equipped to assess the extent to which conclusions and recommendations of studies like H2024 are actually supported by data.]

In short, applying nature’s building blocks─natural intelligence─to drug design can enhance the opportunities now offered by artificial intelligence. [In my view "natural intelligence" appears to be arm-waving that is neither natural nor intelligent.]  

This is a good point to wrap up and to also conclude blogging for the year. My new year wish is for a kinder, happier and more peaceful World in 2025 and I'll leave you with a photo of BB and Coco in the study here in Maraval. They had been helping me with this post before I unwisely decided to explain ligand efficiency to them. Let sleeping dogs lie I guess.


 

Thursday, 8 June 2023

Archbishop Ussher's guide to efficient selection of development candidates

One piece of advice I gave in NoLE is that “drug designers should not automatically assume that conclusions drawn from analysis of large, structurally-diverse data sets are necessarily relevant to the specific drug design projects on which they are working” and the L2021 study that I’m reviewing in this post will give you a good idea of what I was getting at when I wrote that. I see a fair amount of relatively harmless “stamp collecting” in L2021 but there are also some rather less harmless errors of the type that you really shouldn’t be making if cheminformatics is your day job.  

I’ll start the review of L2021 with annotation of the abstract:

"Physicochemical descriptors commonly used to define ‘drug-likeness’ and ligand efficiency measures are assessed for their ability to differentiate marketed drugs from compounds reported to bind to their efficacious target or targets. [I would argue that differentiating an existing drug from existing compounds that bind to the same target is not something that medicinal chemists need to be able to do. It is also incorrect to describe efficiency metrics such as LE and LLE as physicochemical descriptors because they are derived from biological activity measurements such as binding affinity or potency.] Using ChEMBL version 26, a data set of 643 drugs acting on 271 targets was assembled, comprising 1104 drug−target pairs having ≥100 published compounds per target. Taking into account changes in their physicochemical properties over time, drugs are analyzed according to their target class, therapy area, and route of administration. Recent drugs, approved in 2010−2020, display no overall differences in molecular weight, lipophilicity, hydrogen bonding, or polar surface area from their target comparator compounds. Drugs are differentiated from target comparators by higher potency, ligand efficiency (LE), lipophilic ligand efficiency (LLE), and lower carboaromaticity. [I may be missing something but stating that drugs tend to differ in potency from non-drugs that hit the same targets does rather seem to be stating the obvious. The same point can also be made about efficiency metrics such as LE and LLE since these are derived, respectively, by scaling potency with respect to molecular size and offsetting potency with respect to lipophicity (LLE).] Overall, 96% of drugs have LE or LLE values, or both, greater than the median values of their target comparator compounds.” [What is the corresponding figure for potency?]

I must admit to never having been a fan of drug-likeness studies such as L2021 (when I first encountered analyses of time dependency of drug properties about 20 years ago I was left with an impression that some senior medicinal chemists had a bit too much time on their hands) and it is now ten years since the term "Ro5 envy" was introduced in a notorious JCAMD article. My view is that the data analysis presented in L2021 has minimal relevance to drug discovery so I’ll be saying rather less about the data analysis than I’d have done had J Med Chem asked me to review the study.

The L2021 study examines property differences between marketed drugs and compounds reported to bind to efficacious target(s) of each drug. Specifically, the property differences are quantified by difference between the value of the property for the drug and the median of the values of property for the target comparator compounds. If doing this then you really do need to account for the spread in the distribution if you’re going to interpret property differences like these (a large difference in values of a property for the drug and the median property for the target may simply reflect a wide spread in the property distribution for the target).  However, I would argue that a more sensible starting point for analysis like this would be to locate (e.g., as a percentile) the value of each drug property within the corresponding property distribution for the target comparator compounds.

Let’s take a look now at how the authors of L2021 suggest their study be used.  

“This study, like all those looking at marketed drug properties, is necessarily retrospective. Nevertheless, those small molecule drug properties that show consistent differentiation from their target compounds over time, namely, potency, ligand efficiencies (LE and LLE), and the aromatic ring count and lipophilicity of carboaromatic drugs, are those that are most likely to remain future-proof. Candidate drugs emerging from target-based discovery programs should ideally have one, or preferably both, of their LE and LLE values greater than the median value for all other compounds known to be acting at the target.”

I would argue that the L2021 study has absolutely no relevance whatsoever to the selection of compounds for development since the team will have data available that enables them to rule out the vast majority of the project compounds for nomination.  A discovery team nominating a compound for development will have achieved a number of challenging objectives (including potency against target and in one or more cell-based assays) and the likely response of team members to a suggestion that they calculate medians for LE and LLE for comparison with nomination candidate(s) is likely to be bemused eye-rolling. In general, a discovery team nominating a development candidate has access to a lot of unpublished potency measurements (which won’t be in ChEMBL) and it’s usually a safe assumption that the development candidate will be selected from the most potent compounds (LE and LLE values for these compounds are also likely to be above average). In the extremely unlikely event that the discovery team nominates a compound with LE or LLE values below the magic median values then you can be confident that the decision has been based on examination of measured data (consider the likelihood of the discovery team members acting on a suggestion that they should pick another compound with LE or LLE value above the magic median values because doing so will increase the probability of success in clinical development).   

As the start of the post, I did mention some errors that you don’t want to be making if cheminformatics is your day job and regular readers of this blog will have already guessed that I’m talking about ligand efficiency (LE). I should point out l that the problem is with the ligand efficiency metric and not the ligand efficiency concept which is both scientifically sound and useful, especially in fragment-based design where molecular size often increases significantly in the hit-to-lead phase. 

The problem with the LE metric is that perception of efficiency changes when you express affinity (or potency) using a different unit and this is shown clearly in Table 1 in NoLE. Expressing a quantity using a different unit doesn’t change the quantity so any change in perception is clearly physical nonsense. That’s why I appropriate a criticism (it’s not even wrong) usually attributed to Pauli when taking gratuitous pot shots at the LE metric.  The change in perception is also cheminformatic nonsense and that’s why it’s rather unwise to use the LE metric if cheminformatics is your day job. L2021 does cite NoLE but simply notes the LE metric’s “scientific basis and application have provoked a literature debate”.

The L2021 study asserts that “the absolute LE value of a drug candidate is less important” but the problem is that even differences in LE change when you express affinity (or potency) using a different concentration unit. This is shown in Table 2 in NoLE and the problem is that there is no objective way to select a particular concentration unit as ‘better’ than all the other concentration units.  To conclude, can we say that a medicinal chemistry leader’s choice of concentration unit (1 M) is any better (or any worse) than that of Archbishop Ussher (4.004 μM)?