Showing posts with label DDT. Show all posts
Showing posts with label DDT. Show all posts

Sunday, 30 November 2025

Covalent ligand efficiency

It appears that whoever first described Economics as the ‘Dismal Science’ had never encountered a ligand efficiency metric. I’ll be taking a look at the FK2025 study (Covalent ligand efficiency) in this post and the study has already been reviewed by Dan. Something that I’ve observed repeatedly over the years is that authors of ligand efficiency studies exhibit a lack of understanding of units and dimensions associated with physicochemical quantities that would shame a first year undergraduate studying introductory physical chemistry (this is somewhat ironic given that creators of ligand efficiency metrics frequently tout their creations in physicochemical terms). I consider covalent ligand efficiency (CLE) as defined in the FK2025 study to have no value whatsoever for design of drugs that bind irreversibly to their targets through covalent bond formation given that the metric is time-dependent and based on an invalid measure of bioactivity. The formidable Lady Bracknell is clearly unimpressed and I should mention that the photo is from the wikipedia page for English actress Rose Leclercq (1843-1899). Given the serious deficiencies in the FK2025 study this is going to be a long and tedious post (even more so than usual 😁😁😁) so please ensure that you have strong coffee close to hand. As is usually the case in posts here I've used the same reference numbers as were used in FK2025 and quoted text is indented with my comments in red italics. I've organized some of the mathematical material into three tables and references to tables in the post are to these (and not to any of the tables in the FK2025 study).    

Before starting the review of FK2025 it’s worth examining irreversible covalent inhibition from a molecular design perspective and I’ll direct readers to the informative S2016 and McW2021 reviews, and the recent L2025 study which presents COOKIE-Pro for covalent inhibitor binding kinetics profiling on the proteome scale. Covalent bond formation between RNA and ligands can also be exploited (S2025 | K2025 | L2015)  and I generally use 'target' rather than 'protein' in blog posts and journal articles. An irreversible covalent inhibitor acts by first binding non-covalently to its target in the first step with the covalent bond forming in the second step between an electrophilic ligand atom (the term warhead is commonly used) and a nucleophilic target atom such as the sulfur atom of a cysteine. A commonly used measure of activity for irreversible covalent inhibitors is the kinact/KI ratio which can be thought of as the product of affinity (1/KI) and reactivity (kinact). In design of irreversible covalent inhibitors we try to place the electrophilic atom of the warhead within reacting distance of the nucleophilic atom of the target (this is relatively easy if you have a reliable structure of a complex of the target with a relevant ligand that lacks the electrophilic warhead). The non-covalent complex between target and ligand is stabilised by the non-covalent contacts between the target and ligand (the term ‘molecular interactions’ is also used although I prefer to think in terms of ‘non-covalent contacts’ since the latter can be observed experimentally). However, non-covalent contacts also determine reactivity of the non-covalently bound complex by stabilising the transition state (I consider it more correct to think in terms of reactivity of the complex than in terms of reactivity of either the electrophilic warhead or the target nucleophile). In the design context, this means attempting to tune non-covalent contacts to stabilise the transition state to a greater extent than the non-covalent complex.

The LE and CLE metrics share a very serious deficiency in that your perception of efficiency can be altered if you change the value of an arbitrary term in the formula for the metric and I'll start the review of FK2025 by critically examining LE. The meaningless of LE stems from a fundamental misunderstanding of how logarithms work and I'll by point you toward M2011 (Can one take the logarithm or the sine of a dimensioned quantity or a unit? Dimensional analysis involving transcendental functions) that was published in the Journal of Chemical Education. In drug discovery we frequently need to calculate logarithms for quantities and you need to be aware you can’t calculate the logarithm for a dimensioned quantity. Let’s take pIC50 as an example and this quantity is commonly defined as the negative logarithm of the IC50 in mole per litre (M). However, what you actually do when you calculate pIC50 is that you take the negative logarithm of the numerical value of the IC50 when expressed in mole per litre (this is a bit of a mouthful and it can be written more compactly as equation 1 below). While not denying that it is useful to have a convention such as this for expressing potency values logarithmically it should be remembered that the choice of mol per litre (M) is entirely arbitrary and it would be equally correct to use other valid concentration units such as μM or nM. One consequence of choosing mole per litre (M) for expressing IC50 values is that pIC50 values (or at least measured pIC50 values) will generally be positive because of the extreme difficulty of measuring meaningful IC50 values that are greater than 1 M. 


Let’s take a look at the binding free energy ΔG° and you’ll notice that I’ve written it with a degree symbol which indicates that this quantity corresponds to a standard state defined by a concentration value C° (the standard concentration). Equation 2 shows how the binding free energy is defined as the difference in chemical potential between the associating species (target + ligand) and the target-ligand complex with each species at the standard concentration (the degree symbol indicates that that both the binding free energy and chemical potential depend on the value of C° and I’ve also shown this explicitly in the equation although this is not actually necessary). Equation 3 shows the dependence of chemical potential on the concentration C of the species and the standard concentration C°. Taken together, Equation 2 and Equation 3 should clarify the origins of the dependence of binding free energy on the standard concentration comes from (there are two associating species but only one complex). We can’t actually measure binding free energy directly but we can calculate it from the dissociation constant KD using Equation 4 (which can be derived from Equation 2 and Equation 3). It’s important to be aware that if you use Equation 4 to convert ΔG° values between different values of the standard concentration C° you’ll be making the assumption that solutions are dilute (ΔH is independent of concentration) and this is indicated in Equation 5.

The standard concentration is a source of much confusion in the ligand efficiency field and I’ll direct readers to ‘The Nature of Ligand Efficiency’ (p9). While the standard concentration is integral to a valid thermodynamic treatment of target-ligand binding the value of C° is entirely arbitrary (to suggest otherwise would mean that you’ve abandoned thermodynamics). It is conventional in drug discovery (and biochemical) literature to use a C° value of 1 M when reporting ΔG° values. While this convention is certainly beneficial, it is no more (and no less) valid to use a value of 1M than it is to use a value of 1 μM for this purpose. Furthermore a standard state defined by a C° value of 1 M is not biophysically realistic (consider the difficulty of accommodating a mole of target in a volume of 1 litre and the likelihood of a ligand exhibiting aqueous solubility of 1 M). I assume that most biochemists and biophysicists would agree that it is generally not feasible to measure KD values of greater of 1 M (I would be happy to be proven wrong on this point) and this means that drug discovery scientists tend to assume that ΔG° values are necessarily negative.

Let’s now a take a look at ligand efficiency (LE) and you can see from the photo above that some heretics regard the metric as physically nonsensical (if you're interested in how I came to be chatting with fellow blogger Ash then take a look at this post). The LE metric which is regarded as an article of  faith in the fragment-based design community was introduced in the (p5) study with the symbol Δg (see Equation 6 below) and the authors of that study did not actually state that it had to be calculated using a C° value of 1 M (I consider it unlikely that any of the authors were even aware of the dependence of ΔG° on C°). In The Nature of Ligand Efficiency (p9) I defined the quantity ηbind (see Equation 7) by dividing Δg (LE) by RT (when LE values are quoted the molar energy units are usually discarded and T often does not correspond to the temperature at which the assay was run) and by the factor (2.303) used to convert between natural logarithms and base 10 logarithms. The quantity ηbind is directly proportional to Δg (LE) and using it makes it much easier to see how using a different standard concentration can alter your perception of efficiency.  Take a look at Table 1 in (p9) and you’ll see that the three compounds (a fragment, a lead and a clinical candidate) bind with equal efficiency when C° is 1 M. Change C° to 0.1 M and the clinical candidate is binds more efficiently than the fragment but when C° to 10 M the fragment becomes more ligand-efficient than the clinical candidate. As noted in (p9) “In thermodynamic analysis, a change in perception resulting from a change in a standard state definition would generally be regarded as a serious error rather than a penetrating insight.” 

Here's what the authors of FK2025 say about LE: 

LE depends on the choice of the standard concentration (normally 1 M) (p8) (p9) and its maximal available value is size dependent.(p10) (p11) [It's true that LE depends on C° but it’s also true that ΔG° depends on C° and the difference in the two dependencies is that is that LE “depends upon the choice of standard concentration in a nontrivial fashion” (p8). The issue is not the so much that LE depends on C° but that using a different unit to express KD changes how we perceive efficiency. The ΔΔG values that determine perception of affinity don’t change when you use a different value of C° (equivalent to using a different unit to express KD). However, if you use a different value of C° for calculating LE you can see from Table 1 in (p9) that even the ordering of LE values between two ligands can change. I consider the molecular size dependencies of LE observed by the authors of (p10) and (p11) to be artefactual and I’ll point you toward Fig. 1 in (p9) which shows that using a different value of C° can change how we perceive the molecular size dependency of LE.] Nevertheless, LE is an established tool to normalize potency and facilitate the comparison of ligands with a range of potencies and sizes. [It is not uncommon for adherents of religions to consider their beliefs to be established facts.] The usefulness of LE and other efficiency metrics in drug discovery has been extensively analyzed and reviewed elsewhere. (p6) (p12) (p13) (p14) (p15) (p16) (p17) [My view is that nobody has actually demonstrated the usefulness of LE and I’m unconvinced that it would even be possible to do so meaningfully in an objective manner (consider the feasibility of comparing success rates between a group of individuals using LE in discovery projects and a control group of individuals not using LE in discovery projects). Usefulness means that using something provides demonstrable benefits and ‘widely-used’ is not equivalent to ‘useful’ (I’m guessing that more people use homeopathic ‘medicines’ than use ligand efficiency metrics). One piece of advice that I’ll offer to anybody advocating the use of LE in drug design is to ensure that you fully understand the implications of changes in perception resulting from using different units to express quantities not least because you might find yourself lecturing to people who do understand.]

After a lengthy preamble it’s now time to take a look at how the FK2025 study addresses ligand efficiency in the context of irreversible covalent inhibition. One of the challenges in design of drugs that engage their targets irreversibly is that it’s not possible to meaningfully quantify activity with a single parameter. This is particularly relevant to definition of efficiency metrics which are typically derived by either scaling or offsetting a measured activity value by a risk factor such as molecular size or lipophilicity. While you can certainly measure an IC50 value for an irreversible covalent inhibitor the value that you measure will be time-dependent and it’s not generally meaningful to compare two IC50 values that have been measured using different incubation times. While the kinact/KI ratio is time-independent using it as a measure of activity necessarily entails a degree of information loss.

The authors of FK2025 state:

Our starting point is the LE introduced for noncovalent ligands as a useful metric for lead selection. (p5) [LE  was claimed to be useful when it was introduced although no evidence was presented in support of the claim.]

Let’s take a look at Table 2 which shows two equations from FK2025. The first equation, which appears in the text of the article, illustrates two common errors in the efficiency metric field (taking logarithms of dimensioned quantities and discarding units). It should ring alarm bells for the reader when authors make either error (especially if the authors interpret values of the efficiency metrics).


The authors of FK2025 assert that “LE can be decomposed into contributions from the noncovalent recognition and the covalent reaction (Box 2, Equation III)” and this is reproduced in Table 2 as Equation 2. The first term is a commonly-used mathematical formula for LE when inhibition is reversible and it is important to be aware that KI has been divided by an arbitrary concentration value (1 M) in order that it can be expressed as a logarithm (see Equation 1 in Table 1). The argument of the logarithm in the second term is dimensionless although its magnitude does vary with t. Each term in Equation III (Box 2) has a nontrivial dependence on the value of an arbitrary quantity (the 1 M concentration in the first term and t, in the second term). This means that your perception of efficiency when calculated according to Equation III (Box 2) will be altered if you use either a different concentration unit or a different value of t. You can see this effect in Figure 2 (effect of varying of t) and the appearance of Figure 1 will be altered if you use a value of t other than 1 h or a concentration unit other than M for the calculation of LE.

It's now time to examine CLE (defined as Equation II in Box 3 of the FK2025 study) and I’ll direct you to Table 3 below in which I’ve made some comments. Using CLE requires that the IC50 values for the inhibitors of interest all correspond to the same time point (t) and it is not clear whether the authors are suggesting that that the IC50 values should all be measured using the same incubation time or need to be calculated from measured KI and kinact values using Equation III in Box 2. A quantity t is also explicitly present in the argument of the logarithm in Equation II in Box 3 and this is necessary for the argument of the logarithm to be dimensionless (see M2011). The argument of the logarithm in Equation II (Box 3) is clearly time-dependent and this means that your perception of efficiency will be altered if you use a different value of t when calculating CLE (just as your perception of efficiency will be altered if you use a different concentration unit to express IC50 when you calculate LE for reversible inhibitors). It also means the molecular size dependency of CLE will vary with time just as the molecular size dependency of LE varies with the concentration unit used to express affinity as can be seen in Fig. 1 of (p9).

However, there is another difficulty which is that the argument of the logarithm in Equation II (Box 3) is not a valid measure of activity (the same criticism can also be made of the xLE metric introduced in the Z2025 study that Dan has already reviewed).  This problem is a bit more subtle and it’s important to remember that knowing the IC50 value for a reversible inhibitor enables you to generate a concentration response for inhibition.  When you express an IC50 value as a logarithm you need to scale it by a concentration value to ensure that the argument of the logarithm function is dimensionless (see M2011) but it’s important to remember that the concentration unit is still there even though it’s not shown (see Equation 1 in Table 1).

This is a good point at which to wrap up and I’ve argued that CLE has two deficiencies. First, perception of efficiency and its dependency on molecular size both vary with an arbitrary quantity (t) in the argument of the logarithm (this is analogous to the problems caused by the arbitrary nature of the concentration unit used for scaling affinity/potency in the definition of LE for reversible binders). Second, the argument of the logarithm is not a valid measure of activity because it cannot be used to generate a concentration response. Furthermore, I would question the value of aggregating results from multiple assays for analysis even for a valid metric without these deficiencies and I offered the following advice in (p9):   

Drug designers should not automatically assume that conclusions drawn from analysis of large, structurally-diverse data sets are necessarily relevant to the specific drug design projects on which they are working.

I’ve criticized the FK2025 study at length and saying how I might use data like this in drug design projects is a good way to conclude the post.  A general criticism that I have made of drug design efficiency metrics is that they are based on assumptions of relationships between activity and risk factors such as molecular size.  I argued in (p9) that one should use the trend that is actually observed in the data to normalize activity with respect to risk factors and I’ll point you to the relevant section (Alternatives to ligand efficiency for normalization of affinity) in that article. I would start by attempting to model the relationship between kinact and reactivity with glutathione. The objective of this exercise is to identify inhibitors that best exploit their intrinsic reactivity when forming covalent bonds with the target residue (you can quantify this by how far the point for an inhibitor lies above the trend line and the most interesting compounds have the largest positive residuals). I might  also examine the relationship between kinact and KI for inhibitors with the same intrinsic reactivity (e.g., incorporating the same warhead) with a view to identifying the inhibitors for which non-covalent interactions with the target most effectively stabilise the transition state relative to the non-covalent complex. I should stress that there is no suggestion that these analyses would necessarily yield useful insight.

It's been a long post so thanks for staying with me. This will be the the last post until after Christmas and, as I extend best wishes to all for a happy and peaceful festive season, I'm keenly aware that Christmas will be neither happy nor peaceful for many of our fellow human beings.  

Thursday, 12 October 2017

The resurrection of Maxwell's Demon

Sometimes when reading the residence time literature, I get the impression that the off-raters have re-animated Maxwell's Demon. It seems as if a nano-doorman stands guard at at the entrance of the binding site, only opening his nano-door to ligand molecules that want to get in. Microscopic Reversibility? Stop being so negative! With Big Data, Artificial Intelligence, Machine Learning (formerly known as QSAR) and Ligand Efficiency Metrics we can beat Microscopic Reversibility and consign The Second Law to the Dustbin Of History!

There were a number of things that triggered this blog post. First, I saw a recent article that got me thinking about philatelic drug discovery.  Second, some of the off-raters will be getting together in Berlin next week and I wanted to share some musings because I won't be there in person. Third, my former colleague Rutger Folmer has published a useful (dare I say, brave) critique of the residence time concept that is bang on target. 

I'm not actually going to say much about Rutger's article except to suggest that you read it. That's because I really want to examine the article on philatelic drug discovery in a more detail (it's actually about thermodynamic and kinetic profiling but I thought the reference to philately would better grab your attention). My standard opening move when playing chess with an off-rater is to assert that slow binding is equivalent to slow distribution. In what situations would you design a drug to distribute slowly?

Chemical kinetics is all about energy barriers and, the higher the barrier, the slower things will happen. Microscopic reversibility tells us that a barrier to association is a barrier to dissociation and that the ligand will return to solution along the same path that it took to its binding site. Microscopic reversibility tells you that if you got into the parking spot you can get out of it as well although that may not be the experience of every driver. The reason that microscopic reversibility doesn't always seem to apply to parking is that most humans, with the possible exception of tank drivers in the Italian army, are more comfortable in forward gear than in reverse. Molecules, in contrast, have no more concept of forward and reverse than they do of standard states, IUPAC or the opinions the 'experts' who might quantitatively estimate their drug-likeness while judging their beauty. Molecules don't actually do concepts. Put more uncouthly, molecules just don't give a toss.

I've created a graphic to illustrate to show how things might look in vivo when there is a barrier to association (and, therefore, to dissociation). We can think of the ligand molecule having to get over the barrier in order to get to its binding site and we call the top of the barrier the 'transition state'. This is a simplified version of reality (it is actually the system that passes from the unbound state through the transition state to the bound state and for some ligand-protein association there is no barrier) but it'll serve for what I'd like to say. The graphic consists of three panels and the first (A) of these illustrates the situation soon after dosing when the concentration of ligand (L) is relatively high and the target protein (P) has not had sufficient time to respond. If the barrier is sufficiently high, the system can't get to equilibrium before the ligand concentration starts to fall in what a pharmacokineticist might refer to as the elimination phase. Under this scenario the system will be at equilibrium briefly as the ligand concentration falls and I've shown this in panel B. After the equilibrium point is reached, the rate of dissociation exceeds the rate of association and this is shown in panel C. 



There's something else that I'd like you to take a look at in the graphic and that's the free energy (G) of the unbound state (P + L).  See how it goes down relative to the free energy of the bound state (P.L) as the concentration of ligand decreases. When thinking about energetics of these systems, it actually makes a lot of sense to use the unbound state as the reference but you do need to use a reference concentration (e.g. 1 M) to to do this.

When we do molecular design we often think in terms of manipulating energy differences. For example, we try to increase affinity by stabilizing the bound state relative to the unbound state. Once you start trying to manipulate off-rates, you soon realize that you can't change one thing at a time (unless you draft Maxwell's Demon into your project team).  I've created a second graphic which looks similar to the first graphic although there are important differences between the two graphics. In particular, I'm referencing energy to the unbound state (P + L) which means that the ligand concentration is constant in all three panels. Let's consider the central panel as the starting point for design. We can go left from that starting point and stabilize the bound state which is equivalent to optimizing affinity.  Stabilizing the bound state will also result in slower dissociation provided that the transition stare energy remains unchanged. This is a good thing but it's difficult to show that the benefits come from the slower dissociation and not from the increased affinity. If you raise the barrier (i.e. increase the energy of the transition state) to reduce the off-rate you'll find that you have slowed the on-rate to an equal extent.        



Before moving on, it may be useful to sum up where we've got to so far. First, ask yourself why you think off-rates will be relevant in situations where concentration changes on a longer time scale than binding. Second, you'll need to enlist the help of Maxwell's Demon if you want to reduce off-rate without affecting on-rate and/or affinity. Third, if you want to consider binding kinetics in design then it'd be best to use barrier height (referenced to unbound state) and affinity as your design parameters.

Now I'd take a look at the philatelic drug discovery article. This is a harsh term but it does capture a tendency in some drug discovery programs to measure things for the sake of it (or at least to keep the grinning Lean Six Sigma 'belts' grinning).  Some of this is a result of using techniques such as isothermal titration calorimetry (ITC) and surface plasmon resonance (SPR) that yield information in addition to affinity (that is of primary interest) at no extra cost. I really don't want to come across as a Luddite and I must stress that measurements of enthalpy, entropy, on-rate and off-rate are of considerable scientific interest and are also valuable for improving physical models. Furthermore, I am continually awed by the exquisite sensitivity of modern ITC and SPR instruments and would always want the option to be able to measure affinity using at least one of these techniques. However, problems start when the access to enthalpy, entropy, off-rates and on-rates becomes exploited for 'metrication' and drug discovery scientists seek 'enthalpy-driven' binding simply because the binding will be more 'enthalpy-driven'. It is easier to make the case for relevance of binding kinetics although, as Rutger points out, reducing the off-rate may very well make things worse if the on-rate is also reduced. It is much more difficult to assemble a coherent case for the relevance of thermodynamic signatures in drug discovery. Perhaps, some day, a seminal paper from the Budapest Enthalpomics Group (BEG) will reveal that isothermal systems like live humans can indeed sense the enthalpy and entropy changes associated with drugs binding to their targets although I will not be holding my breath.

Unsurprisingly, the thermodynamic and kinetic profiling (aka philatelic drug discovery) article advocates thermodyanamic profiling of bioactive compounds in lead optimization projects. I'm going to focus on the kinetic profiling and it is worrying that the authors don't seem to be aware that on-rates and off-rates have to be seen in a pharmacokinetic context in order to make the connection with drug discovery. The authors may find it instructive to think about how inhibitor concentration would have varied over the course of a typical experiment in their cell-based assays. They are also likely to find Rutger's article to be educational and I recommend that they familiarize themselves with its content.

The following statement suggests that it may be beneficial for the authors to also familiarize themselves with the rudiments of chemical kinetics:


"Association and dissociation rate constants (kon and koff) of compound binding to a biological target are not intrinsically related to one another, although they are connected by dissociation equilibrium constant KD (KD = koff/kon)."

The processes of association and dissociation are actually connected by virtue of taking place along the same path and by having to pass through the same transition states. The difference in barrier heights for association and dissociation is given by the binding free energy. 

Some analysis of relationships between potency in a cell-based assay and  KD, koff and kon were presented in Figure 6 of the article. I have a number of gripes with the analysis. First, it would be better to use logarithms of quantities like KD, IC50, koff and kon when performing analysis of this nature. In part, this because we typically look for linear free energy relationships in these situations. There is another strong rationale for using logarithms because analysis of correlations between continuous variables works best when the uncertainties in data values are as constant as possible. My second gripe is that the authors have chosen to bin their data for analysis and this is a great way to shoot yourself in the foot. When you bin continuous data you both reduce your data analysis options and leave people wondering whether the binning has been done to hide the weakness of the trends in the data.   I have droned at length about why it is naughty to bin continuous data so I'll leave it at that.

It's been a long post and it's time to wrap things up. If you've found the post to be 'cansativo' (sounds so much more soothing in Portguese) then spare a thought for the person who had to write it. To conclude, I'll leave you with a quote that I've taken from the abstract for Rutger's article:
  
"Moreover, fast association is typically more desirable than slow, and advantages of long residence time, notably a potential disconnect between pharmacodynamics (PD) and pharmacokinetics (PK), would be partially or completely offset by slow on-rate."