Showing posts with label units. Show all posts
Showing posts with label units. Show all posts

Sunday, 2 August 2020

Why fragments?


Paramin panorama

Crystallographic fragment screens have been run recently against the main protease (at Diamond) and the Nsp3 macrodomain (at UCSF and Diamond) of SARS-Cov-2 and I thought that it might be of interest to take a closer look at why we screen fragments. Fragment-based lead discovery (FBLD) actually has origins in both crystallography [V1992 | A1996] and computational chemistry [M1991 | B1992 | E1994]. Measurement of affinity is important in fragment-to-lead work because it allows fragment-based structure-activity relationships to be established prior to structural elaboration. Affinity measurement is typically challenging when fragment binding has been detected using crystallography although affinity can be estimated by observation of the response of occupancy to concentration (the ∆G° value of −3.1 kcal/mol reported for binding of pyrazole to protein kinase B was derived in this manner).

Although fragment-based approaches to lead discovery are widely used, it is less clear why fragment-based lead discovery works as well as it appears to. While it has been stated that “fragment hits form high-quality interactions with the target”, the concept of interaction quality is not sufficiently well-defined to be useful in design. I ran a poll which asked about the strongest rationale for screening fragments.  The 65 votes were distributed as follows: ‘high ligand efficiency’ (23.1%), ‘enthalpy-driven binding’ (16.9%), ‘low molecular complexity’ (26.2%) and ‘God loves fragments’ (33.8%). I did not vote.

The belief is that fragments are especially ligand-efficient has many adherents in the drug discovery field and it has been asserted that “fragment hits typically possess high ‘ligand efficiency’ (binding affinity per heavy atom) and so are highly suitable for optimization into clinical candidates with good drug-like properties”. The fundamental problem with ligand efficiency (LE), as conventionally calculated, is that perception of efficiency varies with the arbitrary concentration unit in which affinity is expressed (have you ever wondered why Kd , Ki or IC50 has to be expressed in mole/litre for calculation of LE?). This would appear to be an rather undesirable characteristic for a design metric and LE evangelists might consider trying to explain why it’s not a problem rather than dismissing it as a “limitation” of the metric or trying to shift the burden of proof is onto the skeptics to show that the evangelists’ choice of concentration unit for calculation of LE is not useful.

The problems associated with the arbitrary nature of the concentration unit used to express affinity were first identified in 2009 and further discussed in 2014 and 2019. Specifically, it was noted that LE has a nontrivial dependency on the concentration,  C°, used to define the standard state. If you want to do solution thermodynamics with concentrations defined then you do need to specify a standard concentration. However, it is important to remember that the choice of standard concentration is necessarily arbitrary if the thermodynamic analysis is to be valid. If your conclusions change when you use a different definition of the standard state then you’ll no longer be doing thermodynamics and, as Pauli might have observed, you’ll not even be wrong. You probably don't know it, but when you use the LE metric, you’re making the sweeping assumption that all values of Kd, Ki and IC50 tend to a value of 1 M in the limit of zero molecular size. Recalling the conventional criticism of homeopathy, is there really a difference between a solute that is infinitely small and a solute that is infinitely dilute?

I think that’s enough flogging of inanimate equines for one blog post so let’s take a look at enthalpy-driven binding. My view of thermodynamic signature characterization in drug discovery is that it’s, in essence, a solution that’s desperately seeking a problem. In particular, there does not appear to be any physical basis for claims that the thermodynamic signature is a measure of interaction quality.  In case you’re thinking that I’m an unrepentant Luddite, I will concede that thermodynamic signatures could prove useful for validating physics-based models of molecular recognition and in, in specific cases, they may point to differences in binding mode within congeneric series. I should also stress that the modern isothermal calorimeter is an engineering marvel and I'd always want this option for label-free, affinity measurement in any project.

It is common to see statements in the thermodynamic signature literature to the effect that binding is ‘enthalpy-driven’ or ‘entropy-driven’ although it was noted in 2009 (coincidentally, in the same article that highlighted the nontrivial dependence of LE on C°) that these terms are not particularly meaningful. The problems start when you make comparisons between the numerical values of ∆H (which is independent of C°) and T∆S° (which depends on C°). If I’d presented such a comparison in physics class at high school (I was taught by the Holy Ghost Fathers in Port of Spain), I would have been caned with a ferocity reserved for those who’d dozed off in catechism class.  I’ll point you toward an article which asserts that, “when compared with many traditional druglike compounds, fragments bind more enthalpically to their protein targets”. I have a number of issues with this article although this is not the place for a comprehensive review (although I’ll probably pick it up in ‘The Nature of Lipophilic Efficiency’ when that gets written).

While I don’t believe that the authors have actually demonstrated that fragments bind more enthalpically than ligands of greater molecular size, I wouldn’t be surprised to discover that gains in affinity over the course of a fragment-to-lead (F2L) campaign had come more from entropy than enthalpy. First, the lost translation entropy (the component of ∆S° that endows it with its dependence on C°) is shared over greater number of intermolecular contacts for structurally-elaborated compounds and this article is relevant to the discussion. Second, I’d expect the entropy of any water molecule to increase when it is moved to bulk solvent from contact with molecular surface of ligand or target (regardless of polarity of the molecular surface at the point of contact). Nevertheless, this is something that you can test easily by examining the response of (∆H + T∆S°) to ∆G° (best to not to aggregate data for different targets and/or temperatures when analyzing isothermal titration calorimetry data in this manner). But even if F2L affinity gains were shown generally to come more from entropy than enthalpy, would that be a strong rationale for screening fragments?

This gets us onto molecular complexity and this article by Mike Hann and GSK colleagues should be considered essential reading for anybody thinking about selecting of compounds for screening. The Hann model is a conceptual framework for molecular complexity but it doesn’t provide much practical guidance as to how to measure complexity (this is not a criticism since the thought process should be more about frameworks and less about metrics). I don’t believe that it will prove possible to quantify molecular complexity in an objective manner that is useful for designing compound libraries (I will be delighted to be proven wrong on this point). The approach to handling molecular complexity that I’ve used in screening library design is to restrict extent of substitution (and other substructural features that can be considered to be associated with molecular complexity) and this is closer to ‘needle screening’ as described by Roche scientists in 2000 than to the Hann model.

Had I voted in the poll, ‘low molecular complexity’ would have got my vote.  Here’s what I said in NoLE (it’s got an entire section on fragment-based design and a practical suggestion for redefining ligand efficiency so that perception does not change with C°):

"I would argue that the rationale for screening fragments against targets of interest is actually based on two conjectures. First, chemical space can be covered most effectively by fragments because compounds of low molecular complexity [18, 21, 22] allow TIP [target interaction potential] to be explored [70,71,72,73,74] more efficiently and accurately. Second, a fragment that has been observed to bind to a target may be a better starting point for design than a higher affinity ligand whose greater molecular complexity prevents it from presenting molecular recognition elements to the target in an optimal manner."

To be fair, those who advocate the use of LE and thermodynamic signatures in fragment-based design do not deny the importance of molecular complexity. Let’s assume for the sake of argument that interaction quality can actually be defined and is quantified by the LE value and/or the thermodynamic signature for binding of compound to target. While these are massive assumptions, LE values and thermodynamic signatures are still effects rather than causes.

The last option for poll was ‘God loves fragments’ and more respondents (33.8%) voted for this than any of the first three options. I would interpret a vote for ‘God loves fragments’ in three ways. First, the respondent doesn’t consider any one of the first three options to be a stronger rationale for screening fragments than the other two. Second, the respondent doesn’t consider any of the first three options to be a valid rationale for screening fragments. Third, the respondent considers fragment-based approaches to have been over-sold.

This is a good place to wrap up. While I remain an enthusiast for fragment-based approaches to lead discovery, I do also believe that they have been somewhat oversold. The sensitivity of LE evangelists to criticism of their metric may stem from the use of LE to sell fragment-based methods to venture capitalists and, internally, to skeptical management. A shared (and serious) deficiency in the conventional ways in which LE and thermodynamic signature are quantified is that perception changes when the arbitrary concentration,  C°, that defines the standard state is changed. While there are ways in which this deficiency can be addressed for analysis, it is important that the deficiency be acknowledged if we are to move forward. Drug design is difficult and if we, as drug designers, embrace shaky science and flawed data analysis then those who fund our activities may conclude that the difficulties that we face are of our own making.     

Saturday, 1 April 2017

A concentration of scoring functions

<< previous || next >>

Researchers at The Hungarian Institute Of Thermodynamics have published a number of seminal articles on the interplay of enthalpy and entropy in areas ranging from physical chemistry to socioeconomics. For example, the cause of World War 1 (also known as 'The Great War' although I doubt whether any of its participants thought that it was that great) was traced to a singularity in the Habsburg Partition Function. In a nutshell, the problem was shown to be a surfeit of the wrong type of entropy (which led to Franz Ferdinand's driver getting lost) coupled with a deficit in the right type of entropy (which would have prevented Gavrilo Princip's bullets from finding their targets). However, it is unlikely that any amount of the right type of entropy could have saved the hapless Maximilian I of Mexico, who generously volunteered to be Emperor only to be shot by the ungrateful Mexicans.

The most recent study from BEG (Budapest Enthalpomics Group) is little short of sensational. Unfortunately it's not available online and the poor fax quality, coupled with my rudimentary grasp of Hungarian, have made the going hard. The essence of this seminal study is that the performance of scoring functions can be significantly improved by including the concentration unit (in which affinity is expressed) as a parameter in the fitting process. The casual observer of virtual screening may have wondered why scoring functions are trained with affinity but validated by enrichment. By treating the concentration unit as a parameter in the fitting process, the authors were able to achieve unprecedented accuracy of prediction and the phone call from Stockholm would seem to be a foregone conclusion. Commenting on these seminal findings, Prof. Kígyó Olaj, the director of the Institute said, "Now we no longer need to use ROC plots to mask feeble correlations between predicted and measured affinity".     

Friday, 6 January 2017

Confessions of a Units Nazi

Regular readers (both of them) of this blog will know that I have an interest, which some might term an obsession, with units. At high school in Trinidad, we had the importance of units beaten into us by the Holy Ghost Fathers and, for some of the more refractory cases, the beating was quite literal. I was taught physics by the much loved, although somewhat highly-strung, Fr. Knolly Knox (aka Knox By Night) who, as Dean of the First Form, used to give 'licks' with a cane of hibiscus (presumably chosen for its tensile properties). You quickly learned not to mess with The Holy Ghost Fathers, especially the Principal, Fr. Arthur Lai Fook (aka Jap), and it was a brave student who responded to the request by Fr. Pedro Valdez to define the dyne by answering, "Fah, it what happen after living". Fr. Pedro was a gentle soul although his brother, Fr. Toba, who taught me Latin, would lob a blackboard eraser with reproducible inaccuracy at any student who had the temerity to doze off during the Second Punic War while Hannibal and his elephants were steamrollering the hapless legions of Gaius Flaminius into Lake Trasimene. At least we didn't have detention at my school. Actually we did have detention only it was called 'penance'. Each and every student also had a Judgement Book in which was entered a mark (out of 10) for each subject each and every week. A mark of 5 (or less) or a failure to return one's Judgement Book, duly signed by parent or guardian, by Wednesday morning earned the transgressor a corrective package of Licks and Penance.  As a thoughtful child, I managed to shield my parents from this irksome bureaucracy and, in any case, it was simply safer that The Holy Ghost Fathers were never given the opportunity to familiarize themselves with the authentic parental signatures.


I used to think that 'Virtus et Scientia' was Latin for 'Licks and Penance'  (17-Feb-2018 update)


What we learned from the Holy Ghost Fathers was that most physical quantities have dimensions and if the quantities on the opposite sides of the 'equal sign' in an equation have different dimensions then it is a sign of an unforced error rather than a penetrating insight. For example the dimensions of force are MLT-2 (M = mass; L = length; T = time) and you are free to express forces in newtons, dynes or poundals as you prefer. You can think of a physical quantity as a number multiplied by a unit and, without the unit, the number is meaningless. Units are extremely important but at the same time they are arbitrary in the sense that if your physical insight changes when you change a unit then it is neither physical nor an insight. Here's a good illustration of why dimensional analysis matters.

I have blogged ( 12 | 3 ) about how building the a concentration unit into the definition of ligand efficiency (LE) results in a metric that is physically meaningless (even though it remains a useful instrument of propaganda) and, for the masochists among you, there's also the LE metric critique in JCAMD. The problem can be linked to a lack of recognition of the fact that logarithms can only be calculated for numbers (which lack units). However, LE has another 'units issue' which is connected with the fact that it is a molar energy that is scaled in the definition of LE rather than pIC50 or pKd. This needn't be an issue but, unfortunately, it is. LE is defined by dividing a molar energy by the number of non-hydrogen atoms in the molecular structure and there is nothing in the definition of LE that says that the energy has to be expressed in any particular unit. This means that you can define LE using any energy unit that you want to. Some 'experts' appear to believe that dividing a molar energy by number of non-hydrogen atoms relieves them of the responsibility to report units. I'm referring, of course, to the practise of multiplying pIC50 or pKd by 1.37 when calculating LE. You might ask why people do this, especially given that 'experts' tout the simplicity of LE and they don't multiply pIC50 or pKd by 1.37 when they calculate LipE/LLE. Don't ask me because I'm neither expert nor 'expert'.

Let's take a look at this NRDD article on LE metrics and I'd like you to go straight to Box 1 (Ligand efficiency metrics). Six numbered equations are shown in Box 1 and it is stated towards the end of the first paragraph that "each equation corresponds to a mathematically valid function".  This statement is incorrect because the first equation (1) in Box 1 is not a mathematically valid function. The reason for this is that the logarithm function cannot take as its argument a quantity, such as Kd, that has units. Equation (5), which defines LLEAT, is mathematically valid although it differs from the mathematically ambiguous equation that was originally used to define LLEAT

To be honest, I think that Box 1 is probably beyond repair by conventional erratum and I'll back this opinion with an example:


"Assuming standard conditions of aqueous solution at 300K, neutral pH and remaining concentrations of 1M,
 –2.303RTlog(Kd/C°) approximates to –1.37 × log(Kd) kcal/mol." 

At my school in Trinidad this would have been called a 'ratch' and, once detected, it would have earned its perpetrator a corrective package of Licks and Penance. I don't think even the Holy Ghost Fathers could have exorcised a concentration unit quite this efficiently. 

In some physical chemistry literature, Kd is defined as a dimensionless quantity by including C° in the definition of Kd. However, in the literature of biochemistry, biophysics and medicinal chemistry,  Kd  is usually quoted in units of concentration. Binding free energy has the same value and same dependence on C° regardless of  which of the two conventions is used to define Kd 
(Update 17-Feb-2018) 

I'd now like to talk a bit about the 'p' operator that we use to transform IC50 and Kd values into logarithms. This makes it much easier to perceive structure-activity relationships and provides a better representation of measurement precision than when the IC50 and Kd values themselves are used. To calculate pKd,, first express Kd in molar concentration units, dump the units and calculate minus the logarithm of the number. I realize that this may come across as arm waving but the process of converting  Kd, to  pKd, can actually be expressed exactly in mathematical terms as follows:

 pKd = –log10(Kd/M)

The 'p' operator has a 1 M concentration built into it. Although this choice of unit is arbitrary, it doesn't cause any problems if you're doing sensible things (e.g. subtracting them from each other) with the pKd values. If, however, you're doing silly things (e.g. dividing them by numbers of non-hydrogen atoms) with the pKd values then the plot starts to unravel faster than you can say 'Brexit means Brexit'. 

I'd like you take a look at another article which also has a Box 1 although I won't bother you with another tiresome 'spot the errors' quiz. The equation that I'll focus on is:

pKd = pKH + pKS 

This equation describes the decomposition of affinity into enthalpic and entropic contributions and you might think this means that you can write:

Kd = KH × KS 

As Prof. Pauli would have observed, this is an error in the 'not even wrong' category and it is clear that a difference in opinion as to the importance of units was as much responsible for the unraveling of the Austro-Hungarian empire as that unfortunate wrong turn in pre-SatNav Sarajevo. The 'p' operator implies that each of KdKH and Khas units of concentration. However, multiplying two such quantities will give a quantity that has units of concentration squared. 

It is actually possible to decompose Kd into enthalpic and entropic contributions a valid manner but you need to be thinking carefully about the meaning of the standard state. As noted previously DG° depends on the concentration used to define the standard state. This is a consequence of the dependence of DS° on the standard concentration and DH is independent of the standard concentration (the standard state is assumed to be a dilute solution). This suggests defining KS as quantity with units of concentration and Kas a quantity without units.

This is probably a good point to wrap things up. My advice to all the authors of the featured NRDD and FMC articles is that they read (and make sure that they understand) the section of this article that is entitled '8. Ligand Efficiency and Additivity Analysis of Binding Free Energy'. This advice is especially relevant for those of the authors who consider themselves to be experts in thermodyamics.

May I wish all readers a happy, successful and metric-free 2017.

Saturday, 30 August 2014

Ligand efficiency metrics considered harmful

Next >>
It has been a while since I did a proper blog post. Some of you may have encountered a Perspective article entitled, ‘Ligand efficiency metrics considered harmful’ and I’ll post on it because the journal has made the article open access until Sept 14.  The Perspective has already been reviewed by Practical Fragments and highlighted in a F1000 review. Some of the points discussed in the article were actually raised last year in Efficient Voodoo Thermodynamics, Wrong Kind of Free Energy and ClogPalk : a method for predicting alkane/water partition coefficients.  There has been recent debate about the validity of ligand efficiency (LE) which is summarized in a blog post (make sure to look at the comments as well).  However, I believe that both sides missed the essential point which is that the choice (conventionally 1 M) of concentration that is used to define the standard state is entirely arbitrary.  

In this blog post, I’ll focus on what I call ‘scaled’ ligand efficiency metrics (LEMs). Scaling means that a measure of activity or affinity is divided by a risk factor such as molecular weight (MW) or heavy (i.e. non-hydrogen) atoms (HA).   For example, LE can be calculated by dividing the standard free energy of binding (ΔG°) by HA:
LE  =  (1/HA)´RTloge(Kd/C°)

Now you’ll notice that I’ve written ΔG° in terms of the dissociation constant (Kd) and the standard concentration (C°) and I articulated why this is important early last year.  The logarithm function is only defined for numbers (i.e. dimensionless quantities) and, inconveniently, Kd has units of concentration.  This means that LE is a function of both Kd and C° and I’m going to first redefine LE a bit to make the problem a bit easier to see.   I'll use IC50 instead of Kd to define a new LEM which I’m not going to name (in the LEM literature names and definitions keep changing so much better to simply state the relevant formula for whatever is actually used):

(−1/NHA)´log10(IC50/Cref)

I have a number of reasons for defining the metric in this manner and the most important of these is that the new metric is metric is dimensionless.  Note how I use the number of heavy atoms (NHA) rather than HA to define the metric.  Just in case you’re wondering what the difference is, HA for ethanol is 3 heavy atoms while NHA is 3 (lacking units of heavy atoms). [In the original post I'd counted 2 for this molecule but if you check the comments you'll see that this howler was picked up by an alert reader and I has now been corrected.]  Apologies for being pedantic (some might even call me a units Nazi) but if people had paid more attention to units, we’d never have got into this sorry mess in the first place.  The other point to note is that I’ve not converted IC50 to units of energy, mainly for the reason that it is incorrect to do so because an IC50 is not a thermodynamic quantity. However, there are other reasons for not introducing units of energy. Often units of energy go AWOL when values of LE are presented and there is no way of knowing whether the original units were kcal/mol or kJ/mol. Even when energy units are stated explicitly, this might be masking a situation in which affinity and potency measurements have been combined in a single analysis.  Of course, one can be cynical and suggest that the main reason for introducing energy units is to make biochemical measurements appear to be more physical. 

So let’s get back to that new metric and you’ll have noticed a quantity in the defining equation that I’ve called Cref (reference concentration).  This is similar to the standard concentration in that the choice of its value is completely arbitrary but it is also different because it has no thermodynamic significance. You need to use it when defining the new LEM because, at the risk of appearing repetitive, you can’t calculate a logarithm for a quantity that has units.  Another way of thinking about Cref is as an arbitrary unit of concentration that we’re using to analyze some potency measurements.  Something that is really, really important and fundamental in science is that your perception of a system should not change when you change the units of the quantities that describe the system.  If it does then, in the words of Pauli, you are “not even wrong” and you should avoid claiming penetrating insight so as to prevent embarrassment later. So let’s take a look at how changing Cref affects our perception of ligand efficiency.  The table below is essentially the same as what was given in the Perspective article (the only difference is that I’m using IC50 in the blog post rather than Kd).  I also should point out that Huan-Xiang Zhou and Mike Gilson made a similar criticism of LE back in 2009 (although we cited their article in the context of standard states, we failed to notice the critique at the end of their article and there really is no excuse for having missed it).  When a reference concentration of 1 M is used, the three compounds are all equally ligand efficient according to the new LEM.  If Cref = 0.1 M, the compounds appear to become more ligand efficient as molecular size increases but the opposite behavior is observed for Cref = 10 M.


Here’s a figure that shows the problem from a different angle. See how the three parallel lines respond the scaling transformation (Y=> Y/X).  The line that passes through the origin transforms to a line of zero slope (Y/X is independent of X) while the other two lines transform to curves (Y/X is dependent on X).  This graphic has important implications for fit quality (FQ) because it shows that some of the size dependency of LE is a consequence of the (arbitrary) choice of a value (usually 1 M) for Cref or C°.


A common response that I’ve encountered when raising these points is that LE is still useful.  My counter-response is that Religion is also useful (e.g. as a means for pastors to fleece their flocks efficiently) but that, by itself, neither makes it correct nor ensures that that it is being used correctly (if that is even possible for Religion).  One occasionally expressed opinion is that, provided you use a single value for Cref or C°, the resulting LE values will be consistent.  However, consistency is no guarantee of correctness and we need to remember that we create metrics for the purpose of measuring things with them.  When you advocate use of a metric in drug discovery, the burden of proof is on you to demonstrate that the metric has a sound scientific basis and actually measures what it supposed to measure.
The Perspective is critical of the LEMs used in drug discovery but it does suggest alternatives that do not suffer from the same deficiencies.  This is a good point to admit that it took me a while to figure out what was wrong with LE and I’ll point you towards a blog post from over five years ago that will give you an idea about how my position on LEMs has evolved.  It is often stated that LEMs normalize activity with respect to risk factor although it is rarely, if ever, stated explicitly what is meant by the term ‘normalize’.   One way of thinking about normalization is as a way to account for the contribution of a risk factor, such as molecular size or lipophilicity, to activity.  You can do this by modelling the activity as a function of risk factor and using the residual as a measure of activity that has been corrected for the contribution made by the risk factor in question.  You can also think of an LEM as a measure of the extent to which the activity of a compound beats a trend (e.g. linear response of activity to HA). If you’re going to do this then why not use the trend actually observed in the data rather than some arbitrarily assumed trend?

Although I still prefer to use residuals to quantify extent to which activity beats a trend, the process of modelling activity measurements hints at a way that LE might be rehabilitated to some extent.  When you fit a line to activity (or affinity) measurements, you are effectively determining the value of Cref (or C°) that will make the line of fit pass through the origin and which you can then use to redefine activity (or affinity) for the purpose of calculating LE. I would argue that the residual is a better measure of the extent to which the activity of a compound beats the trend in the data because the residual has sign and the uncertainty in it does not depend explicitly on the value of a scaling variable such as HA. However, LE defined in this manner can at least be claimed to take account of the observed trend in the activity data. Something that you might want to think about in this context is whether or not you'd expect the same parameters (slope and intercept) if you were to fit activity measured against different targets to MW or HA.  My reason for bring this up is that it has implications for the validity of mixing results from different assays in LE-based analyses so see what you think.  


I’ll wrap up by directing you to a presentation that I've been doing lately and it includes material from the earlier correlation inflation study.  A point that I make when presenting this material is that, if we do bad science and bad data analysis, people can be forgiven for thinking that the difficulties in drug discovery may actually be of our own making.