Showing posts with label voodoo thermodynamics. Show all posts
Showing posts with label voodoo thermodynamics. Show all posts

Monday, 1 April 2024

Standard states and solution thermodynamics

<< previous || next >>

Readers of this blog know that, on more than one occasion, I have denounced the ligand efficiency metric as physically meaningless on the grounds that perception of efficiency varies with the concentration value that defines the standard state. As I argue in NoLE this is clearly thermodynamic nonsense (Pauli might even have suggested that it wasn’t even wrong) and the equivalent cheminformatic argument is that perception shouldn’t change when you use a different unit to express a quantity.

A change in perception resulting from using a different standard concentration can also be a problem when analysing thermodynamic signatures. One particular absurdity is that binding can be switched from enthalpy-driven to entropy-driven simply by using a different concentration to define the standard state. This statement in the W2014 article unintentionally highlights the issue:

Consequently, we define the dimensionless ratio (ΔH + TΔS)/ΔG as the Enthalpy–Entropy Index (IE–E) and use it here to indicate the enthalpy content of binding. Its advantageous feature is that it is normalised by the free energy ΔG (= ΔH  – TΔS), and so it can be used to compare compounds with millimolar to nanomolar binding affinities during the course of a hit-to-lead optimisation.

I do indeed think that it makes a lot of sense to use (ΔH + TΔS) and ΔG as parameters for exploring thermodynamic signatures. However, the dimensionless ratio of the two quantities is physically meaningless because of its dependence on the concentration used to define the standard state (this dependence stems from the fact that ΔS depends on the standard concentration while ΔH is invariant to change in the standard concentration).

One article that I’ve been particularly critical of in the past is “The role of ligand efficiency metrics in drug discovery” NRDD 133:105-121 (2014) DOI. Specifically, I have expressed concerns about this sentence in Box 1 (Ligand efficiency metrics) of the article:

Assuming standard conditions of aqueous solution at 300K, neutral pH and remaining concentrations of 1M, –2.303RTlog(Kd/C°) approximates to –1.37 × log(Kd) kcal/mol.

I do need to mention a potential source of confusion when analysing Kd values. In biochemistry, biophysics and drug discovery Kd values are conventionally quoted as dimensioned quantities in units of concentration. However, Kd values may also be quoted as dimensionless ratios and, in these cases, the Kd value depends on the concentration used to define the standard state. There seems to be an error in that the approximation appears to eliminate the dimensions of the standard concentration C°.

I should say that I’ve always been a bit nervous about denouncing the approximation as an error because the authors are all renowned thought leaders in the drug discovery field. Furthermore, the journal impact factor of NRDD is a significant multiple of my underwhelming h-index and any error of such apparent grossness would surely have been detected during the rigorous peer review process applied by this elite journal. It turns out that my nervousness was indeed well placed and, when calculated at 300 K, the product RT actually serves as an annihilation operator that eliminates the dimensionality associated with Kd. This also explains why a temperature of 300 K must be used when calculating the ligand efficiency even though biochemical assays are usually run at human body temperature (310 K). 

I became convinced of the validity of the above approximation recently after examining a manuscript by the world-renowned expert on tetrodotoxin pharmacology, Prof. Angelique Bouchard-Duvalier of the Port-au-Prince Institute of Biogerontology, who is currently on secondment to the Budapest Enthalpomics Group (BEG). The manuscript has not yet been made publicly available although I was able to access it with the help of my associate ‘Anastasia Nikolaeva’ (she decamped last year from Tel Aviv to Uzbekistan and, to Derek’s likely disapproval, is currently running an open access journal out of a van in Samarkand). There is no doubt that this genuinely disruptive study will comprehensively reshape the generative AI landscape, enabling drug discovery scientists, for the very first time, to rationally design novel clinical candidates using only gene sequences as input.

Prof. Bouchard-Duvalier’s seminal study clearly demonstrates that it is indeed possible to eliminate the need to define standard states for the thermodynamic analysis of liquid solutions, provided that the appropriate temperature is used. The math is truly formidable (my rudimentary understanding of Haitian patois didn’t help either) and involves first projecting the atomic isothermal compressibility matrix into the quadrupole-normalized polarizability tensor before applying the Barone-Samedi transformation, followed by hepatic eigenvalue extraction using the algorithm introduced by E. V. Tooms (a reclusive Baltimore resident better known for his research in analytic topology). ‘Anastasia Nikolaeva’ was also able to ‘liberate’ a prepared press release in which a beaming BEG director Prof. Kígyó Olaj explains that, “possibilities are limitless now that we have eliminated the standard state from solution thermodynamics and thereby consigned the tedious and needlessly restrictive Second Law to the dustbin of history." 

Saturday, 1 April 2017

A concentration of scoring functions

<< previous || next >>

Researchers at The Hungarian Institute Of Thermodynamics have published a number of seminal articles on the interplay of enthalpy and entropy in areas ranging from physical chemistry to socioeconomics. For example, the cause of World War 1 (also known as 'The Great War' although I doubt whether any of its participants thought that it was that great) was traced to a singularity in the Habsburg Partition Function. In a nutshell, the problem was shown to be a surfeit of the wrong type of entropy (which led to Franz Ferdinand's driver getting lost) coupled with a deficit in the right type of entropy (which would have prevented Gavrilo Princip's bullets from finding their targets). However, it is unlikely that any amount of the right type of entropy could have saved the hapless Maximilian I of Mexico, who generously volunteered to be Emperor only to be shot by the ungrateful Mexicans.

The most recent study from BEG (Budapest Enthalpomics Group) is little short of sensational. Unfortunately it's not available online and the poor fax quality, coupled with my rudimentary grasp of Hungarian, have made the going hard. The essence of this seminal study is that the performance of scoring functions can be significantly improved by including the concentration unit (in which affinity is expressed) as a parameter in the fitting process. The casual observer of virtual screening may have wondered why scoring functions are trained with affinity but validated by enrichment. By treating the concentration unit as a parameter in the fitting process, the authors were able to achieve unprecedented accuracy of prediction and the phone call from Stockholm would seem to be a foregone conclusion. Commenting on these seminal findings, Prof. Kígyó Olaj, the director of the Institute said, "Now we no longer need to use ROC plots to mask feeble correlations between predicted and measured affinity".     

Friday, 6 January 2017

Confessions of a Units Nazi

Regular readers (both of them) of this blog will know that I have an interest, which some might term an obsession, with units. At high school in Trinidad, we had the importance of units beaten into us by the Holy Ghost Fathers and, for some of the more refractory cases, the beating was quite literal. I was taught physics by the much loved, although somewhat highly-strung, Fr. Knolly Knox (aka Knox By Night) who, as Dean of the First Form, used to give 'licks' with a cane of hibiscus (presumably chosen for its tensile properties). You quickly learned not to mess with The Holy Ghost Fathers, especially the Principal, Fr. Arthur Lai Fook (aka Jap), and it was a brave student who responded to the request by Fr. Pedro Valdez to define the dyne by answering, "Fah, it what happen after living". Fr. Pedro was a gentle soul although his brother, Fr. Toba, who taught me Latin, would lob a blackboard eraser with reproducible inaccuracy at any student who had the temerity to doze off during the Second Punic War while Hannibal and his elephants were steamrollering the hapless legions of Gaius Flaminius into Lake Trasimene. At least we didn't have detention at my school. Actually we did have detention only it was called 'penance'. Each and every student also had a Judgement Book in which was entered a mark (out of 10) for each subject each and every week. A mark of 5 (or less) or a failure to return one's Judgement Book, duly signed by parent or guardian, by Wednesday morning earned the transgressor a corrective package of Licks and Penance.  As a thoughtful child, I managed to shield my parents from this irksome bureaucracy and, in any case, it was simply safer that The Holy Ghost Fathers were never given the opportunity to familiarize themselves with the authentic parental signatures.


I used to think that 'Virtus et Scientia' was Latin for 'Licks and Penance'  (17-Feb-2018 update)


What we learned from the Holy Ghost Fathers was that most physical quantities have dimensions and if the quantities on the opposite sides of the 'equal sign' in an equation have different dimensions then it is a sign of an unforced error rather than a penetrating insight. For example the dimensions of force are MLT-2 (M = mass; L = length; T = time) and you are free to express forces in newtons, dynes or poundals as you prefer. You can think of a physical quantity as a number multiplied by a unit and, without the unit, the number is meaningless. Units are extremely important but at the same time they are arbitrary in the sense that if your physical insight changes when you change a unit then it is neither physical nor an insight. Here's a good illustration of why dimensional analysis matters.

I have blogged ( 12 | 3 ) about how building the a concentration unit into the definition of ligand efficiency (LE) results in a metric that is physically meaningless (even though it remains a useful instrument of propaganda) and, for the masochists among you, there's also the LE metric critique in JCAMD. The problem can be linked to a lack of recognition of the fact that logarithms can only be calculated for numbers (which lack units). However, LE has another 'units issue' which is connected with the fact that it is a molar energy that is scaled in the definition of LE rather than pIC50 or pKd. This needn't be an issue but, unfortunately, it is. LE is defined by dividing a molar energy by the number of non-hydrogen atoms in the molecular structure and there is nothing in the definition of LE that says that the energy has to be expressed in any particular unit. This means that you can define LE using any energy unit that you want to. Some 'experts' appear to believe that dividing a molar energy by number of non-hydrogen atoms relieves them of the responsibility to report units. I'm referring, of course, to the practise of multiplying pIC50 or pKd by 1.37 when calculating LE. You might ask why people do this, especially given that 'experts' tout the simplicity of LE and they don't multiply pIC50 or pKd by 1.37 when they calculate LipE/LLE. Don't ask me because I'm neither expert nor 'expert'.

Let's take a look at this NRDD article on LE metrics and I'd like you to go straight to Box 1 (Ligand efficiency metrics). Six numbered equations are shown in Box 1 and it is stated towards the end of the first paragraph that "each equation corresponds to a mathematically valid function".  This statement is incorrect because the first equation (1) in Box 1 is not a mathematically valid function. The reason for this is that the logarithm function cannot take as its argument a quantity, such as Kd, that has units. Equation (5), which defines LLEAT, is mathematically valid although it differs from the mathematically ambiguous equation that was originally used to define LLEAT

To be honest, I think that Box 1 is probably beyond repair by conventional erratum and I'll back this opinion with an example:


"Assuming standard conditions of aqueous solution at 300K, neutral pH and remaining concentrations of 1M,
 –2.303RTlog(Kd/C°) approximates to –1.37 × log(Kd) kcal/mol." 

At my school in Trinidad this would have been called a 'ratch' and, once detected, it would have earned its perpetrator a corrective package of Licks and Penance. I don't think even the Holy Ghost Fathers could have exorcised a concentration unit quite this efficiently. 

In some physical chemistry literature, Kd is defined as a dimensionless quantity by including C° in the definition of Kd. However, in the literature of biochemistry, biophysics and medicinal chemistry,  Kd  is usually quoted in units of concentration. Binding free energy has the same value and same dependence on C° regardless of  which of the two conventions is used to define Kd 
(Update 17-Feb-2018) 

I'd now like to talk a bit about the 'p' operator that we use to transform IC50 and Kd values into logarithms. This makes it much easier to perceive structure-activity relationships and provides a better representation of measurement precision than when the IC50 and Kd values themselves are used. To calculate pKd,, first express Kd in molar concentration units, dump the units and calculate minus the logarithm of the number. I realize that this may come across as arm waving but the process of converting  Kd, to  pKd, can actually be expressed exactly in mathematical terms as follows:

 pKd = –log10(Kd/M)

The 'p' operator has a 1 M concentration built into it. Although this choice of unit is arbitrary, it doesn't cause any problems if you're doing sensible things (e.g. subtracting them from each other) with the pKd values. If, however, you're doing silly things (e.g. dividing them by numbers of non-hydrogen atoms) with the pKd values then the plot starts to unravel faster than you can say 'Brexit means Brexit'. 

I'd like you take a look at another article which also has a Box 1 although I won't bother you with another tiresome 'spot the errors' quiz. The equation that I'll focus on is:

pKd = pKH + pKS 

This equation describes the decomposition of affinity into enthalpic and entropic contributions and you might think this means that you can write:

Kd = KH × KS 

As Prof. Pauli would have observed, this is an error in the 'not even wrong' category and it is clear that a difference in opinion as to the importance of units was as much responsible for the unraveling of the Austro-Hungarian empire as that unfortunate wrong turn in pre-SatNav Sarajevo. The 'p' operator implies that each of KdKH and Khas units of concentration. However, multiplying two such quantities will give a quantity that has units of concentration squared. 

It is actually possible to decompose Kd into enthalpic and entropic contributions a valid manner but you need to be thinking carefully about the meaning of the standard state. As noted previously DG° depends on the concentration used to define the standard state. This is a consequence of the dependence of DS° on the standard concentration and DH is independent of the standard concentration (the standard state is assumed to be a dilute solution). This suggests defining KS as quantity with units of concentration and Kas a quantity without units.

This is probably a good point to wrap things up. My advice to all the authors of the featured NRDD and FMC articles is that they read (and make sure that they understand) the section of this article that is entitled '8. Ligand Efficiency and Additivity Analysis of Binding Free Energy'. This advice is especially relevant for those of the authors who consider themselves to be experts in thermodyamics.

May I wish all readers a happy, successful and metric-free 2017.

Wednesday, 6 January 2016

Looking back at 2015


I'll start the year by taking a look back at some of the 2015 blog posts. The dynamic range of the bullshitometer was severely tested last year and there was an element of 'pour encourager les autres' to more than one of the posts. I thought that it'd be fun to share some travel pics and the first is of the Danube in Belgrade (I'd dropped by to catch up with friends and deliver a harangue at the university).  The early evening light was quite perfect although I hope that I won't spoil your experience of the photo by telling you that there was a pig carcass floating about 100 m from it was taken.

 

I changed the title of the blog this year. I've not been involved with FBDD for some years now and molecular design was always my main interest. One of the ideas that I try to communicate is that there's more to design than just making predictions. After Belgrade, I dropped in at Fidelta in Zagreb where I delivered another harangue before heading south to Sarajevo.  I'm a keen student of history so it was inevitable that this would be the first photo I'd take in Sarajevo.


It seems so bizarre today. There had already been one assassination attempt for the day when the driver of the car took the fateful wrong turn that gave Gavrilo Princip the opportunity to fire two shots at the royal couple. Back in Vienna, Sophie was not always allowed out in public with Franz Ferdinand so the trip to Sarajevo may have been a special treat for her. What if SatNav had already been invented but, then again, what if Queen Victoria's eldest child had succeeded her to the throne?

Part of the problem was that, as a lowly Czech countess, Sophie was not considered an appropriate match for the Habsburg heir by Franz Josef (the reigning emperor and a puritanical old killjoy) and there were rules (although metrics and Lean Six Sigma 'belts' had, thankfully, not yet been invented). One of the rules was that the children of Sophie and Franz Ferdinand were barred from succession. It is somewhat ironic that poor Franz Ferdinand was never even supposed to be crown prince in the first place and only got the job because his cousin Rudolf had abruptly removed himself from the Habsburg line of succession a quarter of a century previously. 

All this talk of puritanical rules serves as a reminder that, before moving on, I need to point you towards a friend's blog post on roundheads (who were bigger killjoys than Franz Josef or even Lean Six Sigma 'belts') and cavaliers in drug discovery.  I really like the term 'roundhead' and I think you do have to agree that it's a lot politer than 'compound quality jackboot'. Terms like 'roundhead' and 'jackboot' are invariably associated with pain and that brings me to the next topic which is PAINS. My interest in this topic was piqued by a PAINS-shaming post at Practical Fragments and I have to thank my friends there for launching me on what has proven to be a most stimulating, although at times disturbing, line of inquiry.

My first post on PAINS examined some of the basic science and cheminformatics behind the substructural filters used. One observation that I'll make is that cheminformaticians would have done themselves rather more credit if, instead of implementing PAINS filters quite so enthusiastically, they'd first taken a more forensic look at how the filters had been derived. Singlet oxygen is an integral component of the AlphaScreen technology used in all six assays that formed the basis of the original PAINS study and the second post explored some of the consequences of this reliance on singlet oxygen. The third post was written as a 'get out of jail' card for those who need to get their use of PAINS past manuscript reviewers but, on a more serious note, it does pose some questions about how much we actually know about the behavior of PAINS compounds. The final PAINS post emphasized the need to make a clear distinction in science between what we know and what we believe. If we are unable (or unwilling) to demonstrate that we can do this in drug discovery then those who fund our work may conclude that the difficulties we face are of our own making.

There's actually a lot more to Sarajevo than dead Habsburgs and the city hosted the  1984 Winter Olympics. I took a taxi to the top of the bobsled run and walked back down to the city. Here are some photos. 




  


  
  

So I guess you're wondering where the 1984 bobsled run fits into drug discovery.  Ligand efficiency is, in essence, about slopes and intercepts and, like bobsledders, ligand efficiency advocates prefer not to think about intercepts.  I did two posts on ligand efficiency in 2015. The first post was a response to an article in which our criticism of ligand efficiency metrics was denounced as noise although, in the manner of Pravda, the article didn't actually say what the criticism was and I was left with the impression of a panicky batsman desperately trying to fend off a throat ball that had lifted sharply off just short of a length. The second post explored the link between ligand efficiency and homeopathy. 

I have described ligand efficiency as not even wrong and it also fits snugly into the voodoo thermodynamics category. Sometimes I think that if a coiled dog turd could be converted to molar energy units and scaled by coil radius then it would get adopted as a metric (which we might call 'scatological efficiency'). Voodoo thermodynamics is likely to feature more frequently in 2016 although I did manage one post on this topic in 2015. 

I took the train from Sarajevo to Mostar and the next four photos show a guy jumping, as is the local custom, off the reconstructed Stari Most into the Neretva River.


 





Now I guess you're wondering what a guy jumping off a bridge in Herzegovina has to do with molecular design and the quick answer is nothing at all. During the course of the year I jumped off a bridge of sorts (more accurately out of my applicability domain) with a post on Open Access and there'll hopefully be more of this sort of thing this year. This is probably a good point to wrap up the review of 2015 and I look forward to seeing you towards the end of the month when you'll meet the boys who cried wolf.


Friday, 30 October 2015

Voodoo thermodynamics for dummies


Metrics are like  the heads of the Hydra. Dispatch one and two pop up to take its place.


So #RealTimeChem week is over and it's time to return to the topic of metrics and readers of this blog will be aware that this is a recurring theme here. Sometimes, to give them a more 'hard science feel', drug discovery metrics are cast in thermodynamic terms and 'conversion' of IC50 to free energy provides a good example of the problem. Ligand efficiency (LE) was originally defined by scaling free energy of binding by molecular size and it is instructive to observe how toys are ejected from prams when the thermodynamic basis of LE is challenged.


The most important point to note about a metric is that it's supposed to measure something and, regardless of how much you wave your arms and how noisily you assert the metric's usefulness, the metric still needs to measure. That's why we call it a 'metric' and not a 'security blanket for timid medicinal chemists' nor a 'floatation device for self-appointed experts and wannabe thought-leaders'. To be useful, a metric also has to measure something relevant and, in many drug discovery scenarios, that means being predictive of the chemical or biological behavior of compounds. Drug discovery metrics (and guidelines) are often based on trends in data and the strength of the trend tells us how much weight we should give to metrics and how rigidly we should adhere to guidelines. In the metric business, relevance trumps simplicity and even the most anemic of trends can acquire eye-wateringly impressive significance when powered by enough data.


I'll start my review of the article featured in this post by saying that, had the manuscript been sent to me, the response the editor would have been something between 'why have you sent this out for review' and 'this manuscript needs to be put out of its misery as swiftly and mercifully as possible'. The article appears to be the write up for material presented in webinar format which was reviewed less than favorably.  The authors have made a few changes and what was previously called SEEnthalpy (Simplistic Estimate of Enthalpy) is now called PEnthalpy (Proxy for Enthalpy) but the fatal design flaws in the original metric remain and the review of that webinar will show what happens when you wander by mistake into the mess that metrics make.


Before we try to cone these thermodynamic proxies in the searchlights, it may be an idea to ask why we should worry about enthalpy or entropy when drug action is driven by affinity and free concentration. That's a good question and, to be quite honest, I really don't know the answer. Isothermal titration calorimetry (ITC) is an excellent, label-free method for measuring affinity and enthalpy of binding. However, the idea that the thermodynamic signature for binding of a compound to a protein will somehow be predictive of the behavior of the compound is all sorts of situations that do not involve that protein does seem to be entering the realms of wild conjecture.  There is also the question of how isothermal systems like live humans can 'sense' the benefits of an enthalpically-optimized drug. Needless to say, these are questions that some ITC experts and many aspiring thought-leaders would prefer that you didn't think too hard about.


So let's take a look at the thermodynamic proxies which are defined in terms of the total number (HBT) of hydrogen bond donors and acceptors and the number (RB) of rotatable bonds.  The proxies are defined as follows:



 PEnthalpy = HBT/(RB + HBT)                                          (1)

 PEntropy  = RB/(RB + HBT)                                              (2)

 PEnthalpy  +  PEntropy  =  1                                                 (3)

The proxies predict that the enthalpy and entropy changes associated with binding are functions only of ligand structure and therefore are of no value for comparing the thermodynamics for a particular ligand binding to different proteins as one might want to do when assessing selectivity. Equation (3) shows that there is effectively only one metric (what a relief) since the two proxies are perfectly anticorrelated so each is as effective as the other as a predictor of either the enthalpy or entropy changes associated with ligand binding.

Now you may remember in the webinar that one of the authors of the featured article was telling us at 22:43 that "entropy comes from non-direct hydrophobic interactions like rotatable bonds".  At least now they seem to realize that the rotatable bonds represent degrees of freedom although I don't get the impression from reading the article that have a particularly solid grasp of the underlying physicochemical principles. Freezing rotatable bonds is an established medicinal chemistry tactic for increasing affinity and, if successful, we expect it to lead to a more favorable entropy of binding which some self-appointed thought-leaders would assert is a bad way to increase affinity.  Trying to keep an open mind on this issue, I suggest that we might follow the lead of British Rail and try to define right and wrong types of entropy.


One of the criticisms that I made of the webinar was that no attempt was made to validate the metrics against measured values of binding enthalpy and entropy. In the article, the metrics are evaluated against a small data set of measured values.  As I mentioned earlier, there is effectively only one metric because the two metrics are perfectly anti-correlated so you need to look beyond the fit of the data to the metrics if you want to assess what I'll call the 'thermodynamic connection'. This means digging into the supplementary information.  I found the following on page 5 of the SI:


-TdeltaS =  159.80522 − 343.46172*Pentropy(RB/(HBT + RB))              (4)

which implies that:

TdeltaS  =  −159.80522 + 343.46172*Pentropy(RB/(HBT + RB))            (5)

These equations tell us that the change in entropy associated with binding actually increases with RB rather than decreasing with RB as one would expect for degrees of freedom that become frozen when the intermolecular complex forms. When you're assessing proxies for thermodynamic quantities it's a really good idea to take a look at the root mean square error (RMSE) for the fit of the quantity to the proxy.  The RMSE values for fitting ΔH and TΔS are 28.35 kJ/mol and 29.30 kJ/mol respectively and I will leave it to you, the reader, to decide for yourself whether or not you consider these RMSE values to justify PEnthalpy an PEntropy being called thermodynamic proxies.  The alert reader might ask where the units for ΔH and TΔS° came from since neither the article nor the the SI provides this information and the answer is that you need to go to the source from which the ΔH and TΔS° values were taken to find out.

Now you'll recall that these thermodynamic proxies predict constant values of ΔH and TΔS° for a binding of a given compound to any protein (even those proteins to which it does not bind). The  ΔG° values for the compounds in the small data set used to evaluate the thermodynamic proxies lie in a relatively narrow range (i.e. less than the RMSE values mentioned in the previous paragraph) from −37.6 kJ/mol to −57.3 kJ/mol and are not representative of the affinity of these compounds for proteins against which they had not been optimized. Any guesses how the RMSE values for fittling the data would have differed if  ΔH and TΔS° values had been used for each compound binding to each of the protein targets?


Now if you've you've kept up to date with the latest developments in the drug discovery metric field, you'll know that even when the mathematical basis of a metric is fragile, there exists the much-exercised option of touting the metric's simplicity and claiming that it is still useful. Provided that nobody calls your bluff, metrics can prove to be a very useful propaganda instruments. The featured article does present examples of data analysis based on the using the thermodynamic proxies as descriptors and one general criticism that I will make of this analysis is that most of it is based on the significance rather the strength of trends. When you tout the significance of a trend, you're saying as much about the size of your data set as you are about the strength of the trend in it. This point is discussed in our correlation inflation article and I'd suggest taking a particularly close look at what we had to say about the analysis in this much-cited article.


I'd like to focus on the analysis presented in the section entitled 'GSK PKIS Dataset' and which explored correlations between protein kinase % inhibition and a number of molecular descriptors. The authors state,


"In addition to PEnthalpy, we assessed the correlation across a variety for physicochemical properties including molecular weight, polar surface area, and logP in addition to PEnthalpy  (Fig. 5)"   


This statement is actually inaccurate because they have assessed the significance of the correlations rather than the correlations themselves. Although they may have done the assessment for logP and polar surface area, the results of these assessments do not seem have materialized in Fig. 5 and we are left to speculate as to why. The strongest correlation between PEnthalpy  and % inhibition was observed for CDK3/cyclinE and the plot is shown in Fig. 5b. I invite you, the reader, to ask yourself whether the correlation shown in Fig. 5b would be useful in a drug discovery project.


Since the title of the post mentions voodoo thermodynamics, we should take a look at this in the context of the article and the best place to look is in the Discussion section.  We are actually spoiled for choice when looking for examples of voodoo thermodynamics there but take a look at:


"It is assumed in the literature that the "enthalpically driven compound series" with fewer RBs tend to be (generally) lower MW compounds as well. In contrast, in cases where selectivity is steeper among compounds in a series for which activity and selectivity is likely governed by compounds with relatively more RB versus HBA and HBA [sic], than when the entropic contributors are dominating."


So that's about as much voodoo thermodynamics as I can take for a while so, if it's OK with you, I'll finish by addressing a couple of points to the authors of this article. The flagship product of company with which the authors are associated is a database system for integrating chemical and biological data. Although I'm not that familiar with this database system, responses to my questions during the course of a discussion in the FBDD LinkedIn group suggested that a number of cheminformatic issues have been carefully thought through and that the database system could be very useful in drug discovery. One problem with the featured article is that its scientific weaknesses could lead to some customers losing confidence in the database system. Secondly, the folk who created the database system (and keep it running) may have only limited opportunities to publish and scientifically weak publications by colleagues who are perhaps less focused on what actually pays the bills may breed some resentment.


That's where I'll wrap because there is only so much voodoo thermodynamics that one can take in a day so, as we say in Brazil, 'até mais'.

      

Thursday, 28 March 2013

Efficient Voodoo Thermodynamics

So you’ve got a little data analysis problem.  You have some compounds with a range of IC50 values and you’d like to explore the extent that molecular size contributes to potency/affinity.  Here’s one suggestion for starters.  Plot pIC50 against your favourite measure of  molecular size which could be number of number of non-hydrogen atoms, molecular weight and look at the residuals which tell you how much each compound beats the trend (or is beaten by it).  Let’s start by defining  pIC50 which you can calculate from:

 pIC50 = -log(IC50/M)                                                                              1
You might be asking why I’m not suggesting that you use the standard Gibbs free energy of binding which is defined by:

ΔG° = RTln(Kd/C°)                                                                                 2
The main reason for not doing this is that we can’t.  You’ll notice that ΔG° is calculated from Kd and not  IC50 and the two are not the same thing even though they both have units of concentration.  Those of you who have worked on kinase projects may have even used IC50 values measured at different ATP concentrations to get a better idea how much kick your inhibitors will have at physiological ATP concentration.  Put another way, you can measure the concentration of sugar in your coffee and plug this into equation 2 but that does not make what you calculate a standard Gibbs free energy of binding.   If you’ve measured IC50 then you really should use  pIC50 in this analysis.  Converting pIC50 to ΔG° is technically incorrect and arguably pretentious since the converted pIC50 can give the impression that it is somehow more thermodynamic than that from which it was calculated.  Converting pIC50  to ΔG° also introduces additional units (of energy/mole) and there is always a degree of irony when these units, which may have been introduced to just make biological data look more physical, get lost when the results are presented.

So let’s get back to the data analysis problem.  Suppose that we’ve plotted  pIC50 against number of heavy (i.e. non-hydrogen) atoms and the next step is to fit the data.  Best way to start is to fit a straight line although you could also fit a curve if the data justifies this.  Let’s assume that we’re fitting the straight line:
  pIC50 = A + B×NHA                                                                            3

I realise that using an intercept term (A) will cause a few eyebrows to become raised.  Surely the line of fit should go through the origin?   There is a problem with this line of thinking and it’s helpful now to talk instead in terms of affinity and ΔG° to develop the point a bit more.  You might be thinking that in the limit of zero molecular size a compound should have zero free energy of binding.  However, a zero free energy of binding corresponds to Kd being equal to the standard concentration and you’ll remember that the choice of standard state is arbitrary.  If you must derive insights from thermodynamic measurements then the very least that you can do is to ensure that any insights you derive are invariant with respect to the value of the standard concentration.  

When you use ligand efficiency (-ΔG°/NHA) you’re effectively assuming that the value of ΔG° should be directly proportional to the number of heavy atoms in the ligand molecule.  One consequence of defining ligand efficiency in this manner is that relative values of ligand efficiency for compounds with different numbers of heavy atoms will change if you change (as thermodynamics tells us that we are allowed to do) the standard concentration used to define ΔG°.  I've droned on enough though and it's time to check out.  I will however, leave you with the question of whether it makes sense to try to correct ligand efficiency for the effects of molecular size. 

Sunday, 17 March 2013

The wrong kind of free energy

It has been some time since I last blogged.   There has been the distraction of a little drug-likeness project and there is still much to do before I leave Brasil at the end of May.  I have, however, set up a twitter account (@pwk2013) which I use in futile attempts to engage minor celebrities (like God and Richard Dawkins) in conversation.   I’ll be taking a look at some aspects of thermodynamics in this post and I’ll start by writing the relationship between dissociation constant and the standard free energy of binding:

ΔG° = RTln(Kd/C°)

I realise that this might look a bit different to what you’re used to so I’ll try to explain why its been written like this.  First of all Kd has units of concentration and logarithms are only defined for numbers (i.e. unitless quantities) so you might want to consider the possibility that what you normally write is wrong (apologies for hideous pun).   The other thing that you need to remember is that the standard free energy of binding depends on the choice of standard state and writing the equation like this makes this connection explicit.   ΔG° is negative for a 10nM compound when the standard concentration is 1M.  However, if we were to define the standard concentration to be 1nM then   ΔG° would be positive.  If you consider the idea of a 1nM standard concentration to be offensive, think of a 1M solution of your favourite protein...
One of the misconceptions that drug discovery researchers tend to have is that you can’t call it thermodynamics unless you measure both enthalpy and entropy.  One reason for this is the use of isothermal titration calorimetry (ITC) to measure affinity.  ITC is a direct, label-free method for measuring affinity and you get the enthalpy as part of the package.  One of the important things to remember about thermodynamics is that you need to use the most appropriate thermodynamic quantity to describe the phenomenon that you’re interested in and for binding at constant pressure this is the Gibbs free energy.  In contrast, the engineer scaling up a synthetic reaction for manufacturing a drug will usually be more interested in the heat and volume (i.e. gaseous bi-products) generated by the reaction because failure to control these can result in the Big Kaboom.

One idea that has emerged as ITC has become more accessible in drug discovery is that enthalpy-driven binding is somehow better than entropy-driven binding.  I remain sceptical and ask the rhetorical question of how an isothermal system such as a human taking a drug senses the degree to which engagement of the drug’s target(s) is driven by enthalpy changes.   If measuring enthalpy of binding in addition to affinity helps us to predict affinity for compounds yet to be synthesised or gives us clear (i.e. keeping arms stationary) insights into the nature of binding then it makes sense to do it.  However, I’m not seeing this being done convincingly and sometimes wonder whether enthalpy optimisation is just a case of the wrong kind of snow...

We tend to interpret affinity in terms of contacts between protein and ligand although one must always remember that the contribution of a particular contact to affinity is not in general an experimental observable.   Proteins and their ligands associate in an aqueous environment and the non-local nature of the hydrophobic effect further complicates attempts to use structure to rationalise affinity.  I often hear hydrophobic interactions being described as non-directional and in my view this is complete bollocks since interactions are forces and forces are vectors which by definition have direction.  On this subject, I can't resist telling the tale of the Human Resources department getting the Scalar Quantity Award for having magnitude but no direction...
I think this is a good place to leave things and I'll try to be back in a week or some more specific stuff in a follow up to this post.