Showing posts with label MAMO. Show all posts
Showing posts with label MAMO. Show all posts

Monday, 1 April 2019

Enthalpy-driven pharmacokinetics


<< previous || next >>

Drug design is a multi-objective endeavor. Some objectives such as maximization of affinity against target(s) and minimization of affinity against anti-targets are easily defined. Other objectives such as controllability of exposure are much less easily defined and this means that drug design is indirect. Controllability of exposure is the focus of pharmacokinetic optimization and I recently became aware of an exciting new development that will surely reshape the pharmacokinetic field and transform drug discovery beyond all recognition.

Target engagement potential and the multiple objectives of drug design 

The exciting results from another seminal study by the Budapest Enthalpomics Group (BEG) look set to revolutionize the way that we think about pharmacokinetics. The work, funded by Mothers Against Molecular Obesity (MAMO), was described in a book chapter that was prematurely posted online although the error does appear to have been recognized because the article is no longer publicly visible. In a study that will undoubtedly disrupt drug discovery, it is clearly shown that the thermodynamic signature for the binding of a ligand to a protein is predictive of the physicochemical behavior of the ligand even when that protein is absent from the system.

The theoretical treatment introduced in this groundbreaking study is formidable and the starting point is is an eigenvalue decomposition of the entropic field tensor in reciprocal heavy atom space. Machine learning using the Blofeld Optimized Ligand Lipophilicity Of Cryogenic Krypton Solvates algorithm demonstrates unequivocally that the Grayling annihilation operator can be used to eliminate the entropy (and its efficacy-limiting dependence on the definition of the standard state) from any in vivo system. This leads to highly-efficient, enthalpy-driven pharmacokinetics in which the clearance (shown to be strongly correlated with the trace of the entropic field tensor) can be significantly attenuated. "The key to successful pharmacokinetic optimization is to eliminate the elimination", explains institute director Prof Kígyó Olaj, "and we have shown, for the first time, that entropy can be exorcised from the equations of pharmacokinetics with even greater efficiency than if it had been done by Torquemada himself."

Sunday, 30 September 2018

Hydrogen bonding asymmetries

Next >>

Have you ever wondered why the Rule of 5 (Ro5) specifies hydrogen bond (HB) thresholds of 10 acceptors but only 5 donors? This is, perhaps, the prototypical example of what I'll call a 'hydrogen bonding asymmetry' and it is sometimes invoked in support of the folklore that HB donors are somehow 'worse' than HB acceptors in drug design. I have, on occasion, tried to track down the source of this folklore but that trail has always gone cold on me. In any case, I don't think the HB asymmetry in Ro5 has any physical significance since HB acceptors (especially as defined for Ro5) tend to be more common in chemical structures of interest to medicinal chemists than HB donors. This was discussed in our correlation inflation article and the bigger Ro5 question for me is why the high polarity limit is defined by counts of HB donors and acceptors while the low polarity limit is defined in terms of lipophilicity. As may become a blogging habit, I'll include some random photos (these are from a visit to India late in 2013) to break up the text a bit. 

Drum fest at Buland Darwaza

It was this article in JCAMD about the 'polarized' nature of protein-ligand interfaces that got me thinking again about hydrogen bonding asymmetries. The study found that proteins donate twice as many HBs as they accepted. While the observation is certainly interesting, I do think that the authors might be over-interpreting it. For example, the authors suggest that it appears to be an underlying explanation for Ro5 and they may find that there are significant differences in their definitions of HB acceptors and those used to apply Ro5. The authors also state "Peptidyl ligands, on the other hand, showed no strong preference for donating versus accepting H-bonds". This observation would more be consistent with 'polarization' of protein-ligand interfaces being determined by nature of the ligand.

The authors assert that "lone pairs available to accept H-bonds are actually 1.6 times as prevalent as protons available to donate, both on the protein and ligand side of the interface." While it is appropriate to count lone pairs in situations where only one lone pair accepts an HB (e.g. when considering 1:1 hydrogen bonded complexes in low polarity solvents), I would argue that it is not appropriate to do so when considering biomolecular recognition in aqueous media because the acceptance of an HB by one oxygen lone pair makes the other lone pair less able to accept an HB. You can see this effect using molecular electrostatic potential as discussed in this article (see polarization effects section and Table 4). Put another way, how often is a carbonyl oxygen observed to accept two HBs from a binding partner? How many docking tools would explictly penalize a pose in which a carbonyl oxygen accepted two HBs?

As I see it, a typical protein is more likely to have a surplus of HB donors under normal physiological conditions. Some parts (e.g. serine, threonine, tyrosine and histidine side chains and the backbone) of a protein can be regarded as having equal numbers of HB donor and acceptor atoms. While the anionic side chains of aspartate and glutamate cannot donate HBs, the cationic side chains of arginine and lysine have five and three donor hydrogen atoms respectively while lacking HB acceptors. The tryptophan side chain has only a single HB donor (although its p-system is likely to be able to accept HBs) while each side chain of aspargine and glutamine has two donor hydrogen atoms and one acceptor oxygen atom. The histidine side chain is sometimes observed to be protonated in X-ray crystal structures which means that it should be considered to be more HB donor than HB acceptor in the constext of protein-ligand recognition. The tyrosine hydroxyl would be expected to be a stronger HB donor (and weaker HB acceptor) than the hydroxyls of either serine or threonine.  

A magical place

The study considers the "possibility is that nature avoids the presence of chemical groups bearing both H-bond donor and acceptor capacity, such as hydroxyl groups, in the binding sites of proteins or ligands" although it is not clear what glycobiologists would have to say about this. Let's think a bit about what happens when a hydroxyl group donates its hydrogen atom. Let's suppose you've spotted a nice juicy hydrogen bond acceptor at the bottom of a deep binding pocket that is otherwise hydrophobic. The ligandability is eye-wateringly awesome (the ligandometer is beeping loudly and appears to have gone into dynamic range overload). Even the tiresome Mothers Against Molecular Obesity (MAMO) are impressed and have recommended that you deploy a hydroxyl group since this will be great for property forecast index (PFI). What could possibly go wrong?

The main problem is that the hydroxyl HB donor comes with baggage. In order to donate an HB to the acceptor at the bottom of that pocket, you're going to need to force an HB acceptor into contact with the non-polar part of that binding pocket. Although this contact is not inherently repulsive, it is destabilizing. Another factor is that donation of an HB by the hydroxyl group is likely to increase the HB basicity of the oxygen (which will exacerbate the problem). You can think of other neutral HB donors (e.g. amide NH) but the vast majority of them come with baggage the form of an accompanying HB acceptor. Exceptions such as NH in pyrrole (not renowned for stability) and indole (steric demands) come with baggage of their own. In contrast, the drug designer has access to a diverse set (e.g. heteroaromatic N, nitrile N, tertiary amide O, sulfoxide O, ether O) of HB acceptors that are not accompanied by HB donors. If you use one of these, you don't have the problem of having to also accommodate a ligand HB donor.

This is a good place to wrap up. In the next post, I'll talk about a completely different type of hydrogen bonding asymmetry, but for now, I'll leave you with some photos from an afternoon spent admiring asses in the Rann of Kutch. 

Até mais!



Friday, 3 June 2016

Yet more on ligand efficiency metrics

In this post, I'll be responding to a couple of articles in the literature that cited our gentle critique of ligand efficiency metrics (LEMs). The critique has also been distilled into a harangue and  readers may find that a bit more digestible that the article. As we all know, Ligand Efficiency (LE), the original LEM was introduced to normalize affinity with respect to molecular size. Before getting started, I'd like to ask you, the reader, to ask yourself exactly what you take to mean by the term 'normalize'.

The first article which I'll call L2016 states:

Optimisation frequently increases molecular size, and on average there is a trade-off between potency and size gains, leading to little or no gain in LE [42,52] but an increase in SILE [52]. This, and the nonlinear dependence of LE on heavy atom count, together with thermodynamic considerations, has led some authors to question the validity of LE [76,77], while others support its use [52,78,79].

This statement is misleading because the "thermodynamic considerations" are that our perception of efficiency changes when we change the concentration units in which affinity and potency are expressed.  As such, LE is a physicochemically meaningless quantity and, in any case, references 52 and 78 precede our challenge to the thermodynamic validity of LE (although not an equivalent challenge in 2009). Reference 78 uses a mathematically invalid formula for LE when claiming to have shown that LE is mathematically valid and reference 79 creates much noise while evading the challenge. I have responded to reference 79 (aka the 'sound and fury article') in two blog posts ( 1 | 2 ).

This is a good place for a graphic to break up the text a bit and I'll use the table (pulled from an earlier post) that shows how our perception of ligand efficiency changes with the concentration units used to define affinity. I've used base 10 logarithms and dispensed with energy units (which are often discarded) to redefine LE as generalized LE (GLE) so that we can explore the effect of changing the concentration unit (which I've called a reference concentration). Please take special of note how a change in concentration unit can change your perception of efficiency for the three compounds. Do you think it makes sense to try to 'correct' LE for the effects of molecular size? 


Another article also cites our LEM critique.  Let's take a look at how the study, which I'll call M2016, responds to our criticism of LE (reference 69 in this study):

The appeal of LE and GE is in the convenience and rapidity with which these factors can be assessed during lead optimization, but the simplistic nature of these metrics requires an understanding of, and appreciation for, their inherent limitations when interpreting data.[67,68,69,70The relevance of LE as a metric has been challenged based on the lack of direct proportionality to molecular size and an inconsistency of the magnitude of effect between homologous series, both attributed to a fundamental invalidity underlying its mathematical derivation.[65,67] These criticisms have stimulated considerable discussion and provoked discourse that attempts to moderate the perspective and provide guidance on how to use LE and GE as rule-of-thumb metrics in lead optimization.[68,69,70]

To be blunt, I don't think that the M2016 study does actually respond to our criticism of LE as a metric which is that our perception of efficiency changes when we change the concentration unit with which we specify affinity or potency. This is an alarming characteristic for something that is presented as a tool for decision making and, if it were a navigational instrument, we'd be talking about fundamental design flaws rather than "limitations". The choice of 1 M is entirely arbitrary and selecting a particular concentration unit for calculation of LE places the burden of proof on those making the selection to demonstrate that this particular concentration unit is indeed the one that is most fit for purpose. 


The other class of LEM that is commonly encountered is exemplified by what is probably best termed lipophilic efficiency (LipE).  Although the term LLE is more often used, there appears to be some confusion as to whether this should be taken to mean ligand-lipophilicity efficiency or lipophilic ligand efficiency so it's probably safest to use LipE. Let's see what the M2016 study has to say about LipE:

LLE is an offsetting metric that reflects the difference in the affinity of a drug for its target versus water compared to the distribution of the drug between octanol and water, which is a measure of nonspecific lipophilic association.[69,12]

If I knew very little about LEMs, I would find this sentence a bit confusing although I think that it is essentially correct. We used (and possibly even introduced) the term 'offset' in the LEM context to describe metrics that are defined by subtracting risk factor from affinity (or potency). This is in contrast to LE and its variations which are defined by dividing affinity (or potency) by molecular size and can be described as scaled. There is still an arbitrary aspect to LipE in that we could ask whether (pIC50 - 0.5 ´ logP) might not be a better metric than  (pIC50 - logP).  Unlike LE, however, LipE is a quantity that actually has some physicochemical meaning, provided that the compound in question binds to its target in an uncharged form. Specifically, LipE can be considered to quantify the ease (or difficulty) of moving the compound from octanol to its binding site in the target as shown in the figure below:


Let's see what M2016 study has to say:

However, care needs to be exercised in applying this metric since it is dependent on the ionization state of a molecule, and either Log P or Log D should be used when appropriate.

This statement fails to acknowledge a third option which is that there may be situations in which  neither logP nor logD is appropriate for defining LipE. One such situation is when the compound binds to its target in a charged form. When this is the case, neither logP nor logD quantifies the ease (or difficulty) of moving the bound form of compound from octanol to water. As an aside, using logD to quantify compound quality suggests that increasing the extent of ionization will lead to better compounds and I hope that readers will see that this is a strategy that is likely to end in tears.

Let's take a look at LEMs from the perspective of folk who are working in lead optimization projects or doing hit-to-lead work. Merely questioning the value of LEMs is likely incur the wrath of Mothers Against Molecular Obesity (MAMO) so I'll stress that I'm not denying that excessive lipophilicity and  molecular size are undesirable. We even called them "risk factors" in our LEM critique. That said, in the compound quality and drug-likeness literature, it is much more common to read that X and Y are correlated, associated or linked than to actually be shown how strong the correlation, association or linkage is. When you do get shown the relationship between X and Y, it's usually all smoke and mirrors (e.g. graphics colored in lurid, traffic light hues). When reading M2016 you might be asking why can't we see the relationship between PFI and aqueous solubility presented more directly (or even why iPFI is preferred over PFI for hERG and promiscuity). A plot of one against the other perhaps even a correlation coefficient? Is it really too much to ask?

The reason for the smoke and mirrors is that the correlations are probably weak. Does this mean that we don't need to worry about risk factors like molecular size and lipophilicity? No, it most definitely does not! "You speak in more riddles than a Lean Six Sigma belt", I hear you say, "and you tell us that the correlations with the risk factors have been smoked and mirrored and yet we still need to worry about the risk factors".  Patience, dear reader, because the apparent paradox can be resolved once you realize some much stronger local correlations may be lurking beneath an anemic global correlation. What this means is that potencies of compounds in different projects (and different chemical series in the same project) may respond differently to risk factors like lipophilicity and molecular size. You need to start thinking of each LO project as special (although 'unique' might be a better term because 'special projects' were what used to happen to senior managers at ICI before they were put out to pasture).  

Another view of LEMs is that they represent reference lines. For example, we can plot potency against molecular size and draw a line with positive slope from a point corresponding to a 1 M IC50 on the potency axis and say that all points on the line correspond to the same LE. Analogously, we can draw a line of unit slope on a plot of pIC50 against logP and say that all points on the line correspond to the same LipE.  You might be thinking that these reference lines are a bit arbitrary and you'd be thinking along the right lines. The intercept on the potency axis is entirely arbitrary and that was the basis of our criticism of LE. A stronger case can be made for considering  a line of unit slope on a plot of pIC50 against logP to represent constant LipE but only if the compounds bind in uncharged forms to their target.

Let's get back to that project you're working on and let's suppose that you want to manage risk factors like lipophilicity and molecular size. Before you calculate all those LEMs for your project compounds, I'd like you to plot pIC50 against molecular size (it actually doesn't matter too much what measure of molecular size you use). What you now have in front of you is the response of potency to molecular size.  Do you see any correlation between pIC50 and molecular size? Why not try fitting a straight line to your data to get an idea of the strength of the correlation? The points that lie above the line of fit beat the trend in the data and the points that lie below the line are beaten by the trend. The residual for a point is simply the distance above the line for that point and its value tells you how much the activity that it represents beats the trend in the data. Are there structural features that might explain why some points are relatively distant from the line that you've fit? In case you hadn't realized it, you've just normalized your data. Vorsprung durch technik! Here's a graphic to give you an idea how this might work. 



The relationship between affinity and molecular size shown in the plot above is likely to be a lot tighter than what you'll see for a typical project. In the early stages of a project, the range in activity for the project compounds will often be too narrow for the response of activity to risk factor to be discerned. You can make assumptions about the response of affinity (or potency) to risk factor (e.g. that LipE will remain constant during optimization) in order to forecast outcome but it's really important to continually monitor the response of activity to risk factor to check that your assumptions still hold. If affinity (or potency) is strongly correlated with risk factor then you want the response to risk factor to be as steep as possible. Could this be something to think about when trying to prioritize between series?  

So it's been a long post and there are only so many metrics that one can take in a day. If you want to base your decisions on metrics that cause your perception to change with units then as consenting adults you are free to do so (just as you are free to use astrological charts or to seek the face of a deity in clouds). A former prime minister of India drank a glass of his own urine every day and lived to 98. Who would have predicted that? Our LEM critique was entitled 'Ligand efficiency metrics considered harmful' and I now I need to say why. When doing property-based design, it is vital to get as full an understanding as possible of the response of affinity (or potency) to each of the properties in which you're interested. If exploring the relationship between X and Y, it is generally best to analyse the data as directly as possible and to keep X and Y separate (as opposed to looking at the response of a function of Y and X to X). When you use LEMs you're also making assumptions about the response of Y to X and you need to ask yourself whether that's a sensible way to explore the response of Y to X. If you want to normalize potency by risk factor, would you prefer to use the trend that you've actually observed in your data or an arbitrary trend that 'experts' recommend on the basis that it's "simple"?

Next week, PAINS...

Thursday, 10 March 2016

Ligand efficiency beyond the rule of 5


One recurring theme in this blog is that the link between physicochemical properties and undesirable behavior of compounds in vivo may not be as strong as property-based design 'experts' would have us believe. To be credible, guidelines for drug discovery need to reflect trends observed in relevant, measured data and the strengths of these trends tells you how much weight you should give to the guidelines. Drug discovery guidelines are often specified in terms of metrics, such as Ligand Efficiency (LE) or property forecast index (PFI), and it is important to be aware that every metric encodes assumptions (although these are rarely articulated).

The most famous set of guidelines for drug discovery is known as the rule of 5 (Ro5) which is essentially a statement of physicochemical property distributions for compounds that had progressed at least as far as as Phase II at some point before the Ro5 article was published in 1997. It is important to remember (some 'experts' have short memories) that Ro5 was originally presented as a set of guidelines for oral absorption. Personally, I have never regarded Ro5 as particularly helpful in practical lead optimization since it provides no guidance as to how suboptimal ADMET characteristics of compliant compounds can be improved. Furthermore, Ro5 is not particularly enlightening with respect to the consequences of straying out the allowed region and into 'die Verbotenezone'.

Nobody reading this blog needs to be reminded that drug discovery is an activity that has been under the cosh for some time and a number of publications ( 1 | 2 | 3 | 4 ) examine potential opportunities outside the chemical space 'enclosed' by Ro5. Given that drug-likeness is not the secure concept that those who claim to be leading our thoughts would have us believe, I do think that we really need to be a bit more open minded in our views as to the regions of chemical space in which we are prepared to work. That said, you cannot afford to perform shaky analysis when proposing that people might consider doing things differently because that will only hand a heavy cudgel to the roundheads for them to beat you with.

The article that I'll be discussing has already been Pipelined and this post has a much narrower focus than Derek's post. The featured study defines three regions of chemical space:  Ro5 (rule of 5), eRo5 (extended rule of 5) and bRo5 (beyond rule of 5). The authors note that "eRo5 space may be thought of as a buffer zone between Ro5 and bRo5 space".  I would challenge this point because there is a region (MW less than 500 Da and ClogP between 5 and 7.5) between Ro5 and bRo5 spaces that is not covered by the eRo5 specifications. As such, it is not meaningful to compare properties of eRo5 compounds with properties of Ro5 or bRo5 compounds. The authors of featured article really do need to fix this problem if they're planning to carve out niche in this area of study because failing to do so will make it easier for conservative drug-likeness 'experts' to challenge their findings. Problems like this are particularly insidious because the activation barriers for fixing them just keep getting higher the longer you ignore them.   

But enough of Bermuda Triangles in the space between Ro5 and bRo5 because the focus of this post is ligand efficiency and specifically its relevance (or otherwise) to bRo5 cmpounds. I'll write a formula for generalized LE is a way that makes it clear that DG° is a function of temperature, pressure and the standard concentration: 


LEgen = -DG°(T,p,C°)/HA

When LE is calculated it is usually assumed that C° is 1 M although there is nothing in the original definition of LE that says this has to be so and few, if any, users of the metric are even aware that they are making the assumption. When analyzing data it is important to be aware of all assumptions that you're making and the effects that making these assumptions may have on the inferences drawn from the analysis.

Sometimes LE is used to specify design guidelines.  For example we might assert that acceptable fragment hits must have LE above a particular cutoff. It's important to remember that setting a cutoff for LE is equivalent to imposing an affinity cutoff that depends on molecular size. I don't see any problem with allowing the affinity cutoff to increase with molecular size (or indeed lipophilicity) although the response of the cutoff to molecular size should reflect analysis of measured data (rather than noisy sermons of self-appointed thought-leaders). When you set a cutoff for LE, you're assuming (whether or not you are aware of it) that the affinity cutoff is a line that intersects the affinity axis at a point corresponding to Kof 1 M. Before heading back to bRo5, I'd like you to consider a question. If you're not comfortable setting an affinity cutoff as a function of molecular size would you be comfortable setting a cutoff for LE?

So let's take a look at what the featured article has to say about affinity: 


"Affinity data were consistent with those previously reported [44] for a large dataset of drugs and drugs in Ro5, eRo5 and bRo5 space had similar means and distributions of affinities (Figure 6a)"

So the article is saying that, on average, bRo5 compounds don't need to be of higher affinity than Ro5 compounds and that's actually useful information. One might hypothesize that unbound concentrations of bRo5 compounds tend to be lower than for Ro5 compounds because the former are less drug-like and precisely the abominations that MAMO (Mothers Against Molecular Obesity) have been trying to warn honest, god-fearing folk about for years. If you look at Figure 6a in the featured article, you'll see that the mean affinity does not differ significantly between the three categories of compound. Regular readers of this blog will be well aware that that categorizing continuous data in this manner tends to exaggerate trends in data. Given that the authors are saying that there isn't a trend, correlation inflation is not an issue here. 

Now look at Figure 6b. The authors note:


"As the drugs in eRo5 and bRo5 space are significantly bigger than Ro5 drugs, i.e., they have higher molecular weights and more heavy atoms, their LE is significantly lower"

If you're thinking about using these results in your own work, you really need to be asking whether or not the results provide any real insight (i.e. something beyond the the trivial result that 1/HA gets smaller when HA gets larger? This would also be a good time to think very carefully about all the assumptions you're going to make in your analysis. The featured article states:

"Ligand efficiency metrics have found widespread use;[45however, they also have some limitations associated with their application, particularly outside traditional Ro5 drug space. [46We nonetheless believe it is useful to characterize the ligand efficiency (LE) and lipophilic ligand efficiency (LLE) distributions observed in eRo5 and bRo5 space to provide guides for those who wish to use them in drug development"

Given that I have asserted that LE is not even wrong and have equated it with homeopathy, I'm not sure that I agree with sweeping LE's probems under the carpet by making a vague reference to "some limitations".  Let's not worry too much about trivial details because declaring a ligand efficiency metric to be useful is a recognized validation tool (even for LELP which appears have jumped straight from the pages of a Mary Shelley novel).  There is a rough analogy with New Math where "the important thing is to understand what you're doing rather to get the right answer" although that analogy shouldn't be taken too far because it's far from clear whether or not LE advocates actually understand what they are doing. As an aside, New Math is what inspired "the rule of 3 is just like the rule of 5 if you're missing two fingers" that I have occasionally used when delivering harangues on fragment screening library design.

So let's see what happens when one tries to set an LE threshold for for bRo5 compounds. The featured article states:

"Instead, the size and flexibility of the ligand and the shape of the target binding site should be taken into account, allowing progression of compounds that may give candidate drugs with ligand efficiencies of ≥0.12 kcal/(mol·HAC), a guideline that captures 90% of current oral drugs and clinical candidates in bRo5 space"

So let's see how this recommended LE threshold of 0.12 kcal/(mol.HA) translates to affinity thresholds for compounds with molecular weights of 700 Da and 3000 Da. I'll assume a temperature of 298 K and C° of 1 M when calculating DG°and will use 14 Da/HA to convert  molecular weight to heavy atoms.  I'll conclude the post by asking you to consider the following two questions? 


  • The recommended LE threshold transforms to pKD threshold of 4.4 at 700 Da. When considering progression of compounds that may give candidate drugs, would you consider a recommendation that KD should be less than 40 mM to be useful?



  • The recommended LE threshold transforms to a pKD threshold of 19 at 3000 Da. How easy do you think it would be to measure a pKD value of 19? When considering progression of compounds that may give candidate drugs, would you consider a recommendation that  pKD be greater than 19 to be useful?

      

Monday, 29 February 2016

The boys who cried wolf

So it's back to blogging and it's taken a bit longer to get into it this year since I had to finish a few things before leaving Brazil. This is a long post so make sure to have some strong coffee to hand.

This post features an article, 'Molecular Property Design: Does Everyone Get It?' by two unwitting 'collaborators' in our correlation inflation Perspective. There are, however, a number of things that the authors of this piece just don't 'get' which makes their choice of title particularly unfortunate.  The first thing that they don't 'get' is that doing questionable data analysis in the past means that people in the present are less likely to heed your warnings about the decline in quality of compounds in today's pipelines. As has been pointed out more than once by this blog, rules/guidelines in drug discovery are typically based on trends observed in measured data and the strength of the trend tells you how rigidly you should adhere to the rule/guideline. Correlation inflation (see also voodoo correlations) is a serious problem in drug discovery because it causes drug discovery scientists to to give more weight to rules/guidelines (and 'expert' opinion) than is justified by the data. In drug discovery, we need to make a distinction between what we believe and what we know. If we can't (or won't) make this distinction then those who fund our activities may conclude that the difficulties that we face are actually of our own making and that's something else that the authors of the featured article just don't seem to 'get'. "Views obtained from senior medicinal chemistry leaders..." does come across as arm-waving and I'm surprised that the editor and reviewers (if there were any) let them get away with it.  

If you're familiar with the correlation inflation problem, you'll know that one of the authors of the featured article did some averaging of groups of data points prior to analysis which was presented in support of an assertion that, "Lipophilicity plays a dominant role in promoting binding to unwanted drug targets". This may indeed be the case but it is not correct to suggest that the analysis supports this opinion because the reported correlations are between promiscuity and median lipophilicity rather than lipophilicity itself. The author concedes that the analysis has been criticized but does not make any attempt to rebut the criticism. Readers can draw their own conclusions from the lack of rebuttal.

The other author of the featured article also 'contributed' to our correlation inflation study although it would be stretching it to term that contribution as 'data analysis'. The approach used there was to first bin the data and then to plot bar charts which were compared visually. You might wonder how a bar chart of binned data can be used to quantify the strength of a  trend and, if attempting to do this, keep your arms loose because you'll be waving them a lot. Here are a couple of examples of how the approach is applied:    

The clearer stepped differentiation within the bands is apparent when log DpH7.4 rather than log P is used, which reflects the considerable contribution of ionization to solubility. 

This graded bar graph (Figure 9) can be compared with that shown in Figure 6b to show an increase in resolution when considering binned SFI versus binned c log DpH7.4 alone. 

This second approach to data 'analysis' is actually more relevant than the first to this blog post because it is used as 'support' ('a crutch' might be a more appropriate term) for SFI (Solubility Forecast Index), which is the old name for PFI (Property Forecast Index) which the featured article touts as a metric. If you're thinking that it's rather strange to 'convert' one form of continuous data (e.g. measured logD) into another form of continuous data (values of metrics) by first making it categorical and turning it into pictures, you might not be alone. What 'senior medicinal chemistry leaders' would make of such data 'analysis' is open to speculation.  

But enough of voodoo correlations and 'pictorial' data analysis because I should make some general comments on property-based design. Here's a figure that provides an admittedly abstract view of property-based design.
 
One challenge for drug-likeness advocates analyzing large, structurally heterogenous data sets is to make the results of analysis relevant to the medicinal chemists working on one or two series in a specific lead optimization project. Affinity (for association with both therapeutic target and antitargets) and free concentration at site of action are the key determinants of drug action. In general, the response of activity to lipophilicity depends on chemotype and, in the case of affinity, also on the relevant protein target (or antitarget). If you're going to tell medicinal chemists how to do their jobs then you can't really afford to have any data-analytic skeletons rattling around in the closet and that's something else that the authors of the featured article just don't 'get'.

The featured article asserts:

The principle of minimal hydrophobicity, proposed by Hansch and colleagues in 1987 states that “without convincing evidence to the contrary, drugs should be made as hydrophilic as possible without loss of efficacy.” This hypothesis is surviving the test of time and has been quantified as lipophilic ligand efficiency (LLE or LipE).

A couple of points need to be made here. Firstly, when Hansch et al refer to 'hydrophobicity', they mean octanol/water logP (as opposed to logD). Secondly, the observation that excessive lipophilicity is a bad thing doesn't actually justify using LLE/LipE in lead optimization. The principle proposed by Hansch et al suggests that a metric of the following functional form may be useful for normalization of activity with respect to liophilicity:


pIC50  -  (l  ´ logP)

However, the principle does not tell us what value of l is most appropriate (or indeed whether a single value of l is appropriate for all situations).  The 'sound and fury' article reviewed in an earlier post makes a similar error with ligand efficiency. 

So it's now time to take a look at PFI and the featured article asserts:


The likelihood of meeting multiple criteria, a typical requirement for a candidate drug, increases substantially with  ‘low fat, low flat’ molecules where PFI is <7, versus >7. In considering a portfolio of drug candidates, the probabilistic argument hypothesizes that successful outcomes will increase as the portfolio’s balance of biological and physicochemical properties becomes more similar to that of marketed drugs.

The first thing that a potential user of PFI should be asking him/herself is where this magic value of 7 comes from since the featured article does imply that the likelihood of good things will increase substantially when PFI is reduced from 7.1 to 6.9. Potential users also need to ask whether this step jump in likelihood is backed by statistical analysis of experimental data or by 'clearer stepped variation' in pictures created using an arbitrary binning scheme. It's also worth remembering that thresholds used to apply guidelines often reflect the binning schemes used to convert continuous data to categorical data and the correlation inflation Perspective discusses the 4/400 rule in this context. Something that molecular property design 'experts' really do need to 'get' is that simple yes/no guidelines are of limited use in practical lead optimization even when these are backed by competent analysis of relevant experimental data. Molecular property 'experts' also need to 'get' that measured lipophilicity is not actually a molecular property. 

PFI is defined as the sum of chromatographic logD (at pH 7.4) and the number of aromatic rings: 


PFI = Chrom logDpH7.4 +  Ar rings

Now suppose that you're a medicinal chemist in a department where the head of medicinal chemistry has decreed that that 80% of compounds synthesized by departmental personnel must have PFI less than 7.  When senior medicinal chemistry leaders set targets like these, the primary objective (i.e topic of your annual review) is to meet them. Delivering clinical candidates is secondary objective since these will surely materialize in the pipeline as if by magic provided that the compound quality targets are met. 

There is a difference between logD (whether measured  by shake-flask or chromatographically) and logP and one which it is important for compound quality advocates to 'get'. When we measure lipophilicity, we determine logD rather than logP and so it is not generally valid to invoke Hansch's principle of minimal hydrophobicity (which is based on logP) when using logD. If the compound in question is not significantly ionized under experimental conditions (pH) then logP and logD will be identical. However, this is not the case when ionization is significant as is usually the case for amines and carboxylic acids at a physiological pH like 7.4. If ionization is significant then logD will typically be lower than logP and we sometimes assume that only the neutral form of the compound partitions into the organic phase for the purposes of prediction or interpretation of log D values. If this is indeed the case we can write logD as a function of logP and the fraction of compound existing in neutral form(s):

  log D(pH) = log P + log Fneut(pH)
Ionized forms can sometimes partition into the organic phase although measuring the extent to which this happens is not easy and the effective partition coefficient for a charged entity depends on whatever counter ion is present (and its concentration). 

So let's get back to the problem of reducing logD so our medicinal chemist can achieve those targets and get an A+ rating in the annual review.  Two easy ways to lower logD are to add ionizable groups (if compound is neutral) and to increase extent of ionization (if compound already has ionizable groups). Increasing the extent of ionization will generally be expected to increase aqueous solubility but I hope readers can see why we wouldn't expect this to help when a compound binds in an ionized form to an antitarget such as hERG (see here for a more complete discussion of this point).  Now I'd like you to take a close look at Figure 2(a) in the featured article. You'll notice that the profiles for the last two entries (hERG and promiscuity) have actually been generated using intrinsic PFI (iPFI) rather than PFI itself and you may be wondering what iPFI is and why it was used instead of PFI. In answer to the first question, iPFI is calculated using logP rather than logD:


 iPFI = logP +  Ar rings

This definition of iPFI is not quite complete because the authors of the featured article don't actually say what they mean by logP.  Is it actually obtained directly from experimental measurements (e..g. logD/pH profile) or is it calculated (in which case it should be stated which method was used for the calculation).

Some medicinal chemists reading this will be asking what iPFI was even doing in the article in the first place and my response would be, as I say frequently in Brazil, 'boa pergunta'.  My guess is that using PFI rather than iPFI for the hERG row of Figure 2(a) would have the effect of shifting the cells in this row one or two cells to the left (based on the assumption that logP will be 1 to 2 units greater than logD at pH 7.4).  Such a shift would make compounds with PFI less than 7 look 'dirtier' than the PFI advocates would like you to think.

There is another term in PFI and that's the number of aromatic rings (# Ar rings) which is meant to measure how 'flat' a molecular structure is.  That it might do but then again it might not because two 'flat' aromatic rings will look a lot less flat when linked by a sulfonyl group and their rigidity could prove to be a liability when trying to pack them into a crystal lattice. However, number of aromatic rings will also quantify molecular size (especially in typical Pharma compound collections) and this is something my friends at Practical Fragments have also noted. Molecular size had been recognized as a pharmaceutical risk factor for at least a decade before people started to tout PFI (or SFI) as a compound quality metric and we can legitimately ask whether or not using a more conventional measure of molecular size (e.g. molecular weight, number of non-hydrogen atoms or molecular volume) would have resulted in a more predictive (or useful) metric.

So let's assume for a moment that you're a medicinal chemist in a place where the 'senior medicinal chemistry leaders' actually believe that optimizing PFI is useful. In case you don't know, jobs for medicinal chemists don't exactly grow on trees these days and so it makes a lot of sense to adopt an appropriately genuflectory attitude to the prevailing 'wisdom' of your 'leaders'. The problem is that your 1 nM enzyme inhibitor with the encouraging pharmacokinetic profile has a PFI of 8 and your lily-livered manager is taking some flak from the Compound Respository Advisory Panel for having permitted you to make it in the first place. Fear not because, you have two benzene rings at the periphery of the molecular structure which will make the synthesis relatively easy. Basically you need to think of a metric like PFI as a Gordian knot that needs to be cut efficiently and you can do this either by eliminating rings or by eliminating aromaticity. Substitution of benzoquinone (either isomer) or cyclopentadiene for the offending benzene rings will have the desired effect.

It's been a long post and I really do need to start wrapping things up. One common reaction when you criticize of drug discovery metrics is the straw man defense in which your criticism is interpreted as an assertion that one doesn't need to worry about physicochemical properties.  In other words, this is precisely the sort of deviant behavior that MAMO (Mothers Against Molecular Obesity) have been trying to warn about. To the straw men, I will say that we described lipophilicity and molecular size as pharmaceutical risk factors in our critique of ligand efficiency metrics. In that critique, we also explain what it means to normalize activity to with respect to risk factor and that's something that not even the NRDD ligand efficiency metric review does. There's a bit more to defining a compound quality metric than dreaming up arbitrary functions of molecular size and lipophilicity and that's something else that the authors of the featured article just don't seem to 'get'. When you use PFI you're assuming that a one unit decrease in chromatographic logD is equivalent to eliminating an aromatic ring (or the aromaticity of a ring) from the molecular structure. 

The essence of my criticism of metrics is that the assumptions encoded by the metrics are rarely (if ever) justified by analysis of relevant measured data. A plot of pIC50 against the relevant property for your project compounds is a good starting point for property-based design and it allows you to use the actual trend observed in your data for normalization of activity values (see the conclusion to our ligand efficiency metric critique for a more detailed discussion of this). If you want to base your decisions on 'clearer stepped differentiation' in pictures or on the blessing of 'senior medicinal chemistry leaders', as a consenting adult, you are free to do so.