Showing posts with label FBLD. Show all posts
Showing posts with label FBLD. Show all posts

Sunday, 24 July 2022

HB Donor Fragment Selection Themes

[This post was updated on 11-Dec-2024 to reflect the publication of the 'HBDs in drug design' preprint as the K2022 article]

In this post I’ll look at a couple of fragment selection themes with a hydrogen bond donor (HBD) focus. The material has been taken from the recent ‘HBDs in drug design’ preprint (HBD3; this would subsequently be published as the K2022 article) which introduced the term ‘hydrogen bond donor-acceptor asymmetry’ and suggested that we need to think differently about HBDs and hydrogen bond acceptors (HBAs) in drug design. One example of these hydrogen bond donor-acceptor asymmetries is that HBAs are typically more strongly solvated than HBDs in aqueous media and this is especially relevant to lead optimization (as shown in the graphical abstract for HBD3 below). 

However, this post is about fragment selection, rather than fixing ADME, and so I’ll say something about differences between HBDs and HBAs in the context of binding to targets. Let’s suppose that you’d like to exploit an HBD in the binding site of your target. All you need to do is place an HBA at a point in space where it can form a good hydrogen bond (taking care to address issues like steric footprint and conformational energy) and you’ve got it sorted. However, life is not quite so simple if you’re trying to exploit an HBA in the binding site because the HBD (e.g., amide NH) that you present to it will almost invariably be accompanied by an HBA (e.g., amide carbonyl O). In contrast, it is relatively easy to design an HBA (e.g., pyridine N) into a ligand structure that is not accompanied by an HBD.   

In HBD3, I describe the HBA that accompanies pretty much every neutral HBD as ‘co-occurring’. The problem for designers is that the co-occurring HBA, which is likely to come with a larger desolvation penalty than that for the HBD, needs to be accommodated and this places constraints on design. It’s also more difficult to achieve ‘line-of-sight’ access with HBDs than is the case for HBAs (you’re likely to need line-of-sight access when targeting a polar atom at the bottom of a relatively narrow binding pocket). The following figure should give you a better idea of what I’m getting at and let’s assume that we’re trying to donate an HB to HBA sitting at the bottom of a narrow and otherwise non-polar binding pocket. Although each of the three structures has appropriate geometry for line-of-sight access, things are not likely to end well if you try to exploit this line-of-sight access in a real-life design situation.

Let’s start with the phenol and, although not pertinent to this discussion, it’s worth mentioning that hydroxyl groups are prone to conjugation in phase 2 metabolism (drugs get hydroxylated in phase 1 metabolism in order to facilitate clearance). Donation of an HB by a ligand hydroxyl to a target HBA also brings the hydroxyl oxygen (the co-occurring HBA) into proximity with the molecular surface of the target. This increases the likelihood of an energetic penalty resulting from desolvation of the phenolic oxygen. One subtle point is that donation of an HB by the phenolic hydroxyl increases the HB basicity of the oxygen which effectively increases the energetic cost of desolvating it.

The co-occuring HBA of the primary amide is an even bigger problem than for phenol because the high polarity of the carbonyl oxygen means that it carries a large desolvation penalty (bad news if you’re trying to hit an HBA at the bottom of a narrow and otherwise non-polar binding pocket). If this is not enough of a problem, you also need to worry about desolvation penalties associated with the second HBD (the primary amide has two HBDs and methyl-capping will take out the one that you need for hitting that HBA at the bottom of the binding pocket). As Lady Bracknell might have observed, “One desolvated polar atom may be regarded as a misfortune; to lose solvation of two polar atoms looks like carelessness”.

The last of the trio of structures is pyrazole linked at C4 and this avoids problems that might result from biasing the tautomeric preference. Pyrazole is a great warhead if you’re targeting a proximal HBD and HBA (as is the case when trying to hit a kinase hinge). However, pyrazole’s HBA may become a liability when trying to hit the HBA at the bottom of that otherwise non-polar binding pocket. Why not just take out pyrazole’s HBA, you might ask? The problem is that pyrroles are very electron rich and tend to be quite reactive.  One tactic is to move the co-occurring HBA from the ring to the linker (1 and 2) in a way that makes the linker electron-withdrawing and pray for a less destabilizing contact between the co-occurring HBA and the binding site. Alternatively, you can take out the co-occurring HBA and modify the linker to make it more electron-withdrawing (3). I’ve included Hammett σ values in the graphic and these will give you an idea how the substituents vary in their ability to suck electron density out of the pyrrole ring (beneficial both for making the pyrrole ring more rugged and increasing the HB acidity of its NH HBD).  I see these fragments as being of about the right size to be screened crystallographically but you might want something a bit larger than methyl if you’re using another detection method.

If you’re designing (or trying to improve the coverage of) a fragment library then another selection theme that you might want to think about is fragments that can present a high ‘density’ of HBDs to a target while minimizing the number of co-occurring HBAs. One way to do this is to use the guanidine substructure although this will cause some medicinal chemists to roll their eyes (concerns about permeability) while Ro3’s adherents would be likely to denounce you for heresy (actually not such a bad thing and I think that the late, great Denis Healey might have likened this to “being savaged by a dead sheep”). Guanidine itself is extremely basic (pKa = 13.6 | ref) which means very little of the neutral form for diffusing across membranes. However, the pKa of guanidine is also extremely sensitive to substitution and a number of approved drugs incorporate this substructure. I should also point out that, even in the neutral form of guanidine, the amide-like nitrogen atoms do not function as HBAs (even though they’d be counted as such when applying Ro5).

I’ve made a small selection of substituted guanidines that I think may of interest for screening as fragments. The pKa values that I quote in this post are from an article by two former colleagues (Peter Taylor and Alan Wait who are sadly both deceased) and this is an excellent source of measured guanidine pKa values.   Two of these (4 and 5) will be predominantly protonated at neutral pH although there’ll still be a significant amount of the neutral form that you’ll need for permeability. The other two guanidines will be predominantly neutral at neutral pH although 6 is sufficiently basic to protonate in lysosomes.  As for the pyrroles, I see these as about the right size to be screened crystallographically but you might want something a bit bigger than methyl if you plan to use a different detection method.


  


Friday, 1 April 2022

Enthalpic fragments

<< previous || next >>

Enthalpy-driven binding has been presented as a rationale for screening fragments although some have argued that thermodynamic signature is actually a 'red herring' in the context of drug discovery.  Binding of a ligand grown from a fragment hit incurs a translational entropy penalty that is similar to that of the original fragment hit and it is therefore it is hardly surprising that synthetic elaboration results in binding that is more driven by entropy.

A recent collaborative study between researchers in the Budapest Enthalpomics Group (BEG) and Prof Wilhelmina Wiplasch, well known for her seminal study ‘The Ecstasy and Agony of Recreational PAINS’, shows this view to be hopelessly naïve. The mathematical treatment used in the study is formidable and was originally developed by Prof Wiplasch during a sabbatical at the Port-au-Prince Institute of Biogerontology. Briefly, deep learning was used to model the time-dependent covariance and kurtosis of the polarizability tensor for a series of rhodanines, showing that the enthalpic nature of fragment binding is caused by their greater ligand efficiencies. “This model comprehensively outperforms all competitors”, explains Group Leader Prof Kígyó Olaj, “and we have shown for the very first time that the Sackur-Tetrode equation can be safely consigned to the dustbin of History”.

Tuesday, 12 January 2021

Tom Lehrer's guide to design of SARS-CoV-3 main protease inhibitors for treatment of COVID-32

<< previous || next >>

It’s been ages since my last COVID-19 post (How not to repurpose a 'drug') and I’ll kick blogging off for 2021 with a follow up to an even older post (SARS-CoV-2 main protease. Crowdsourcing, peptidomimetics and fragments). I consider it unlikely that a SARS-CoV-2 main protease inhibitor, designed from scratch, will be available in time to have real impact on the current pandemic (in saying this, I’m making the huge assumption that defeat does not get snatched from the jaws of victory on the vaccination front). While many grinning Lean Six Sigma ‘belts’ (and their synchronously smiling allies in Human Resources) would denounce this as negative and defeatist, what I’m really getting at is that we need to think about targeting SARS-CoV-3 main protease when designing inhibitors for SARS-CoV-2 main protease. As Tom Lehrer advises in the intro to So Long, Mom, “If any songs are going to come from World War III, we better start writing them now”.

Happy New Year (this orchid opened during night of Dec 31/Jan 1)

If we’re designing a SARS-CoV-2 main protease inhibitor to also hit SARS-CoV-3 main protease then it’d be a good idea to engineer it to have greater affinity than necessary for the current target. In the fourth of his rules for air fighting, ‘Sailor’ Malan (readers may also be interested in his insights into fragment screening library design) asserts that “height gives you the initiative” which can be adapted for drug design as “affinity gives you the initiative”.  We should anticipate that inhibitors optimized against the current target will have lower affinity for the future target(s) although it’s obviously not a problem if this proves not to be the case. In any case, high affinity allows you to use a lower dose and that’s an important consideration if you’re planning for healthy people such as nurses and doctors to take the drug prophylactically in order to remain healthy. For a SARS-CoV-2 main protease inhibitor, I’d be looking at a target affinity of 1 nM (or better) which I believe would be achievable without causing too many self-appointed arbiters of 'compound quality' to spit feathers. Pfizer began a phase I study of the SARS-CoV-2 main protease inhibitor PF-00835321 (Ki = 0.27 nM; dosed intravenously as the phosphate pro-drug PF-07304814) in September 2020 although this compound had actually come from a discontinued SARS-CoV project. 

If we want to maximize the chances that a SARS-CoV-2 main protease inhibitor will exhibit comparative affinity for SARS-CoV-3 (or even SARS-CoV-4) then we need to exploit protein structural features that are likely to be conserved between the different main proteases. This points to milking as much activity as possible out of the core substructure of the inhibitor as a design strategy. With this in mind, I suggest that we really do need to exploit the catalytic cysteine if we’re serious about treating COVID-32 (or worried about SARS-CoV-2 main protease mutations). 

In drug design, we typically exploit a catalytic cysteine by forming a covalent bond between the thiol sulfur and an electrophilic atom in the molecular structure of the inhibitor (PF-00835321 uses the carbon of a carbonyl group to engage the catalytic cysteine). The functional group containing the electrophilic atom is commonly referred to as a “warhead” and covalent bond formation between cysteine can either be reversible or irreversible. Geometric constraints associated with covalent bond formation are typically a lot more stringent than for hydrogen bonds and you’ll make life much easier for yourself by getting the warhead into structures as early as possible in hit-to-lead. I generally recommend using reversible warheads in design of cysteine protease inhibitors (PF-00835321 binds reversibly to SARS-CoV-2) and present my reasoning in this document. In essence, irreversible inhibition adds complexity to design (both Ki and kinact need to be controlled) while placing greater technical demands on the design team (e.g. for generation of the structural models for transition states required for structure-based design).

The argument typically presented in support of irreversible inhibition (and slow binding kinetics) is that it leads to longer duration of action. This argument emphasizes benefits of slow (or zero) off-rate during the elimination phase while ignoring disadvantages of slow on-rate during the distribution phase and I’ll point you to an insightful article by my former colleague Rutger Folmer. While there will be situations in which irreversible inhibition really is the best option, the decision as to whether to go for reversible or irreversible inhibitors is one that should be carefully considered at the start of the project. In drug discovery, it usually ends in tears once the tail starts wagging the dog as would be the case if choice of screening tactics (covalent fragment screening typically finds irreversible binders) were to dictate lead optimization strategy. In particular, I wouldn't really recommend the laissez faire approach to project management (“once the rockets are up, who cares where they come down”) chronicled by Tom Lehrer.

Here's some information that may be of interest if you're selecting or designing warheads to form covalent bonds with catalytic cysteines. First, a couple of comparative studies of reversible and irreversible warheads. Second, some papain inhibition data taken from the literature ( B1977 | W1972 | L1971 ), summarized in the graphic below, that are relevant to fragment library design.

Off-target activity is always a concern in drug design since this can cause toxicity (it’s often considered politer to say “adverse drug reaction” rather than use the uncouth T-word although Tom Lehrer provides a useful perspective) and that’s a strong rationale for trying to achieve a low therapeutic dose. It’s my understanding (still wading through literature) that SARS-CoV-2 main protease functions in the endoplasmic reticulum which means that the relevant physiological pH is close to neutral. Many proteases (potential anti-targets for SARS-CoV-2 main protease inhibitors) function in acidic compartments such as lysosomes and the presence of a basic center in the molecular structure of an inhibitor will tend to draw it into these acidic compartments. When designing SARS-CoV-2 inhibitors, the safest option is simply to avoid basic centers (see F2005) . In particular, to link a ‘gratuitous’ basic center and an irreversible warhead would be to tempt launchpad misadventure.

I'll conclude the post with an observation that the COVID-19 pandemic seems to have triggered a parallel pandemic in scholarly publishing which is forcing scientists to be more creative in finding new ways to getting their messages to stand out. I'll let Tom Lehrer have the last word.   

Saturday, 18 July 2020

SARS-CoV-2 main protease. Crowdsourcing, peptidomimetics and fragments

<< previous || next >>

“Just take the ball and throw it where you want to. Throw strikes. Home plate don’t move.”

Satchel Paige (1906-1982) 

The COVID Moonshot and OSC19 are examples of what are sometimes called crowdsourced or open source approaches to drug discovery. While I’m not particularly keen on the use of the term ‘open source’ in this context, I have absolutely no quibble with the goal of seeking cures and treatments for diseases that are ignored by commercial drug discovery organizations. Open source drug discovery originated with OSDD in India and it should be noted that the approach has also been pioneered for malaria by OSM.  I see crowdsourcing primarily as a different way to organize and resource drug discovery rather than as a radically different way to do drug discovery.

One point that’s not always appreciated by cheminformaticians, computational chemists and drug discovery scientists in academia is that there’s a bit more to drug discovery than making predictions. In particular, I advise those seeking to transform drug discovery to ensure that they actually know what a drug needs to do and understand the constraints under which drug discovery scientists work. Currently, it does not appear to be possible to predict the effects of compounds in live humans from molecular structure with the accuracy needed for prediction-driven design and this is the primary reason that drug discovery is incremental in nature. A big part of drug discovery is generation of the information needed in order to maintain progress and there are gains to be had by doing this as efficiently as possible. Efficient generation of information, in turn, requires a degree of coordination that may prove difficult to achieve in a crowdsourced project.

The SARS-CoV-2 main protease (Mpro) is one of a number of potential targets of interest in the search for COVID-19 therapies. Like the cathepsins that are (or, at least, have been) of interest to the pharma/biotech industry as potential targets for therapeutic intervention, Mpro is a cysteine protease. If I’d been charged with quickly delivering an inhibitor of Mpro as a candidate drug then I’d be taking a very close look at how the pharma/biotech industry has pursued cysteine protease targets. Balacatib, odanacatib (cathepsin K inhibitors) and petesicatib (cathepsin S inhibitor) can each be described as a peptidomimetic with a warhead (nitrile) that forms a covalent bond reversibly with the catalytic cysteine.

A number of peptidomimetic Mpro inhibitors have been described in the literature and this blog post by Chris Southan may be of interest. I’ve been looking at the published inhibitors shown below in Chart 1 (which exhibit antiviral activity and have been subjected to pharmacokinetic and toxicological evaluation) and have written some notes on mapping the structure-activity relationship for compounds like these. I should stress that compounds discussed in these notes are not expected to be dramatically more potent than the two shown in Chart 1 (in fact, I expect at least one to be significantly less potent). Nevertheless, I would argue that assay results for these proposed synthetic targets would inform design.

My assessment of these compounds is that there is significant room for improvement and I think that it would be relatively easy to achieve a pIC50 of 8 (corresponding to an IC50 of 10 nM) using the aldehyde warhead. I’d consider taking an aldehyde forward (there are options for dosing as a prodrug) although it really would be much better if there was also the option to exchange this warhead for the nitrile (a warhead that is much-loved by industrial medicinal chemists since it’s rugged, polar and contributes minimally to molecular size). While I’d anticipate that replacement of aldehyde with nitrile will lead to a reduction in potency, it’s necessary to quantify the potency loss to enable the potential of nitriles to be properly assessed. The binding mode observed for 1 is shown below in Figure 1 and it’s likely that the groove region will need to be more fully exploited (this article will give you an idea of the sort of thing I have in mind) in order to achieve acceptable potency if the aldehyde warhead is replaced by nitrile.

The COVID Moonshot project currently appears to be in what many industrial drug discovery scientists would call the hit-to-lead phase.  In my view the principal objective of hit-to-lead work is to create options since having options will give the lead optimization team room to manoeuvre (you can think of hit-to-lead work as being a bit like playing in midfield). The COVID Moonshot project is currently focused on exploitation of hits from a fragment screen against MPro and, while I’d question whether this approach is likely to get to a candidate drug more quickly than the conventional structure-based design used in industry to pursue cathepsins, it’s certainly an interesting project that I’m happy to contribute to. It’s also worth mentioning that fragment screens have been run against SARS-CoV-2 Nsp3 macrodomain at UCSF and Diamond since there are no known inhibitors for this target.

Here’s a blog post by Pat Walters in which he examines the structure-activity relationships emerging for the fragment-derived inhibitors. Specifically, he uses a metric known as the Structure-Activity Landscape Index (SALI) to quantify the sensitivity of activity to structural changes. Medicinal chemists apply the term ‘activity cliff’ to situations where a small change in structure results in a large change in activity and I’ve argued that the idea of quantifying the sensitivity of a physicochemical effect to structural modifications goes all the way back to Hammett.  One point that comes out of Pat’s post is that it’s difficult to establish structure-activity relationships for low affinity ligands with a conventional biochemical assay. When applying fragment-based approaches in lead discovery, there are distinct advantages to being able to measure low binding affinity (~ 1 mM) since this allows fragment-based structure-activity relationships to be explored prior to synthetic elaboration of fragment hits. As Pat notes, inadequate solubility in assay buffer clearly places limits on the affinity that can be reliably measured in any assay although interference with the readout of a biochemical assay can also lead to misleading results. This is one reason that biophysical detection of binding using methods such as surface plasmon resonance (SPR) are favored in fragment-based lead discovery. Here’s an article by some of my former colleagues which shows how you can assess the impact of interference with the readout of a biochemical assay (and even correct for it if the effect isn’t too great).     

My first contribution to the COVID Moonshot project is illustrated in Chart 2 and the fragment-derived inhibitor 3 from which I started is also featured in Pat’s post. From inspection of the crystal structure, I noticed that the catalytic cysteine might be targeted by linking a ‘reversible’ warhead from the amide nitrogen (4 and 5). Although this might look fine on paper, the experimental data in this article suggest that linking any saturated carbon to the amide nitrogen will bias the preferred amide geometry away from trans to cis. Provided that the intrinsic gain in affinity resulting from linking the warhead is greater than the cost of adopting the bound conformation, the structural modification will lead to a net increase in affinity and the structures could be locked (here's an article that shows how this can work) into the bound conformation (e.g. by forming a ring).


In addition to being accessible to a warhead linked from the amide nitrogen of 3, the catalytic cysteine is also within striking distance of the carbonyl carbon and it would be prudent to consider the possibility that 3 and its analogs can function as substrates for Mpro. There is precedent for this type of behavior and I’ll point you toward an article that notes that a series of esters identified as cruzain inhibitors can function as substrates and more recent article that presents cruzain inhibitors that I’d consider to be potential substrates. A crystal structure of the protein-ligand complex is potentially misleading in this context since the enzyme might not be catalytically active. I believe that 6 could be used to explore this possibility since the carbonyl carbon would be expected to be more electrophilic and 3-hydroxy, 4-methylpyridine would be expected to be a better leaving group than its 3-amino analog.

This is a good point to wrap things up. I think that Satchel Paige gave us some pretty good advice on how to approach drug discovery and that's yet another reason that Black Lives Matter.

Friday, 30 November 2018

Ligand efficiency and fragment-to-lead optimizations


The third annual survey (F2L2017) of fragment-to-lead (F2L) optimizations was published last week. Given that it was the second survey (F2L2016) in this series, that prompted me to write 'The Nature of Ligand Efficiency' (NoLE), I thought that some comments would be in order. F2L2017 presents analysis of data that had been aggregated from all three surveys and I'll be focusing on the aspects of this analysis that relate to ligand efficiency (LE).

As noted in NoLE, perception of efficiency changes when affinity is expressed in different concentration units and I have argued that this is an undesirable feature for a quantity that is widely touted as useful for design. At very least, it does place a burden of proof on those who advocate the use of LE in design to either show that the change in perception of efficiency with concentration unit is not a problem or to justify their choice of the 1 M concentration unit. One difficulty that LE advocates face is that the nontrivial dependency of LE on the concentration unit only came to light a few years after LE was introduced as "a useful metric for lead selection" and, even now, some LE advocates appear to be in a state of denial. Put more bluntly, you weren't even aware that you were choosing the 1 M concentration unit when you started telling medicinal chemists that they should be using LE to do their jobs but you still want us to believe that you made the correct choice?

I'm assuming that the authors of F2L2017 would all claim familiarity with the fundamentals of physical chemistry and biophysics while some of the authors may even consider themselves to be experts in these areas. I'll put the following question to each of the authors of F2L2017: what would your reaction be to analysis showing that the space group for a crystal structure changed if the unit cell parameters were expressed using different units? I can also put things a bit more coarsely by noting that to examine the effect on perception of changing a unit is, when applicable, a most efficacious bullshit detector.

The analysis in F2L2017 that I'll focus on is the the comparison between fragment hits and leads. As I showed in NoLE, it is meaningless to compare LE values because LE has a nontrivial dependency on the concentration unit used to express affinity. LE advocates can of course declare themselves to be Experts (or even Thought Leaders) and invoke morality in support of their choice of the 1 M concentration unit. However, this is a risky tactic because physical science can't accommodate 'privileged' units and an insistence that quantities have to be expressed in specific units might be taken as evidence that one is not actually an Expert (at least not in physical science).

So let's take a look at what F2L2017 has to say about LE in the context of F2L optimizations.

"The distributions for fragment and lead LE have also remained reasonably constant. On average there is no significant change in LE between fragment and lead (ΔLE = 0.004, p ≈ 0.8). Figure 5A shows the distribution of ΔLE, which is approximately centered around zero, although interestingly there are more examples where LE increases from fragment to lead (40) than where a decrease is seen (25). Some caution is warranted when interpreting these data, as our minimum criterion for 100-fold potency improvement may have introduced some selection bias. Nevertheless, there is no clear evidence in this data set that LE changes systematically during fragment optimization. Although the average change in LE from fragment to lead is small, Figure 5B shows that the correlation between fragment and lead LE is modest (R2 = 0.22), with a mean absolute difference between fragment and lead LE of 0.08."

This might be a good point at which to remind the authors of F2L2017 about some of the more extravagant claims that have been made for LE. It has been asserted that “fragment hits typically possess high ‘ligand efficiency’ (binding affinity per heavy atom) and so are highly suitable for optimization into clinical candidates with good drug-like properties”.  It has also been claimed that "ligand efficiency validated fragment-based design".  However, the more important point is that it is completely meaningless to compare values of LE of hits and leads because you will come to different conclusions if you express affinity using a different concentration unit (see Table 2 in NoLE). It is also worth noting that expressing affinity in units of 1 M introduces selection bias just as does the requirement for 100-fold potency improvement. 

Had I been reviewing F2L2017, I'd have suggested that the authors might think a bit more carefully about exactly why they are analyzing differences between LE values for fragments and leads. A perspective on fragment library design (reviewed in this post) correctly stated that a general objective of optimization projects is “ensuring that any additional molecular weight and lipophilicity also produces an acceptable increase in affinity". If you're thinking along these lines then scaling the F2L potency increase by the corresponding increase in molecular size makes a lot more sense than comparing LE for the fragments and leads. This quantifies how efficiently (in terms of increased molecular size) the potency gains for the F2L project have been achieved. This is not a new idea and I'll direct readers toward a 2006 study in which it was noted that a tenfold increase in affinity corresponded to a mean increase in molecular weight of 64 Da (standard deviation = 18 Da) for 73 compound pairs from FBLD projects. This is how group efficiency (GE) works and I draw the attention of the two F2L2017 authors from Astex to a perceptive statement made by their colleagues that GE is “a more sensitive metric to define the quality of an added group than a comparison of the LE of the parent and newly formed compounds”.

The distinction between a difference in LE and a difference in affinity that has been scaled by a difference in molecular size becomes a whole lot clearer if you examine the relevant equations. Equation (1) defines the F2L LE difference and first thing that you'll notice is that is that it is algebraically more complex than equation (2). This is relevant because LE advocates often tout the simplicity of the LE metric. However, the more significant difference between the two is that the concentration that defines the standard state is present in equation (1) but absent in equation (2). This means that you get the same answer when you scale affinity difference by the corresponding molecular size difference regardless of the units in which you express affinity.


So let's see how things look if you're prepared to think beyond LE when assessing F2L optimizations. Here's a figure from NoLE in which I've plotted the change in affinity against the change in number of non-hydrogen atoms for the F2L optimizations surveyed in F2L2016. The molecular size efficiency for each optimization can be calculated by dividing the change in affinity by the change in in number of non-hydrogen atoms. I've drawn lines corresponding to minimum and maximum values of molecular size efficiency and have also shown the quartiles.

So now it's time to wrap things up. A physical quantity that is expressed in a different unit is still the same physical quantity and I presume that all the authors of F2L2017 would have been aware of this while they were still undergraduates. LE was described as thermodynamically indefensible in comments on Derek's post on NoLE and choosing to defend an indefensible position usually ends in tears (just as it did for the French at Dien Bien Phu in 1954). The dilemma facing those who seek to lead opinion in FBDD is that to embrace the view that the 1 M concentration unit is somehow privileged requires that they abandon fundamental physicochemical principles that they would have learned as undergraduates.   

Sunday, 22 May 2016

Sailor Malan's guide to fragment screening library design


Today I'll take a look at a JMC Perspective on design principles for fragment libraries that is intended to provide advice for academics. When selecting compounds to be assayed the general process typically consists of two steps. First, you identify regions of chemical space that you hope will be relevant and then you sample these regions. This applies whether you're designing a fragment library, performing a virtual screen or selecting analogs of active compounds with which to develop structure-activity relationships (SAR). Design of compound libraries for fragment screening has actually been discussed extensively in the literature and the following selection of articles, some of which are devoted to the topic, may be useful: Fejzo (1999), Baurin (2004), Mercier (2005), Schuffenhauer (2005), Albert (2007) Blomberg (2009), Chen (2009), Law (2009), Lau (2011), Schulz (2011); Morley (2013). This series of blog posts ( 1 | 2 | 3 | 4) on fragment screening library design that may also be helpful.

The Perspective opens with the following quote:

"Rules are for the obedience of fools and the guidance of wise men"

Harry Day, Royal Air Force (1898-1977)


It wasn't exactly clear what the authors are getting at here since there appears to be no provision for wise women. Also it is not clear how the authors would view rules that required darker complexioned individuals to sit at the backs of buses (or that swarthy economists should not solve differential equations on planes). That said, the quote hands me a legitimate excuse to link Malan's Ten Rules for Air Fighting and I will demonstrate that the authors of this Perspective can learn much from the wise teachings of 'Sailor' Malan.

My first criticism of this Perspective is that the authors devote an inordinate amount of space to topics that are irrelevant from the viewpoint of selecting compounds for fragment screening. Whatever your views on the value of ligand efficiency metrics and thermodynamic signatures, these are things that you think about once you've got the screening results. The authors assert, "As a result, fragment hits form high-quality interactions with the target, usually a protein, despite being weak in potency" and some readers might consider the 'concept' of high-quality interactions to be pseudoscientific psychobabble on par with homeopathy, chemical-free food and the wrong type of snow. That said, discussion of some of these peripheral topics would have been more acceptable if the authors had articulated the library design problem clearly and discussed the most relevant literature early on. By straying from their stated objective, the authors have broken the second of Malan's rules ("Whilst shooting think of nothing else, brace the whole of your body: have both hands on the stick: concentrate on your ring sight").


The section on design principles for fragment libraries opens with a slightly gushing account of the Rule of 3 (Ro3). This is unfortunate because this would have been the best place for the authors to define the fragment library design problem and review the extensive literature on the subject. Ro3 was originally stated in a short communication and the analysis that forms its basis is not shared. As an aside, you need to be wary of rules like these because the cutoffs and thresholds may have been imposed arbitrarily by those analyzing the data. For example, the GSK 4/400 rule actually reflects the scheme used to categorize continuous data and it could just have easily been the GSK 3.75/412 rule if the data had been pre-processed differently. I have written a couple ( 1 | 2 ) of blog posts on Ro3 but I'll comment here so as to keep this post as self-contained as possible. In my view, Ro3 is a crude attempt to appeal to the herding instinct of drug discovery scientists by milking a sacred cow (Ro5). The uncertainties in hydrogen bond acceptor definitions and logP prediction algorithms mean that nobody knows exactly how others have applied Ro3. It also is somewhat ironic that the first article referenced by this Perspective actually states Ro3 incorrectly. If we assume that Ro5 hydrogen bond acceptor definitions are being used then Ro3 would appear to be an excellent way to ensure that potentially interesting acidic species such as tetrazoles and acylsulfonamides are excluded from fragment screening libraries. While this might not be too much of an issue if identification of adenine mimics is your principal raison d'etre, some researchers may wish to take a broader view of the scope of FBDD. It is even possible that rigid adherence to Ro3 may have led to the fragment starting points for this project being discovered in Gothenburg rather than Cambridge. Although it is difficult to make an objective assessment of the impact of Ro3 on industrial FBDD, its publication did prove to be manna from heaven for vendors of compounds who could now flog milligram quantities of samples that had previously been gathering dust in stock rooms.



This is a good point to see what 'Sailor' Malan might have made of this article. While dropping Ro3 propaganda leaflets, you broke rule 7 (Never fly straight and level for more than 30 seconds in the combat area) and provided an easy opportunity for an opponent to validate rule 10 (Go in quickly - Punch hard - Get out). Faster than you can say "thought leader" you've been bounced by an Me 109 flying out of the sun. A short, accurate (and ligand-efficient) burst leaves you pondering the lipophilicity of the mixture of glycol and oil that now obscures your windscreen. The good news is that you have been bettered by a top ace whose h index is quite a bit higher than yours. The bad news is that your cockpit canopy is stuck. "Spring chicken to shitehawk in one easy lesson."

Of course, there's a lot more to fragment screening library design than counting hydrogen bonding groups and setting cutoffs for molecular weight and predicted logP. Molecular complexity is one of the most important considerations when selecting compounds (fragments or otherwise) and anybody even contemplating compound library design needs to understand the model introduced by Hann and colleagues. This molecular complexity model is conceptually very important but it is not really a practical tool for selecting compounds. However, there are other ways to define molecular complexity in ways that allow the general concept to be distilled into usable compound selection criteria. For example, I've used restriction of extent of substitution (as detailed in this article) to control complexity and this can be achieved using SMARTS notation to impose substructural requirements. The thinking here is actually very close to the philosophy behind 'needle screening' which was first described in 2000 by researchers at Roche although they didn't actually use the term 'molecular complexity'.


As one would expect, the purging of unwholesome compounds such as PAINS is discussed. The PAINS field suffers from ambiguity, extrapolation and convolution of fact with opinion. This series ( 1 | 2 | 3 | 4) of blog posts will give you a better idea of my concerns. I say "ambiguity" because it's really difficult to know whether the basis for labeling a compound as a PAIN (or should that be a PAINS) is experimental observation, model-based prediction or opinion. I say "extrapolation" because the original PAINS study equates PAIN with frequent-hitter behavior in a panel of six AlphaScreen assays and this is extrapolated to pan-assay (which many would take to mean different types of assays) interference. There also seems to be a tendency to extrapolate the frequent-hitter behavior in the AlphaScreen panel to reactivity with protein although I am not aware that any of the compounds identified as PAINS in the original study were shown to react with any of the proteins in the AlphaScreen panel used in that study. This is a good point to include a graphic to break the text up a bit and, given an underlying theme of this post, I'll use this picture of a diving Stuka.



One view of the fragment screening mission is that we are trying to present diverse molecular recognition elements to targets of interest. In the context of screening library design, we tend to think of molecular recognition in terms of pharmacophores, shapes and scaffolds. Although you do need to keep lipophilicity and molecular size under tight control, the case can be made for including compounds that would usually be considered to be beyond norms of molecular good taste. In a fragment screening situation I would typically want to be in a position to present molecular recognition elements like naphthalene, biphenyl, adamantane and (especially after my time at CSIRO) cubane to target proteins. Keeping an eye on both molecular complexity and aqueous solubility, I'd select compounds with a single (probably cationic) substituent and I'd not let rules get in the way of molecular recognition criteria. In some ways compound selections like those above can be seen as compliance with Rule 8 (When diving to attack always leave a proportion of your formation above to act as top guard). However, I need to say something about sampling chemical space in order to make that connection a bit clearer.

This is a good point for another graphic and it's fair to say that the Stuka and the B-52 differed somewhat in their approaches to target engagement. The B-52 below is not in the best state of repair and, given that I took the photo in Hanoi, this is perhaps not totally surprising. The key to library design is coverage and former bombardier Joseph Heller makes an insightful comment on this topic. One wonders what First Lieutenant Minderbinder would have made of the licensing deals and mergers that make the pharma/biotech industry such an exciting place to work.  


The following graphic, pulled from an old post, illustrates coverage (and diversity) from the perspective of somebody designing a screening library.  Although I've shown the compounds in a 2 dimensional space, sampling is often done using molecular similarity which we can think of inversely related to distance. A high degree of molecular similarity between two compounds indicates that their molecular structures are nearby in chemical space.  This is a distance-geometric view of chemical space in which we know the relative positions of molecular structures but not where they are.  When we describe a selection of molecular structures as diverse, we're saying that the two most similar ones are relatively distant from each other. The primary objective of screening library design is to cover relevant chemical space as effectively as possible and devil is in the details like 'relevant' and 'effectively'. The stars in the graphic below show molecular structures that have been selected to cover the chemical space shown. When representing a number of molecular structures by a single molecular structure it is important, as it is in politics, that what is representative not be too distant from what is being represented. You might ask, "how far is acceptable?" and my response would be, as it often is in Brazil, "boa pergunta". One problem is scaffolds differ in their 'contributions' to molecular similarity and activity cliffs usually provide a welcome antidote to the hubris of the library designer.         


I would argue that property distributions are more important than cutoff values for properties and it is during the sampling phase of library design that these distributions are shaped. One way of controlling distributions is to first define regions of chemical space using progressively less restrictive selection criteria and then sample these in order, starting with the most restrictively defined region. However, this is not the only way to sample and might also try to weight fragment selection using desirability functions. Obviously, I'm not going to provide a comprehensive review of chemical space sampling in a couple of paragraphs of a blog post but I hope to have shown that the sampling of chemical space is an important aspect of fragment screening library design. I also hope to have shown that failing to address the issue of sampling relevant chemical space represents a serious deficiency of the featured Perspective

The Perspective concludes with a number of recommendations and I'll conclude the post with comments on some of these. I wouldn't have too much of a problem with the proposed 9 - 16 heavy atom range as a guideline although I would consider a requirement that predicted octanol/water logP be in the range 0.0 - 2.0 to be overly restrictive. It would have been useful for the authors to say how they arrived at these figures and I invite all of them to think very carefully about exactly what they mean by "cLogP" and "freely rotatable bonds" so we don't have a repeat of the Ro3 farce. There are many devils in the details of the statement:"avoid compounds/functional groups known to be associated with high reactivity, aggregation in solution, or false positives". My response to "known" is that it is not always easy to distinguish knowledge from opinion and "associated" (like correlated) is not a simple yes/no thing. It is not cleat how "synthetically accessible vectors for fragment growth" should be defined since there is also a conformational stability issue if bonds to hydrogen are regarded as growth vectors.   

This is a good point at which to wrap things up and I'd like to share some more of Sailor Malan's wisdom before I go. The first rule (Wait until you see the whites of his eyes. Fire short bursts of 1 to 2 seconds and only when your sights are definitely 'ON') is my personal favorite and it provides excellent, practical advice for anybody reviewing the scientific literature. I'll leave you with a short video in which a pre-Jackal Edward Fox displays marksmanship and escaping skills that would have served him well in the later movie. At the start of the video, the chemists and biologists have been bickering (of course, this never really happens in real life) and the VP for biophysics is trying to get them to toe the line. Then one of the biologists asks the VP for biophysics if they can do some phenotypic screening and you'll need to watch the video to see what happens next...