Showing posts with label Hammett. Show all posts
Showing posts with label Hammett. Show all posts

Saturday, 18 July 2020

SARS-CoV-2 main protease. Crowdsourcing, peptidomimetics and fragments

<< previous || next >>

“Just take the ball and throw it where you want to. Throw strikes. Home plate don’t move.”

Satchel Paige (1906-1982) 

The COVID Moonshot and OSC19 are examples of what are sometimes called crowdsourced or open source approaches to drug discovery. While I’m not particularly keen on the use of the term ‘open source’ in this context, I have absolutely no quibble with the goal of seeking cures and treatments for diseases that are ignored by commercial drug discovery organizations. Open source drug discovery originated with OSDD in India and it should be noted that the approach has also been pioneered for malaria by OSM.  I see crowdsourcing primarily as a different way to organize and resource drug discovery rather than as a radically different way to do drug discovery.

One point that’s not always appreciated by cheminformaticians, computational chemists and drug discovery scientists in academia is that there’s a bit more to drug discovery than making predictions. In particular, I advise those seeking to transform drug discovery to ensure that they actually know what a drug needs to do and understand the constraints under which drug discovery scientists work. Currently, it does not appear to be possible to predict the effects of compounds in live humans from molecular structure with the accuracy needed for prediction-driven design and this is the primary reason that drug discovery is incremental in nature. A big part of drug discovery is generation of the information needed in order to maintain progress and there are gains to be had by doing this as efficiently as possible. Efficient generation of information, in turn, requires a degree of coordination that may prove difficult to achieve in a crowdsourced project.

The SARS-CoV-2 main protease (Mpro) is one of a number of potential targets of interest in the search for COVID-19 therapies. Like the cathepsins that are (or, at least, have been) of interest to the pharma/biotech industry as potential targets for therapeutic intervention, Mpro is a cysteine protease. If I’d been charged with quickly delivering an inhibitor of Mpro as a candidate drug then I’d be taking a very close look at how the pharma/biotech industry has pursued cysteine protease targets. Balacatib, odanacatib (cathepsin K inhibitors) and petesicatib (cathepsin S inhibitor) can each be described as a peptidomimetic with a warhead (nitrile) that forms a covalent bond reversibly with the catalytic cysteine.

A number of peptidomimetic Mpro inhibitors have been described in the literature and this blog post by Chris Southan may be of interest. I’ve been looking at the published inhibitors shown below in Chart 1 (which exhibit antiviral activity and have been subjected to pharmacokinetic and toxicological evaluation) and have written some notes on mapping the structure-activity relationship for compounds like these. I should stress that compounds discussed in these notes are not expected to be dramatically more potent than the two shown in Chart 1 (in fact, I expect at least one to be significantly less potent). Nevertheless, I would argue that assay results for these proposed synthetic targets would inform design.

My assessment of these compounds is that there is significant room for improvement and I think that it would be relatively easy to achieve a pIC50 of 8 (corresponding to an IC50 of 10 nM) using the aldehyde warhead. I’d consider taking an aldehyde forward (there are options for dosing as a prodrug) although it really would be much better if there was also the option to exchange this warhead for the nitrile (a warhead that is much-loved by industrial medicinal chemists since it’s rugged, polar and contributes minimally to molecular size). While I’d anticipate that replacement of aldehyde with nitrile will lead to a reduction in potency, it’s necessary to quantify the potency loss to enable the potential of nitriles to be properly assessed. The binding mode observed for 1 is shown below in Figure 1 and it’s likely that the groove region will need to be more fully exploited (this article will give you an idea of the sort of thing I have in mind) in order to achieve acceptable potency if the aldehyde warhead is replaced by nitrile.

The COVID Moonshot project currently appears to be in what many industrial drug discovery scientists would call the hit-to-lead phase.  In my view the principal objective of hit-to-lead work is to create options since having options will give the lead optimization team room to manoeuvre (you can think of hit-to-lead work as being a bit like playing in midfield). The COVID Moonshot project is currently focused on exploitation of hits from a fragment screen against MPro and, while I’d question whether this approach is likely to get to a candidate drug more quickly than the conventional structure-based design used in industry to pursue cathepsins, it’s certainly an interesting project that I’m happy to contribute to. It’s also worth mentioning that fragment screens have been run against SARS-CoV-2 Nsp3 macrodomain at UCSF and Diamond since there are no known inhibitors for this target.

Here’s a blog post by Pat Walters in which he examines the structure-activity relationships emerging for the fragment-derived inhibitors. Specifically, he uses a metric known as the Structure-Activity Landscape Index (SALI) to quantify the sensitivity of activity to structural changes. Medicinal chemists apply the term ‘activity cliff’ to situations where a small change in structure results in a large change in activity and I’ve argued that the idea of quantifying the sensitivity of a physicochemical effect to structural modifications goes all the way back to Hammett.  One point that comes out of Pat’s post is that it’s difficult to establish structure-activity relationships for low affinity ligands with a conventional biochemical assay. When applying fragment-based approaches in lead discovery, there are distinct advantages to being able to measure low binding affinity (~ 1 mM) since this allows fragment-based structure-activity relationships to be explored prior to synthetic elaboration of fragment hits. As Pat notes, inadequate solubility in assay buffer clearly places limits on the affinity that can be reliably measured in any assay although interference with the readout of a biochemical assay can also lead to misleading results. This is one reason that biophysical detection of binding using methods such as surface plasmon resonance (SPR) are favored in fragment-based lead discovery. Here’s an article by some of my former colleagues which shows how you can assess the impact of interference with the readout of a biochemical assay (and even correct for it if the effect isn’t too great).     

My first contribution to the COVID Moonshot project is illustrated in Chart 2 and the fragment-derived inhibitor 3 from which I started is also featured in Pat’s post. From inspection of the crystal structure, I noticed that the catalytic cysteine might be targeted by linking a ‘reversible’ warhead from the amide nitrogen (4 and 5). Although this might look fine on paper, the experimental data in this article suggest that linking any saturated carbon to the amide nitrogen will bias the preferred amide geometry away from trans to cis. Provided that the intrinsic gain in affinity resulting from linking the warhead is greater than the cost of adopting the bound conformation, the structural modification will lead to a net increase in affinity and the structures could be locked (here's an article that shows how this can work) into the bound conformation (e.g. by forming a ring).


In addition to being accessible to a warhead linked from the amide nitrogen of 3, the catalytic cysteine is also within striking distance of the carbonyl carbon and it would be prudent to consider the possibility that 3 and its analogs can function as substrates for Mpro. There is precedent for this type of behavior and I’ll point you toward an article that notes that a series of esters identified as cruzain inhibitors can function as substrates and more recent article that presents cruzain inhibitors that I’d consider to be potential substrates. A crystal structure of the protein-ligand complex is potentially misleading in this context since the enzyme might not be catalytically active. I believe that 6 could be used to explore this possibility since the carbonyl carbon would be expected to be more electrophilic and 3-hydroxy, 4-methylpyridine would be expected to be a better leaving group than its 3-amino analog.

This is a good point to wrap things up. I think that Satchel Paige gave us some pretty good advice on how to approach drug discovery and that's yet another reason that Black Lives Matter.

Thursday, 13 September 2018

On the Nature of QSAR

With EuroQSAR2018 fast approaching, I'll share some thoughts from Brazil since I won't be there in person. I've not got any QSAR related graphics handy so I'll include a few random photos to break the text up a bit.



East of Marianne River on north coast of Trinidad

Although Corwin Hansch is generally regarded as the "Father of QSAR", it is helpful to look further back to the work of Louis Hammett in order to see the prehistory of the field. Hammett introduced the concept of the linear free energy relationship (LFER) which forms the basis of the formulation of QSAR by Hansch and Toshio Fujita. However, the LFER framework encodes two other concepts that are also relevant to drug design. First, the definition of a substituent constant relates a change in a property to a change in molecular structure and this underpins matched molecular pair analysis (MMPA). Second, establishing an LFER allows the sensitivity of physicochemical behavior to structural change to be quantified and this can be seen as a basis for the activity cliff concept.


Kasbah cats in Ouarzazate 

As David Winkler and the late Prof. Fujita noted in this 2016 article, QSAR has evolved into "two QSARs":

Two main branches of QSAR have evolved. The first of these remains true to the origins of QSAR, where the model is often relatively simple and linear and interpretable in terms of molecular interactions or biological mechanisms, and may be considered “pure” or classical QSAR. The second type focuses much more on modeling structure–activity relationships in large data sets with high chemical diversity using a variety of regression or classification methods, and its primary purpose is to make reliable predictions of properties of new molecules—often the interpretation of the model is obscure or impossible.

I'll label the two branches of QSAR as "classical" (C) and "machine learning" (ML). As QSAR evolved from its origins into ML-QSAR, the descriptors became less physical and more numerous. While I would not attempt to interpret ML-QSAR models, I'd still be wary of interpreting a C-QSAR model if there was a high degree of correlation between the descriptors. One significant difficulty for those who advocate ML-QSAR is that machine learning is frequently associated with (or even equated to) artificial intelligence (AI) which, in turn, oozes hype. Here are a couple of recent In The Pipeline posts (don't forget to look at the comments) on machine learning and AI.

One difference between C-QSAR models and ML-QSAR models is that the former are typically local (training set compounds are closely related structurally) while the the latter are typically non-local (although not as global as their creators might have you believe). My view is that most 'global' QSAR models are actually ensembles of local models although many QSAR modelers would have me dispatched to the auto-da- for this heresy. A C-QSAR model is usually defined for a particular structural series (or scaffold) and the parameters are often specific (e.g. p value for C3-substituent) to the structural series. Provided that relevant data are available for training, one might anticipate that, within its applicability domain, local model will outperform a global model since the local model is better able to capture the structural context of the scaffold.

I would guess that most chemists would predict the effect on logP of chloro-substituting a compound more confidently than they would predict logP for the compound itself. Put another way, it is typically easier to predict the effect of a relatively small structural change (a perturbation) on chemical behavior than it is to predict chemical behavior directly from molecular structure. This is the basis for using free energy calculations to predict relative affinity and it also provides a motivation for MMPA (which can be seen as the data-analytic equivalent of free energy perturbation). This suggests viewing activity and properties in terms of structural relationships between compounds. I would argue that C-QSAR models are better able than ML-QSAR models to exploit structural relationships between compounds.


Down the islands with Venezuela in the distance 

ML-QSAR models typically use many parameters to fit the data and this means that more data is needed to build them. One of the issues that I have with machine learning approaches to modeling is that it is not usually clear how many parameters have been used to build the models (and it's not always clear that the creators of the models know). You can think of number of parameters as the currency in which you pay for the quality of fit to the training data and you need to account for number of parameters when comparing performance of different models. This is an issue that I think ML-QSAR advocates need to address.

Overfitting of training data is an issue even for C-QSAR models that use small numbers of parameters. Generally, it is assumed that if a model satisfies validation criteria it has not been over-fitted. However, cross-validation can lead to an optimistic assessment of model quality if the distribution of compounds in the training space is very uneven. An analogous problem can arise even when using external test sets. Hawkins advocated creating test sets by removing all representatives of particular chemotypes from training sets and I was sufficiently uncouth to mention this to one of the plenaries at EuroQSAR 2016. Training set design and model validation do not appear to be solved problems in the context of ML-QSAR.


The Corniche in Beirut 

I get the impression that machine learning algorithms may be better suited for classification than QSAR and it is common to see potency (or affinity) values classified as 'active' or 'inactive' for modeling. This creates a number of difficulties and I'll also point you towards the correlation inflation article that explains why gratuitous categorization of continuous data is very, very naughty. First, transformation of continuous data to categorical data throws away huge amounts of information which would seem to be the data science equivalent of shooting yourself in the foot. Second, categorization distorts your perception of the data (e.g. a pIC50 value of 6.5 might be regarded as more similar to one of 9.0 than one of 5.5). Third, a constant uncertainty in potency translates to a variable uncertainty in the classification. Fourth, if you categorize continuous data then you need to demonstrate that conclusions of analysis do not depend on the categorization scheme.

In the machine learning area not all QSAR is actually QSAR. This article reports that "the performance of Naïve Bayes, Random Forests, Support Vector Machines, Logistic Regression, and Deep Neural Networks was assessed using QSAR and proteochemometric (PCM) methods". However, the QSAR methods used appear to be based on categorical rather than quantitative definitions of activity. Even when more than two activity categories (e.g. high, medium, low) are defined, analysis might not be accounting for the ordering of the categories and this issue was also discussed in the correlation inflation article. Some clarification from the machine learning community may be in order as to which of their offerings can be used for modelling quantitative activity data.


I'll conclude the post by taking a look at where QSAR fits into the framework of drug design. Applying QSAR methods requires data and one difficulty for the modeler is that the project may have delivered its endpoint (or been put out of its misery) by the time that there is sufficient data for developing useful models. Simple models can be useful even if they are not particularly predictive. For example, modelling the response of pIC50 to logP makes it easy to see the extent to which the activity of each compound beats (or is beaten by) the trend in the data. Provided that there is sufficient range in the data, a weak correlation between pIC50 and logP is actually very desirable and I'll leave it to the reader to ponder why this might be the case. My view is that ML-QSAR models are unlikely to have significant impact for predicting potency against therapeutic targets in drug discovery projects.  

So that's just about all I've got to say. Have an enjoyable conference and make sure keep the speakers honest with your questions. It'd be rude not to.


Early evening in Barra 

Saturday, 7 April 2018

Hammett

I first became aware of Louis Hammett during the third term of my first year as an undergraduate at the University of Reading. Hammett was a pioneer in physical-organic chemistry and is widely regarded as one of the founders of that field. He would have been 124 today and was less than a year younger than Christopher Ingold, another pioneer in the field. Hammett passed away in 1987 at the age of 92 (here is an excellent obituary).



Today Hammett is remembered primarily for the parameters that describe electronic interactions between aromatic rings and their substituents. He also introduced linear free energy relationships which form the basis of classical QSAR. These days, QSAR has evolved away from its origins in physical-organic chemistry into what many call machine learning and parameters have become less physical (and considerably more numerous). Hammett's work provided an early lesson to wannabe molecular designers in how to think about molecules.

Jens Sadowski and I introduced matched molecular pair analysis (MMPA) in a chapter of a cheminformatics book that was conceived and edited by my dear friend (and favorite Transylvanian) Tudor Oprea. Here's a photo of Tudor and me at an OpenEye meeting (I think CUP II in 2001) during which our props (Tudor is wearing a PoD cape) were provided by the session chair (the formidable Janet Newman who intimidates proteins to the extent that they 'voluntarily' crystallize).


Now you might be wondering what MMPA has to do with Hammett. The short answer is that our book chapter included a table of what are effectively substituent constants for aqueous solubility and these have Hammett's fingerprints all over them. The longer answer is that Hammett introduced the idea of associating parameters with structural relationships (e.g. X is chloro analog of Y) between compounds. This is an important idea because much pharmaceutical design is focused on understanding and predicting the effects of structural modifications on the activity and properties of compounds. One rationale for this focus is the belief that it is easier to predict differences (e.g. relative affinity) in chemical behavior between structurally-related compounds than it is to predict chemical behavior directly from molecular structure.

At first, I didn't see the deeper connection between Hammett's work and pharmaceutical design. The main focus of our book chapter was preparing chemical structures in databases for virtual screening so the full extent of Hammett's influence on MMPA was not immediately recognized. As is often the case, we think we've discovered something really new only to find out later that somebody had been thinking along similar lines many years before. 

Happy 124th birthday, Louis Hammett.