“Quantity has a quality all of its own”
Original source unknown
The quote with which I’ve opened the post on the OpenADMET initiative is often attributed to Joseph Stalin and some suggest that he might have been referring to the T-34. While the safety and comfort of crews were not priorities for Soviet tank designers the T-34 was a better tank than the quote might be taken to imply. In particular, the distinctive sloped armour presented challenges for Wehrmacht anti-tank gunners (at least until the introduction of the the formidable 88 mm PAK 43) while the wider tracks of the T-34 enabled it to operate effectively in conditions that the more refined (and heavier) Tiger could not. Here are some photos of T-34s taken at the Brest Fortress and the War Memorial of Korea:
The principal objective of the OpenADMET initiative appears to be generation of measured data to enable machine learning (ML) models to be built for prediction of toxicity and absorption, distribution, metabolism, and excretion (ADME). Just as the success of the T-34 on the Eastern Front was not just down to being available in huge numbers there is a bit more to data generation than simply generating massive data sets. A significant challenge for initiatives such as OpenADMET is covering chemical space at a sufficiently fine resolution and with sufficient spread in the measured data to enable building of ML models that can predict accurately across a diverse range of chemotypes. As I’ve noted previously drug discovery project teams have delivered (and continue to deliver) clinical development candidates without ever having sufficient data for building ML models that can accurately predict all the quantities of interest to the project teams. While I'll be criticising aspects of the OpenADMET initiative in this post it must be stressed that I do see great value to drug discovery in making relevant data freely available in the public domain (and Open Science in general).
Previously I suggested that drug design can be thought of in terms of three objectives and the OpenADMET initiative addresses the second (minimize off-target bioactivity) and third (maximize controllability of exposure) of these. As I argued in that previous post Absorption, Distribution, Metabolism, Excretion (ADME) and Toxicity (T) should be seen as separate issues in the design context (my view is that using the term ADMET projects a lack of familiarity with the practical realities of drug design). Put another way toxicity is something that drugs do to the human body while ADME determines what the human body does to drugs (pharmacy students typically encounter the distinction between pharmacodynamics and pharmacokinetics early in their training). This is a good point at which to mention the Avoid-ome and here’s a post from a couple of years ago. While I agree that toxicity fits naturally into an Avoid-ome framework, I'm unconvinced that the introduction of the term is the Great Leap Forward that some believe it to be. However, ADME issues generally cannot be accommodated within an Avoid-ome framework because ADME-based design is ultimately about control of exposure (concentration of drug in contact with the target or anti-target) and not about avoidance.
This is a good point at which to take a look at the recent F2026 article (Mapping the avoid-ome: a systematic open-science approach to predictive ADMET). As is customary for posts at this blog quoted text is indented with my comments enclosed within square brackets in red italics, and I’ve used the F2026 reference numbers.
I’ll start with the abstract:
Drug discovery often fails due to unpredictable ADMET issues, which account for 30% of clinical setbacks. [While fully agreeing that toxicity and poor ADME are important issues that that do need to be addressed more effectively I do not consider references (3) (4) (5) to support what the authors have stated here or their claim that “more than 90% of molecules created during discovery fail to meet basic ADME standards”. For example, reference (3) states that “the major causes of attrition in the clinic in 2000 were lack of efficacy (accounting for approximately 30% of failures) and safety (toxicology and clinical safety accounting for a further approximately 30%)”. All that said, decisions to take compounds into clinical development are based on measurements made in a range of assays and failure in clinical development reflects an inability of these assays to predict clinical outcomes. To more effectively address attrition we actually need new assays that are more predictive of outcomes in clinical development as opposed to new ML models that are more predictive of quantities that will need to be measured anyway.] Conventional methods lack the atomistic detail needed to navigate the “Avoid-ome”—a finite set of proteins acting as “anti-targets”. OpenADMET is an open-science initiative addressing this by creating pre-competitive, mechanistic datasets. [With respect to to "atomistic detail" it's important to bear in mind that structures for transition states (relevant when the quantity of interest is rate of turnover by metabolic enzymes) cannot actually be observed in experimental protein structural studies.] Using high-throughput structural biology, active learning, and community challenges, it builds generalizable models grounded in structural “ground truth”. [I would question the wisdom of invoking “ground truth” in scientific studies because it brings to mind stirring sermons on themes like "Pastor needs an additional private jet" and invocation of “ground truth” will endow your arguments with a distinctly pastoral odour. Building truly generalizable ML models for the quantities of interest to drug designers would certainly be of great value in drug discovery but it will take much more than anointing models with "ground truth" to achieve this.] By directly studying the Avoid-ome, OpenADMET facilitates an era of rational, multi-parameter drug design. [My view is that drug design should be described as ‘multi-objective’ rather than ‘multi-parameter’ because design against a single objective such as affinity maximisation can still be multi-parameter in nature, and I consider it tautological to describe drug design as ‘rational’.]
Understanding and navigating the Avoid-ome is the central universal challenge of modern drug discovery. [My view, which might be shared by a few others, is that the principal challenge for drug discovery is (and has always been) the uncertainty that results from the complexity of human biology and here’s a relevant blog post by "That Dude That Says That AI Drug Discovery Isn't So Amazing".] By creating open, structural, and mechanistic datasets and benchmarking predictive models through blind challenges, OpenADMET provides a practical foundation for a new era of rational drug design. [I am a big fan of of Open Science and see a significant value in making data relevant to drug discovery freely available. At the risk of repetition the term “rational drug design” is tautological and the promise of (yet another) “new era” will trigger the eye-rolling reflex for many experienced drug hunters. Given the aspirational nature of the initiative at this stage I think it would be more accurate to state that "OpenADMET aims to provide" rather than "OpenADMET provides".] The best way to increase the effectiveness of drug discovery in the coming decade is to stop avoiding the Avoid-ome and instead study it directly. [While I certainly see benefits from a deeper understanding of the proteins that cause toxicity and influence ADME I don’t see doing so as quite the panacea that the authors of F2026 would have us to believe it to be. I argued over two years ago that the Avoid-ome is generally not a useful concept for consideration of ADME issues and my view is that shackling OpenADMET to the Avoid-ome has actually reduced the scientific credibility of the initiative. My advice to those leading the OpenADMET initiative would actually be to drop the Avoid-ome before it's too late.]
To be fair to the authors of F2026 do concede ensuring that that there is a lot more to ADME optimization than ensuring that anti-targets are not engaged although I would still challenge their assertion that "understanding and navigating the Avoid-ome is the central universal challenge of modern drug discovery".
While our framework heavily emphasizes the specific protein anti-targets of the Avoid-ome, we recognize that fundamental physicochemical and integrative properties—such as aqueous solubility, membrane permeability (logD), and metabolic stability are major drivers of ADMET outcomes, particularly for absorption and excretion [absorption and excretion would be more accurately described as 'ADME outcomes' than 'ADMET outcomes']. Although these factors are not mediated by a single anti-target, [these factors are not mediated by anti-targets] they are critical bulk molecular properties [I consider the term "bulk molecular properties" to be an oxymoron] that often dictate whether a compound succeeds or fails.
Having got the general stuff out of the way I’ll examine the OpenADMET initiative from the perspectives of both ADME and toxicity, starting with the latter. My view is that any off-target bioactivity is undesirable given the complexity of human biology although bioactivity against known anti-targets such as hERG is clearly unacceptable. It’s also important to take account of the concentration at which off-target effects are observed and a weakness shared by many (most?) studies of pharmacological promiscuity is that bioactivity thresholds are set far too permissively to have any physiological relevance whatsoever (LS2007 classifies compounds that exhibit >30% inhibition at 10 µM as 'active' and, more recently, FOM2025 states that "Mestres et al. (p4) anticipated that the average number of proteins with which a drug interacts with potentially relevant bioactivity (<10 µm) was close to six").
It’s perhaps appropriate to take a general look at QSAR/QSPR approaches given that the main focus of OpenADMET appears to be generation of data for training what could be referred to as 'QSAR-like' or 'QSPR-like' ML models. In my view, the impact of QSAR/QSPR modelling on real world drug discovery was limited and claims to the contrary are generally not verifiable. A difficulty faced by those advocating the use of QSAR/QSPR approaches was that projects had either delivered or been put out of their misery by the time there was sufficient data for building predictively useful models. My view is that modern ML models, like the QSAR and QSPR models that preceded them, can’t generally extrapolate out of the chemical spaces in which they were trained. While I certainly wouldn’t claim ground truth for this view, I’m not aware of any studies in which a QSAR or QSPR model built using only data from one structural series was convincingly shown to be usefully predictive of for compounds in a different structural series. Medicinal chemists typically perform their optimizations within specific structural series and this means that structure-activity/property relationships (SARs/SPRs) tend to be local in nature. For users of ML models of bioactivity and other properties of compounds it is important to know whether chemical structures for which predictions are being made lie within the applicability domains of the models. Put another way, medicinal chemists who use ML models are generally much more interested in how well the models will predict for the structural series that they're working on and much less interested in how well the models have fit the training data (anybody who has received financial advice will be familiar with the "past performance is not indicative of future results" disclaimer). The selection criteria for assays and compounds by the OpenADMET initiative are not currently clear.
I see a degree of overlap between the OpenADMET and OpenBind initiatives in that safety assessment will often require prediction of binding affinity of anti-targets for compounds being considered for synthesis. Indeed, there is no reason that structures for complexes of anti-targets with ligands should be excluded from data sets when the objective is to build universal ML models for prediction of binding affinity. Nevertheless, categorical models for prediction of off-target bioactivity still have value in hit-to-lead work and lead optimization whereas categorical models for predicting on-target bioactivity are generally only useful during hit identification.
While it will often prove feasible to build ML models for binding affinity of anti-targets for ligands, many users will want to know whether the chemical structures of interest to them lie within the applicability domains of the models. As discussed in my post on the OpenBind initiative, using geometric features in protein-ligand complexes as descriptors to train ML models can potentially enable accurate affinity predictions to be made for chemotypes that are not represented in the training data. My view is that those leading the OpenADMET initiative will need to be more explicit about how (or even whether) they propose to use protein structural information for affinity prediction. Not all off-target effects of drugs can be specified in terms of reversible binding (as is the case for PXR activation which forms the basis for the current OpenADMET challenge) and this is generally more likely to be an issue when attempting to build models for off-target bioactivity.
Let's now take a look ADME from the perspective of ML modelling. As noted earlier this post, ADME and toxicity are completely different issues in drug design and using the term ‘ADMET’ conveys (at least to me) an impression that some of the practical realities of drug discovery have not been properly understood. Drug action is driven by the concentration of the drug at its site(s) of action (the term exposure is commonly used) and one of the practical realities of drug discovery is that drug concentration at sites of action generally cannot be measured in vivo unless binding sites are directly in contact with plasma (see post on the objectives of drug design). I also suggest that readers take a look at SR2019 (Smith and Rowland, Intracellular and Intraorgan Concentrations of Small Molecule Drugs: Theory, Uncertainties in Infectious Diseases and Oncology, and Promise | DMD 2019 47:665-672) that I recommend to everybody working in drug discovery and chemical biology. While maximization of affinity is a legitimate design objective, exposure is something that needs to be carefully controlled rather than simply maximized (I'm guessing that Paracelsus might have cautioned against maximization of exposure and he’s been dead for almost half a millennium).
Pharmacokinetic/pharmacodynamic (PK/PD) modelling (DM1999 | D2008 | R2008 | NCF2011 | W2015 | SS2017 | BL2023 | HR2024 | C2024 ) is used in the later stages of drug discovery to predict the (dose-dependent) effects of drugs on humans. The inputs for PK/PD modelling are predicted human pharmacokinetic profiles (typically generated from the results from PK profiles observed in animal studies) and bioactivity measurements. In PK/PD modelling it is usual to invoke the free drug hypothesis (SYF2022 | W2025) by assuming that the concentration of a drug at its site of action is equal to the unbound concentration of the drug in plasma (the terms ‘free drug principle’ and ‘free drug theory’ are also used although I prefer ‘free drug hypothesis’ because it’s an assumption that is being made). There are two scenarios under which this assumption is known to break down. First, there is active transport at one or more points in the path between the drug’s site of administration and its site of action. Second, the drug is ionizable and its site of action is within a compartment where the pH differs from plasma pH (basic centres not required for binding to the target are generally not recommended when targeting lysosomal enzymes if you’re concerned about selectivity).
The ADME acronym refers to in vivo phenomena and physicochemical properties (lipophilicity, aqueous solubility, passive membrane permeability) and in vitro biochemical quantities (turnover by metabolic enzymes, active efflux) should actually be described as ‘ADME predictors’ and not ‘ADME properties’. My view (which I’ll be happy to change in the light of compelling evidence) is that it isn’t currently possible to predict in vivo plasma concentration to the level of accuracy required for PK/PD modelling if you’re only using in vitro ADME predictors. That said, I’m certainly not suggesting that genuinely predictive models for aqueous solubility, membrane permeability and turnover by metabolic enzymes are without value in drug discovery projects.
In a previous post I argued that drug designers should aim to maximize controllability of exposure and to do so requires a focus on pharmacokinetic profile rather than individual ADME predictors. In many cases the ideal pharmacokinetic profile would be that resulting from intravenous infusion whereby the plasma concentration of a drug is maintained at the minimum level required for therapeutically useful effects (the variation in plasma concentration over the dosing interval for an orally-dosed drug is actually undesirable from the perspective of simultaneously achieving efficacy and safety). While ML models for quantities such as aqueous solubility, membrane permeability and turnover by metabolic enzymes certainly have a place in drug design, decisions as to whether a compound should be evaluated in vivo will generally be based on in measured values rather than predictions from ML models.
Let’s now take a look at ADME in the context of the Avoid-ome. To be fair, the authors of F2026 do concede that ADME doesn’t quite fit into the Avoid-ome framework although this does rather beg the question as to why they assert that “understanding and navigating the Avoid-ome is the central universal challenge of modern drug discovery”. One fundamental problem with the F2026 article is that its authors have got themselves in a bit of a tangle with respect to how they’ve defined anti-targets. In drug discovery, the term ‘anti-target’ generally refers to a protein which is associated with the risk of toxicity when engaged in vivo. The authors of F2026 appear to be broadening the 'anti-target’ definition to include proteins that they believe have detrimental effects on the ADME behaviour of compounds and, in my view, it is neither valid nor useful to do so.
Let’s take a look at Fig. 1 (The set of protein anti-targets that comprise the Avoid-ome) in F2026 and you’ll see that a number of transporters are included as anti-targets in the graphic. While I certainly agree that active efflux is generally undesirable from the perspective of achieving adequate exposure, transporters would not automatically be regarded as anti-targets from the safety perspective. That said, inhibition of bile salt export pump (BESP) is considered a risk factor for drug-induced liver injury (DILI) and here's a link to a relevant article.
The pitfalls associated the broadening the definition of ‘anti-target’ are brought more sharply into focus when metabolic enzymes enter the picture. It is widely accepted that CYP inhibition is a safety issue because of the potential for drug-drug interactions (LL1998 | H2020 | L2024) but CYPs have been labelled in Fig. 1 of F2026 as metabolism (M) anti-targets. While I certainly agree that high clearance makes it difficult to maintain exposure at levels required for therapeutic benefits, a recommendation that clearance be entirely avoided would generally be regarded by drug safety scientists as very bad advice indeed (it can be instructive to consider why CYP inhibition is considered to be a safety issue). As a reviewer of F2026 I would have pressed its authors to explain why they had classified the aromatic hydrocarbon receptor (AhR) in Fig. 1 as a metabolism (M) anti-target given that the toxicity of dioxin results from engagement of this receptor (M2005 | S2014).
The authors of F2026 include serum albumin (HSA) in the Avoid-ome and it’s fair to say that plasma protein binding (PPB) is a source of confusion even for medicinal chemists (there would probably be much less confusion if unbound plasma concentration was measured directly in PK studies and I’ll direct readers to the SDK2010 article). The recent B2025 study (What Do Oral Drugs Really Look Like? Dose Regimen, Pharmacokinetics, and Safety of Recently Approved Small-Molecule Oral Drugs) reports that "most drugs are >95% plasma protein bound (58%), with a large fraction >99% bound (29%)" which appears to contradict the view expressed in F2026 that HSA should be considered as an anti-target. I would argue that PPB is potentially beneficial in that it protects the drug to some extent during first pass metabolism (distribution from plasma also protects of drug from plasma only occurs after).
This is a good point at which to wrap up. While OpenADMET is likely to generate some useful data and models, I don’t see it as a game-changer and to claim that the initiative “provides a practical foundation for a new era of rational drug design” is fanciful in my view. Something that comes across from my reading of F2026 is a general lack of expertise in the areas of drug safety, pharmacokinetics and PK/PD modelling, and I urge those leading OpenADMET to address these deficiencies. I have argued that the Avoid-ome is neither valid nor even useful as a framework for ADME-based drug design and my advice to those leading OpenADMET is to quietly drop the Avoid-ome.


No comments:
Post a Comment