Thursday, 28 March 2013

Efficient Voodoo Thermodynamics

So you’ve got a little data analysis problem.  You have some compounds with a range of IC50 values and you’d like to explore the extent that molecular size contributes to potency/affinity.  Here’s one suggestion for starters.  Plot pIC50 against your favourite measure of  molecular size which could be number of number of non-hydrogen atoms, molecular weight and look at the residuals which tell you how much each compound beats the trend (or is beaten by it).  Let’s start by defining  pIC50 which you can calculate from:

 pIC50 = -log(IC50/M)                                                                              1
You might be asking why I’m not suggesting that you use the standard Gibbs free energy of binding which is defined by:

ΔG° = RTln(Kd/C°)                                                                                 2
The main reason for not doing this is that we can’t.  You’ll notice that ΔG° is calculated from Kd and not  IC50 and the two are not the same thing even though they both have units of concentration.  Those of you who have worked on kinase projects may have even used IC50 values measured at different ATP concentrations to get a better idea how much kick your inhibitors will have at physiological ATP concentration.  Put another way, you can measure the concentration of sugar in your coffee and plug this into equation 2 but that does not make what you calculate a standard Gibbs free energy of binding.   If you’ve measured IC50 then you really should use  pIC50 in this analysis.  Converting pIC50 to ΔG° is technically incorrect and arguably pretentious since the converted pIC50 can give the impression that it is somehow more thermodynamic than that from which it was calculated.  Converting pIC50  to ΔG° also introduces additional units (of energy/mole) and there is always a degree of irony when these units, which may have been introduced to just make biological data look more physical, get lost when the results are presented.

So let’s get back to the data analysis problem.  Suppose that we’ve plotted  pIC50 against number of heavy (i.e. non-hydrogen) atoms and the next step is to fit the data.  Best way to start is to fit a straight line although you could also fit a curve if the data justifies this.  Let’s assume that we’re fitting the straight line:
  pIC50 = A + B×NHA                                                                            3

I realise that using an intercept term (A) will cause a few eyebrows to become raised.  Surely the line of fit should go through the origin?   There is a problem with this line of thinking and it’s helpful now to talk instead in terms of affinity and ΔG° to develop the point a bit more.  You might be thinking that in the limit of zero molecular size a compound should have zero free energy of binding.  However, a zero free energy of binding corresponds to Kd being equal to the standard concentration and you’ll remember that the choice of standard state is arbitrary.  If you must derive insights from thermodynamic measurements then the very least that you can do is to ensure that any insights you derive are invariant with respect to the value of the standard concentration.  

When you use ligand efficiency (-ΔG°/NHA) you’re effectively assuming that the value of ΔG° should be directly proportional to the number of heavy atoms in the ligand molecule.  One consequence of defining ligand efficiency in this manner is that relative values of ligand efficiency for compounds with different numbers of heavy atoms will change if you change (as thermodynamics tells us that we are allowed to do) the standard concentration used to define ΔG°.  I've droned on enough though and it's time to check out.  I will however, leave you with the question of whether it makes sense to try to correct ligand efficiency for the effects of molecular size. 

Sunday, 17 March 2013

The wrong kind of free energy

It has been some time since I last blogged.   There has been the distraction of a little drug-likeness project and there is still much to do before I leave Brasil at the end of May.  I have, however, set up a twitter account (@pwk2013) which I use in futile attempts to engage minor celebrities (like God and Richard Dawkins) in conversation.   I’ll be taking a look at some aspects of thermodynamics in this post and I’ll start by writing the relationship between dissociation constant and the standard free energy of binding:

ΔG° = RTln(Kd/C°)

I realise that this might look a bit different to what you’re used to so I’ll try to explain why its been written like this.  First of all Kd has units of concentration and logarithms are only defined for numbers (i.e. unitless quantities) so you might want to consider the possibility that what you normally write is wrong (apologies for hideous pun).   The other thing that you need to remember is that the standard free energy of binding depends on the choice of standard state and writing the equation like this makes this connection explicit.   ΔG° is negative for a 10nM compound when the standard concentration is 1M.  However, if we were to define the standard concentration to be 1nM then   ΔG° would be positive.  If you consider the idea of a 1nM standard concentration to be offensive, think of a 1M solution of your favourite protein...
One of the misconceptions that drug discovery researchers tend to have is that you can’t call it thermodynamics unless you measure both enthalpy and entropy.  One reason for this is the use of isothermal titration calorimetry (ITC) to measure affinity.  ITC is a direct, label-free method for measuring affinity and you get the enthalpy as part of the package.  One of the important things to remember about thermodynamics is that you need to use the most appropriate thermodynamic quantity to describe the phenomenon that you’re interested in and for binding at constant pressure this is the Gibbs free energy.  In contrast, the engineer scaling up a synthetic reaction for manufacturing a drug will usually be more interested in the heat and volume (i.e. gaseous bi-products) generated by the reaction because failure to control these can result in the Big Kaboom.

One idea that has emerged as ITC has become more accessible in drug discovery is that enthalpy-driven binding is somehow better than entropy-driven binding.  I remain sceptical and ask the rhetorical question of how an isothermal system such as a human taking a drug senses the degree to which engagement of the drug’s target(s) is driven by enthalpy changes.   If measuring enthalpy of binding in addition to affinity helps us to predict affinity for compounds yet to be synthesised or gives us clear (i.e. keeping arms stationary) insights into the nature of binding then it makes sense to do it.  However, I’m not seeing this being done convincingly and sometimes wonder whether enthalpy optimisation is just a case of the wrong kind of snow...

We tend to interpret affinity in terms of contacts between protein and ligand although one must always remember that the contribution of a particular contact to affinity is not in general an experimental observable.   Proteins and their ligands associate in an aqueous environment and the non-local nature of the hydrophobic effect further complicates attempts to use structure to rationalise affinity.  I often hear hydrophobic interactions being described as non-directional and in my view this is complete bollocks since interactions are forces and forces are vectors which by definition have direction.  On this subject, I can't resist telling the tale of the Human Resources department getting the Scalar Quantity Award for having magnitude but no direction...
I think this is a good place to leave things and I'll try to be back in a week or some more specific stuff in a follow up to this post. 

Thursday, 18 October 2012

Screening library size

<< previous || next >>

This post got prompted by a survey of screening library size by Teddy at Practical Fragments. I commented briefly there but it soon became clear that it was going to take a really long comment to say everthing. While at AstraZeneca, I was involved in putting together a 20k screening library for fragment-based work and I thought that some people might be interested in reading about some of the experiences. The project started at the beginning of 2005 although the library (GFSL2005) was not assembled until the folllowing year. Some aspects of the library design were also described in our contribution to the JCAMD special issue on FBDD. The NMR screening libraries with which I've been involved have been a lot smaller (~2k) than this so why would one want to assemble a fragment screening library of this size? One reason was diversity of fragment-based activities, which included a high concentration component when running HTS In screening, you soon learn the importance of logistics and, in particular, compound managment. If you're going to assemble a screening library that can be easily used, this usually means getting the compounds into dimethyl sulfoxide (DMSO) so that samples can be dispensed automatically. Fragments will usually be screened at high concentration which means that stock solutions also need to be more concentrated (we used 100µM) because most assays don't particularly like DMSO. Another reason for maintaining stock solutions is that there are minimum volume limits for dissolving solids in DMSO so starting from solids each time you run a screen is wasteful of material as well as time. The essence of the GFSL2005 project was ensuring that there were enough high concentration stock solutions in the liquid store to support diverse fragment-based activities on different Discovery sites. While the entire library might be screened as a component of HTS, individual groups could screen subsets of GFSL05 that were significantly smaller (at least when I left AZ in 2009). It's important to rememember that while GFSL05 was a generic library we were also trying to anticipate the needs of people who might be doing directed fragment screening. The compounds were selected using Core and Layer (which I've discussed before). This approach is not specific to fragment screening libraries and I've used it to put together a compound library for phenotypic screening. The in house software, some of which had been developed in response to emergence of HTS, was actually in place by the beginning of 1996. Here (from one of my presentations) is a slide that illustrates the approach.
 
One of the comments on the Practical Fragments post was about molecular diversity and the anonymous commentator asked: "Are some companies more keen because of their bigger libraries or just lazy and adding more and more compounds without looking at diversity?" This is a good challenge and I'll try to give you an idea of what were trying to do. Core and Layer is all about diversity in that the approach aims to drive compound selection away (in terms of molecular similarity) from what has already been selected. However, when designing the library we did try to ensure that as many compounds as possible had a near neighbour (within the library). Here's a slide showing (in pie chart format) the fractions of the library corresponding to each number of neighbours. There are actually three pie charts because counting neighbours depends on the similarity threshold that you use and I've done the analysis for three different thresholds.
 
 You can do a similar analysis to ask about neighbours of library compounds that are available. This is of interest because you'll generally want to be able to follow up hits with analogues in a process that some call 'SAR-by-Catalogue' although I regard the term as silly and refuse to use it. Availability is not constant and the analysis shown in the following slide was a snapshot generated for a 2008. You'll notice that there is an extra row of pie charts since one can define availability more (> 20mg) or less (>10 mg) conservatively. If you require plenty of sample for availability and require a neighbours to be very similar then you'll have less neighbours.
 
 There was also some commentary on solubility and ideally this should be measured in assay bufffer for library compounds. We used of our in house high throughput solubility assay when putting GFSL2005 together. This solubility assay had original been designed for assessing hits from conventional HTS and had a limited dynamic range (1-100 µM) which was not ideal for fragments. The assay used standard (i.e. 10mM) stock solutions which meant that we could only do measurements for samples that were in this format (library compounds acquired from external sources were only made up as 100mM stocks). Nevertheless, we used the assay since we believed that the information would still be valuable. The graphic below illustrates the relatonship between solubility and lipophilicity (ClogP) for compounds that are neutral under assay conditions. We used ClogP to bin the data which allowed us to make use of out of range data by plotting percentiles. This allowed us to assess the risk of poor solubily as a function of ClogP. It's worth pointing out that the binning was done so as to be able to include in range and out of range data in a single analysis. We just showed the lowest percentiles because we were only interested in the least soluble compounds. Even so, you should be able to get an idea of the variation in solubility for each bin.
 
Staying with the solubility theme I'll finish off taking a look at with a look at the aromatic rings and carbon saturation as determinants of aqueous solubility. The two articles at which I'll be looking were both featured in a post at Practical Fragments and, given that the analyses in the two articles were not exactly rose-scented, the title of the post struck me as somewhat ironic. I'll start with Escape from Flatland which introduces carbon bond saturation, defined by fraction sp3 (Fsp3) as a molecular descriptor. Figure 5 is the one most relevant to the discussion and this is captioned, "Fsp3 as a function of logS..." although the caption is not totally accurate. LogS is binned and it is actually average Fsp3 that is plotted and not Fsp3 itself. I'm guessing that the average is the mean (rather than the median) Fsp3 for each logS bin. If you look at the plot there appears to be a good correlation between the mid-point of each logS bin and average Fsp3 although it would have been helpful if they'd presented a correlation coefficient and shown a standard deviation for each bin. In fact, it would have been helpful if they'd shown some standard deviations for the other figures as well but that is peripheral to this discussion. The problem with Figure 5 is that the intra-bin variation in Fsp3 is hidden (i.e. no error bars) and without being able to see this variation it is very difficult to know how strongly Fsp3 (as opposed to its average value) is correlated with logS. Those readers who were awake (it was immediately after lunch) at my RACI talk last December will know what I'm talking about but hopefully some of the rest of you will at least be wondering why the authors didn't simply fit logS to Fsp3. Anyone care to speculate as to what the correlation coefficient between logS and Fsp3 might be?

Impact of aromatic ring count presents a box plot of solubility as a function of number of aromatic rings. At least the variation in solubility is shown for each value of aromatic ring count is shown (even though I'd have preferred to see log of solubility plotted). The authors also looked at the correlation between number of aromatic rings cLogP (which I prefer to call ClogP) and judged it to be excellent (although it is not clear whether they bothered to calculate a correlation coefficient to support their assertion of excellence). Correlations between descriptors are important because the effect of one can be largely due to the extent to which is correlated with another. Although there are ways that you can model the dependence of a quantity on two descriptors that are correlated with each other, the authors chose to do this graphically using the pie chart array in Figure 6. If you look at the pie chart array, you can sort of convince yourself that aromatic ring count has an effect on solubility that is not just due to its effect on ClogP. However, there is established statistical methodology for dealing with this sort or problem. I couldn't help wondering why the authors didn't use this to analyse their data. What puzzled me even more was that they didn't seem to have considered the possiblity that the number of aromatic rings might be correlated with molecular weight (Dan also picked up on this in his post) since I'd guess that this correlation might be even stronger than the one with ClogP. I do believe that a strong case can be made for looking beyond aromatic rings when putting screening libraries and I'll point you to a recent intiative that may be of interest. However, this case is based on considerations of molecular recognition and molecular diversity. I don't believe that either of these studies (which both imply that substituting cyclohexadiene for benzene would be a good thing) strengthens that case. Possibly if the analysis had been done differently I might have arrived at a different conclusion but a lack of error bars and a failure to acccount for the effect of molecular weight leave me with too many doubts.

Literature cited
Blomberg et al, Design of compound libraries for fragment screening. JCAMD 2009, 23 513-525 DOI
Colclough et al, High throughput solubility determination with application to selection of compounds for fragment screening. Bioorg. Med. Chem. 2008, 16, 6611-6616 DOI
Lovering, Bikker & Humblet, Escape from Flatland: Increasing Saturation as an Approach to Improving Clinical Success. J. Med. Chem. 2009, 52, 6752-6756 DOI
Ritchie & MacDonald, The impact of aromatic ring count on compound developability: Are too many aromatic rings a liability in drug design? Drug Discov. Today 2009, 14, 1011-1020 DOI

Monday, 8 October 2012

Viajem à Belém (minha tarefa)

Eu cheguei no Brasil ao fim de abril e tento aprender Português. Minha professor se chama Thalita e eu tenho aula de Português com Mihong, que é coreana. Eu faço um ‘blog post’ de minha viajem à Belém em Português para tarefa. Nao posso usar computador para traduzir mas tenho ajuda para conjugar os verbos irregulars. Esta semana, nós falamos sobre a comida e espero descobrir a receta coreana para cachorro-quente.  Também, eu assisto televisão para ajudar meu compreensivo de  Português e eu gosto do Canal Rural (muitas vacas).

Eu fiz uma palestra ao O IV Simpósio de Simulação Computacional e Avaliação Biológica de Biomoléculas na Amazônia (SSCABBA). O simpósio teve lugar na Universidade Federal do Pará (UFPA) e tenho duas fotos do campus. 




Jerônimo (à esquerda) foi o coordenador do evento e eu gostei muito da palestra de Ernesto (ao lado de Jerônimo).  Ademir (camisa de cor de laranja) fez um curso de Car-Parinello e eu compareci uma das seus palestras.



Minha palestra foi em inglês mas eu fiz uma piada em Português dos Argentinos (los hermanos). Uma molécula de água tem menos energia se ela está longe da superficie hidrofóbico.  Na termodinâmica energia pequena e felidade são equivalentes. Talvez posso fazer esta palestra em Buenos Aires?

Thursday, 20 September 2012

FBLD 2012 Preview

FBLD 2012 is about to happen. Although I'll not be there (and not even nearby), I thought that a preview, like the one posted recently for EuroQSAR (which I also didn't go to) might be in order.

Rod Hubbard will kick things off and he always does a great talk.  Following his contribution will be three talks from Big Pharma.  I've only ever been to one fragment conference and the Big Pharma contributions at that meeting tended to be strategy-heavy and results-light.  Hopefully things will have moved on a bit in the last three and a half years...

Were I attending the meeting, I'd be paying particularly close attention to the membrane protein talks (and might even take the opportunity to ask one of the creators of the Rule of 3 how hydrogen bond acceptors are defined when applying Ro3).  I would also be trying to learn as much as possible about newer technologies for measuring affinity.  Remember that the power of a fragment screening assay is defined by the weakness of the binding that can be detected. Although always desirable, high throughput is a secondary consideration in these assays since one rationale for screening fragments is that one doesn't need to assay as many compounds.  Expect to see at least one speaker get reminded that in SPR molecules are 'tethered' rather than 'immobilised'.

Thermodynamics and kinetics provide the focus of a few of the talks and usually one will hear about the benefits of 'enthalpy-driven' binding and slow off-rates.  My stock question for those who assert the benefits of binding that is 'enthalpy-driven' is, "how do isothermal systems sense enthalpy changes associated with binding?" and I have also made this point in print.  For a fixed Kd reducing the off-rate will also reduce the on-rate and binding kinetics have to be seen in the broader context of distribution.  If the binding is faster than distribution then on-rates and off-rates become irrelevant, except to the extent to which they determine Kd.

At a conference like this, I'd have hoped to see something along the lines of 'unanswered questions and unsolved problems'.  How predictive are fragment properties of the properties of structurally-elaborated molecules?  Just how strong is the correlation of promiscuity with lipophilicity?  Is getting structures for protein-ligand complexes still a bottleneck?

So best wishes for an enjoyable and successful meeting.  Make sure to test the wits of the speakers with some tough questions.  It's character-building for them and loads of fun for everybody else.  Meanwhile here in Brasil it is 'imunização para insetos' at my place and I've cooked up a local response to enthalpy optimisation: Termodinâmica Macumba.



Tuesday, 18 September 2012

Ligand deconstruction

Molecular interactions are of interest in molecular design because the functional behaviour of a compound is determined by how strongly its molecules interact with the different environments in which they exist.  Although I'm talking primarily about non-covalent interactions, reversible covalent bond formation, for example between the catalytic cysteine of a protease and nitrile carbon, can also fit into this framework. Molecular design can be hypothesis-driven or prediction-driven and you'll have guessed from my last post which approach I favor. Hopefully at some point in the future we'll be able to predict well enough to do molecular design and when we do get there I think that we'll find that the models will have a strong physical basis.  Until then,hypothesis-driven molecular design will continue to have an important role.

Molecular interactions are relevant to both prediction-driven and hypothesis-driven molecular design. Design hypotheses are often framed in terms of molecular interactions and a predictive model for affinity that fails to capture the physics of molecular interactions will choke when used outside narrowly-defined congeneric series. Although we think of affinity in terms of contacts between protein and ligand, it is important to remember that the contribution of a particular contact to affinity is not strictly an experimental observable.

In FBDD we think of ligands in terms of their component fragments and are particularly interested in the extent to which the properties of fragments determine the properties of structurally elaborated compounds. Comparing the affinties of ligands with the fragments from which they might have been derived is one way in which this question can be addressed and one will occasionally encounter the term 'deconstruction' in the FBDD field. Just as you need to be careful how you link fragments when assembling a ligand, you also need to be careful how you decompose a ligand into component fragments. In this blog post I'm going to use a well-known deconstruction study to highlight some of the things that you need to think about when deconstructing.



I like to think of molecular interactions in terms of generic atom types such as 'neutral hydrogen bond acceptor' or cationic hydrogen bond donor'. This is a pharmacophoric view of molecular recognition which is also relevant to bioisosterism and scaffold-hopping. Those who take a more physical view of molecular recognition would say that pharmacophoric atom-typing is just cheminformatics and somewhat uncouth. However, you can capture a lot of physics with atom-typing and it's not like placing atomic charges on nuclei is such great physics anyway...

When deconstructing a ligand molecule you want to minimise changes to the way in which a binding site might see atoms in the ligand. For example breaking the carbon-nitrogen bond of an amide and adding hydrogens is not a great idea because you’ll turn hydrogen bond donor into a cation (at physiological pH) and a strong hydrogen bond acceptor into a weaker one. Deconstructions that add or remove hydrogen atoms from nitrogen or oxygen atoms are usually not a good idea.   In the featured ligand deconstruction study, fragments 2 and 3 were derived structure from 1. The acyl sulfonamide group of 1 would be expected to be predominantly deprotonated under normal physiological conditions (a pKa of 5.4 has been reported for sulfacetamide). In contrast, fragment 2 would be expected to be predominantly neutral at normal physiological pH (benzenesulfonamide pKa is 10.1). This means that the deconstruction of the acylsulfonamide transforms an anionic nitrogen into a neutral one that is bonded to a donor hydrogen.  This makes it difficult to draw conclusions from the observation that the fragment does not bind to the target. Is the interaction between the relevant part of the parent ligand very weak or has the deconstruction changed the pharmacophoric character of interacting atoms?

The deconstruction of 1 to 3 effectively creates a cationic center (a pKa of 5.1 has been reported for dimethylaniline) and having two nitrogen atoms in the piperazine ring does introduce complications.  Two pKa values are observed for piperazine and in a recent study these were found to be 9.7 and 5.5 at 298K.  This tells us that protonation of one of the nitrogen atoms makes it more difficult to protonate the other one (which makes sense).  These measured pKa values also tell us that piperazine will exist predominantly as a monocation at normal physiological pH and the corresponding values for 1,4-dimethylpiperazine are 8.4 and 3.8.  If you take a look at the source that I used for the benzenesulfonamide pKa, you'll see that attaching a phenyl ring to a carbon that is bonded to a basic nitrogen will make that nitrogen less basic by about one log unit. Bringing this all together for compound 1 suggests that protonation of piperazine will occur preferentially at the left hand nitrogen (see figure above) and the that the relevant pKa will be about 7.4. Deconstruction to fragment 3 is expected to shift the preferred site of protonation to the other nitrogen.

So what is the protonation state of piperazine when compound 1 binds to its target? The closeness of the likely pKa to normal physiological pH makes it difficult to say and if the relevant proton/lone-pair is directed away from the protein surface then cationic and neutral forms may have similar affinity. At this point, I should mention that I couldn't find the value(s) of the pH at which the NMR experiments were performed (if this information is indeed there, I'll invoke the 'reading PDF on my computer defense') and the information really needs to be communicated in a study such as this one.

I'll mention a couple of other deconstructions to illustrate the point that changing an element may sometimes result in less perturbation of the relevant substructures.  Fragment 4 gets round the problem of deconstruction shifting the preferred site of protonation. The nitrogen atom in the parent molecule that is mutated into carbon will be a weak hydrogen bond acceptor because it is linked directly to an aromatic ring. It can be argued that mutating a weak hydrogen bond acceptor into a hydrophobic atom represents a smaller perturbation than mutating it into a cationic center. However, piperidine is more basic than piperazine so there will be less neutral form (which may or may not be relevant). Deconstruction to fragment 5 preserves (actually is likely to strengthen) the hydrogen bond acceptor character of the less basic piperazine nitrogen but is likely to decrease the amount of cationic form because morpholine is less basic than piperazine.

Hopefully this will have got you thinking in a bit more depth about ligand deconstruction and I'll finish off with a cartoon of how we might use deconstruction in lead optimisation.  First we check that we can can actually measure affinity for a fragment that is obtained by deconstructing the lead compound.  Then we assemble SAR (could be a good way to explore bioisosteric replacements) before incorporating the best fragments into the lead structure.  Essentially, the fragment assay allows us to assemble SAR in a more accessible region of chemical space. 



Literature cited 

Barelier, Pons, Marcillat, Lancelin, Krimm, Fragment-Based Deconstruction of Bcl-xL Inhibitors.  J. Med. Chem. 2010, 53, 2577-2588. DOI

Cabot, Fuguet, Ràfols, Rosés, Fast high-throughput method for the determination of acidity constants by capillary electrophoresis. II. Acidic internal standards. J. Chromatography A 2010, 1217, 8340–8345. DOI

Milletti, Storchi, Goracci, Bendels, Wagner, Kansy, Cruciani, Extending pKa prediction accuracy: High-throughput pKa measurements to understand pKa modulation of new chemical series. Eur. J. Med. Chem. 2010, 45, 4270-4279. DOI

Fickling, Fischer, Mann, Packer & Vaughan, Hammett Substituent Constants for Electron-withdrawing Substituents : Dissociation of Phenols, Anilinium Ions and Dimethylanilinium Ions. JACS 1959,81, 4226-4230. DOI

Khalili, Henni, East, pKa Values of Some Piperazines at (298, 303, 313, and 323) K. J. Chem. Eng. Data 2009, 54, 2914-2917. DOI



 

    

Monday, 20 August 2012

QSAR: Nailed to its perch?

I must confess that I’ve never been a big fan of QSAR. When I started in Pharma 24 years ago, QSAR was seen as as something that would solve all our problems and, over the years, a number of other panaceas would follow in its wake. I find it useful to classify molecular design as either hypothesis-driven or prediction-driven and will discuss this a bit more in a future post. QSAR fits into the prediction-driven category and, to get you thinking a bit about the subject, I'll share a couple of slides from my RACI talk last December.



So EuroQSAR is due to happen again and this time there'll be a session to commemorate QSAR's founding father Corwin Hansch, who died last year.  So will 'Grand Challenges for QSAR' deliver?  Were I going to be there, I'd be checking out Maggiora's talk (Activity Cliffs, Information Theory, and QSAR) since people in the field really need to start thinking more about QSAR in terms of relationships between structures.  Although it's not part of the Hansch session, I'd also be checking out 'The Power of Matched Pairs in Drug Design' by my good friend (and former colleague) Jonas Boström since Matched Molecular Pairs represent one way to recognise and articulate relationships between structures.  And of course I wouldn't miss the Hansch Awardee's talk, the title of which reminded me of an Austrian who struggled, although you won't find that one stocked in the local book shops...

I would like to have seen something on training set design and validation in the 'Grand Challenges' session.  Generally building and validating multivariate models work best when the compounds are distributed evenly in the relevant descriptor space.  Clustering in descriptor space can result in validation giving an optimistic view of model quality and that's one way to end up over-fitting yout data.  Maybe this was one Grand Challenge that the Organising Committee just didn't have the stomach for...

So that's all from me for now.  Why not print out 'QSAR: dead or alive?' (it infuriates those who would seek to lead your opinion) to read on the plane and think up some nasty questions on validation for the experts while waiting in Passkontrolle?

Literature Cited

Doweyko, QSAR: dead or alive? JCAMD 2008, 22, 81-89 DOI