Author Archives: Fergus Imrie

The First OpenBind Blind Challenge

At OpenBind, we’re working to generate large-scale, open experimental datasets of protein-ligand structures and binding measurements to support the development and evaluation of computational methods for drug discovery. One thing I’m particularly excited about is using these data not just for retrospective benchmarking, but for prospective, blind evaluation.

So, I’m very pleased that we’re launching the first OpenBind blind challenge on 14 October 2026, in partnership with OpenADMET and the ASAP Discovery Consortium. The interim submission deadline is 11 November, with final submissions due by 16 December 2026 at 23:59 UTC.

The challenge asks participants to predict protein-ligand complexes of 400 compounds bound to Zika virus NS2B-NS3 protease, ranging from small fragments to larger, lead-like molecules.

A key challenge will be using the structural information that already exists to accurately predict the binding modes of new, unseen ligands. There are already a number of publicly available structures of Zika NS2B-NS3 protease in the PDB that participants can use.

We’re also setting a high bar for success. While many protein-ligand benchmarks use an RMSD threshold of 2 Å, we’re tightening this to 1.5 Å, along with additional checks on interaction accuracy (lDDT-PLI > 0.8) and physical validity (PoseBusters validity checks). I’m particularly excited to see how docking, cofolding, and other approaches compare!

If you work on protein-ligand structure prediction, I’d strongly encourage you to join the challenge. You can find all the details, including how to register and join the challenge Discord, in the OpenBind announcement.

Fragment-to-Lead Successes in 2024 – 10th Anniversary Edition

In what I have to admit is now becoming an annual tradition ([2023] [2019]), I’d like to highlight the 2024 edition of the fragment-to-lead success stories, published in J. Med. Chem. at the end of 2025 [Paper].

Continue reading →

Fragment-to-Lead Successes in 2023

Back in 2021, I highlighted the annual fragment-to-lead (F2L) success stories from 2019 [Blog post] [Paper]. This is one of my favourite annual publications, and I’m delighted to see that it’s still going strong. In this post, I’ll discuss the 2023 edition that was published in at the start of 2025 [Paper].

Continue reading →

Reflections on GRC CADD 2025: A Week of Insight, Innovation, and Baseball

Henry

Back in July, some very lucky OPIGlets ventured across the pond to discover life in Southern Maine (and Boston!). For someone visiting Boston for the first time, no trip would be complete without a Red Sox game—a thoroughly enjoyable highlight (see Figure 1). While we were there, we also went to Gordon Research Conference (GRC) on Computer Aided Drug Design (CADD).

A flock of OPIGlets taking in the Fenway Park experience at a Red Sox game.
Continue reading →

Le Tour de Farce v11.0

“They don’t make them like they used to!”

With much experience of all things farcical, it was my delight to have returned just in time for the 2024 edition of OPIG’s Tour de Farce, which took place on 11th July. This year’s route was 8 miles long and encompassed four of the finest establishments Oxford has to offer (nothing “unusually conservative” to see here Eoin).

Continue reading →

Fragment-to-Lead Successes in 2019

In this blogpost, I want to highlight the excellent work by Jahnke and collaborators. For the past 5 years, they have published an annual perspective covering fragment-to-lead success stories from the previous year. Very helpfully, their work includes a table detailing the hit fragment(s) and lead molecule, together with key experimental results and parameters.

Continue reading →

NeurIPS 2020: Chemistry / Biology papers

Another blog post, another look at accepted papers for a major ML conference. NeurIPS joins the other major machine learning conferences (and others) in moving virtual this year, running from 6th – 12th December 2020. In a continuation of past posts (ICML 2020, NeurIPS 2019), I will highlight several of potential interest to the chem-/bio-informatics communities

The list of accepted papers can be found here, with 1,903 papers accepted out of 9,467 submissions (20% acceptance rate).

In addition to the main conference, there are several workshops highly related to the type of research undertaken in OPIG: Machine Learning in Structural Biology and Machine Learning for Molecules.

The usual caveat: given the large number of papers, these were selected either by “accident” (i.e. I stumbled across them in one way or another) or through a basic search (e.g. Ctrl+f “molecule”). If you find any I have missed, please reach out and I will update accordingly.

Continue reading →

Learning from Biased Datasets

Both the beauty and the downfall of learning-based methods is that the data used for training will largely determine the quality of any model or system.

While there have been numerous algorithmic advances in recent years, the most successful applications of machine learning have been in areas where either (i) you can generate your own data in a fully understood environment (e.g. AlphaGo/AlphaZero), or (ii) data is so abundant that you’re essentially training on “everything” (e.g. GPT2/3, CNNs trained on ImageNet).

This covers only a narrow range of applications, with most data not falling into one of these two categories. Unfortunately, when this is true (and even sometimes when you are in one of those rare cases) your data is almost certainly biased – you just may or may not know it.

Continue reading →

ICML 2020: Chemistry / Biology papers

ICML is one of the largest machine learning conferences and, like many other conferences this year, is running virtually from 12th – 18th July.

The list of accepted papers can be found here, with 1,088 papers accepted out of 4,990 submissions (22% acceptance rate). Similar to my post on NeurIPS 2019 papers, I will highlight several of potential interest to the chem-/bio-informatics communities. As before, given the large number of papers, these were selected either by “accident” (i.e. I stumbled across them in one way or another) or through a basic search (e.g. Ctrl+f “molecule”).

Continue reading →