Finally solving drug discovery with a Fly

As I was sitting down this morning, trying to enjoy my rather excellent breakfast, I was suddenly attacked by a fly. Unfortunately for me, my lazy hand swats proved no match for the nimble little thing, who was evidently pretty hungry too. And as I watched the acrobatics this tiny creature could effortlessly pull off in pursuit of my sandwich, I was left wondering: if only we could put these skills to better use.

Open-sourcing a fly

Luckily, I’m not the only one who has been staring at a fly and wondering. A couple of weeks ago, HHMI Janelia, the University of Cambridge and Google Research published a paper in the Cell describing the complete wiring diagram of a male fruit fly’s central nervous system: the brain and the nerve cord, the whole thing, down to individual synapses. To be fair, this isn’t the first time someone has done this. The FlyWire project gave us a complete female fly brain back in 2024, and Janelia had already released half a female brain in 2020. But this one comes with the body wiring attached, and, for whatever reason, this is the one the internet decided to notice.

What the internet did with it

Now, the video games. My question at the top was whether these skills could be put to better use, so here they are:

Beat Saber

Crypto Trading

People even started feeling bad for this fly, so one developer began building it its own heaven: an open world full of grass, trees and an endless supply of fruit, where it can just fly around and eat whatever it wants, forever. Honestly, not a bad retirement.

Taping a molecule to the fly

All of which brings me back to breakfast. The one thing the fly is undeniably, infuriatingly good at is flying: a six-degrees-of-freedom control problem, three translations and three rotations, solved in real time by a brain the size of a poppy seed. And it turns out there is another six-degrees-of-freedom problem that a great deal of expensive software spends a great deal of expensive compute on: docking. Rigid docking is this: here is a protein, here is a drug molecule, find the pose in which the molecule sits in the binding pocket. Three translations, three rotations. If you squint, it is the problem the fly solves on the way to my sandwich, just with a worse view.

So I taped a molecule to the fly.

More precisely, I built a small simulator where the “body” the fly controls is a rigid drug molecule and the “world” is the inside of a protein. The molecule is staurosporine, a famously indiscriminate kinase inhibitor that has been crystallised bound to something like 86 different proteins, which means there are 86 known correct answers lying around in the Protein Data Bank. The protein is CDK2, a cell-cycle kinase, from PDB entry 1AQ1. Each step, the fly nudges the molecule by up to half an ångström and turns it by up to about nine degrees. It is rewarded for getting closer to the crystal pose, penalised for pushing atoms into the protein, and if it gets within 2 Å of the real thing, it has docked.

The wiring in the middle is the real thing: the central-brain part of MaleCNS, 38,128 neurons and 3 million synapses, with each neuron’s outgoing synapses made excitatory or inhibitory according to its predicted neurotransmitter. The fly’s senses go into its 4,656 sensory neurons, the ones that in life would be smelling and touching things: a “smell” that points toward the pocket and says how far away it is, a “touch” that says how close the nearest protein atom is and from which direction. The motor commands, six little nudges, are read out of its 1,312 descending neurons, the cells that in life carry orders from the brain to the wings and legs. Nothing in between is trained. The only things that learn are the two thin layers at the edges, which is the same trick as the Doom fly, and training is ordinary reinforcement learning: about a hundred thousand nudges, half an hour on a laptop.

It did not work the first time. Or the second. For most of a day the fly’s preferred strategy was to grab the molecule and leave, ending up a hundred ångströms from the protein, which in molecular terms is the next postcode. Some of that was my fault: my first version of the protein was a maze the molecule had to tunnel through, and no amount of wiring fixes a bad map. But the interesting failure was in the read-out. The descending neurons fire away at a steady hum whatever the fly is sensing, and the part of their activity that actually depends on the senses turned out to be about 0.1% of the total. The connectome was doing its job; I just had the volume knob wrong. Once each neuron was read relative to its own usual hum, that fraction went to 91%, and twenty thousand nudges later the fly docked staurosporine into CDK2 from every starting position I gave it, including ones where the molecule began nine ångströms out and upside down.

Here it is doing that.

So, have we finally solved drug discovery with a fly? Obviously not. What we have is a dead fly’s brain, untrained and unbothered, parking a molecule about as well as a network built for the job, plus a to-do list that starts with “learn what a pocket is” and ends somewhere around “chemistry”. Still, every demo in this post is made-up physics bolted onto a real map, and it keeps being the map that works. Eighty-five more kinases are sitting in the Protein Data Bank waiting for a pilot. After a hundred thousand attempts at parking, the fly has earned a snack. It can have the sandwich.

One Year Down: Wisdom from a Former OPIG Newbie

With a new cohort of PhD students arriving imminently and my first year coming to a close, I am about to achieve an important milestone: I will no longer be the newest full-time OPIGlet in the group!

Naturally, this makes me extremely qualified to impart wisdom.

In all seriousness, the first year of a PhD has been a strange and wonderful adjustment. I have learned a lot about nanobodies, of course, but I have also learned quite a bit about how to actually do a PhD. So, for the incoming OPIGlets—or anyone else embarking on their first year—here are a few things I wish I had fully appreciated when I started.

Give some structure to all that freedom

One of the strangest things about starting a PhD is suddenly having an almost completely self-guided schedule. It is not something most of us have had to deal with before. Undergrad comes with lectures, tutorials, deadlines, and exams. Most jobs come with meetings, working hours, and someone telling you what needs to be done.

Continue reading

Is the race for AGI a scam? Are the tech bros con men or visionaries?

To find out, we formally introduce the Stinkometer

In research, it can be tough even for experts to tell fact from fiction. This is especially true for more speculative fields like AI. It can feel like every model is “State-of-the-Art” and that AGI (whatever that means) will be “6 months away” for at least the next 6 months. How do we know we can trust these people?

What is a visionary? What is a con man?

Visionaries sell a fringe idea of the future, usually with strong self-belief and charisma. Con men are visionaries who don’t believe what they’re saying. It’s hard to tell the difference because we can’t read minds and can’t see the future.

I solved this by adapting the insightful “New Political Compass” from Harper O’Connor, an American political YouTuber. Harper found a similar problem. He’s active in local politics and has to decide who he should build alliances with to tackle certain issues. Interestingly, the “left-right” axis wasn’t a helpful guide. Most people he canvassed were reasonable but didn’t follow politics closely enough to have a robust ideology. He realised that someone’s psychology can be more important than their stated ideology.

Therefore, it was more useful to ask: Am I talking to a reasonable person? Here, I expand his framework into the Belief Compass (or the “Stinkometer”). We can answer our question by answering three simpler ones:

Continue reading

Predicting ADME Properties with Machine Learning: 82% of the Performance From Two Descriptors

A Nature survey of ~1500 scientists reported that more than 70% had failed to reproduce another scientist’s experiments, and 50% had failed to replicate their own (Baker, 2016). This study wasn’t specific to machine learning, but the crisis has its own flavour in computational drug discovery.

One aspect of this is that available datasets have known quality issues, and results built on them can be fragile. MoleculeNet and the Therapeutic Data Commons (TDC) opened drug discovery to a wider ML community, but they’re no longer sufficient for driving further advances (Wognum et al., 2024).

Amongst a number of issues, 71% of molecules in one MoleculeNet dataset (BACE) contained at least one undefined stereocenter (a point where the same atoms can sit in two different 3D arrangements), making it unclear what chemical entity is actually being modelled (Li et al., 2026).

The stakes are high. In 2017, a research group found their cancer-target inhibitor inactive from one vendor and highly active from another. Eventually they traced this to vendors selling different mixtures of the compound’s two 3D forms where only one was active on the target and mechanism they were investigating (Baker, 2017).

Most approved pharmaceuticals are relatively small chemical molecules, typically weighing under 900 g/mol, with most of the rest being biologics. Whilst in Oxford in OPIG over the summer as a UNIQ+ student, I focused on ADME, an aspect of early stage drug discovery, where assays are measured in vitro as a stand-in to predict how a compound will behave in vivo (in the human body) before clinical trials. Absorption (does it enter the body), Distribution (where does it go), Metabolism (how quickly it’s transformed) and Excretion (how quickly it’s eliminated) are processes critical to whether a candidate succeeds. 

Continue reading

Project Cybersyn: Bayesian Socialist Algocracy in 1970s Chile

Palantir and Peter Thiel have found themselves in the headlines a lot this year, very often in the same sentence as the word “surveillance”. The idea behind Palantir is a simple one: give one company, or one control room, a real-time algorithmic view of how a government or an economy is actually functioning, and let them help run things better.

That idea is not new. In 1971, Salvador Allende’s socialist government in Chile built almost exactly that system, out of telex machines and a room full of fibreglass chairs, on practically no budget. There’s evidence it worked, until a coup destroyed it before anyone found out whether it would have worked at scale. This was known as “Project Cybersyn”.

A modern reconstruction of Project Cybersyn’s hexagonal operations room, with seven chairs arranged in a ring and display screens on the walls.
The space-age themed Opsroom, where decisions were made in response to Bayesian forecasting.
Continue reading

The cost of a better pose: balancing GNINA sampling and runtime

If you have ever set up a docking experiment or tuned a docking workflow for GNINA, you may have found yourself asking:

What are the “best” settings for running GNINA?

Unfortunately, there is no single objectively correct answer. The optimal settings will depend on the input data, the goal of the docking experiment, and the computational resources available. One reasonable approach is simply to use the standard GNINA settings, dock each molecule once using a single conformer, and leave it at that. The default settings already perform well in many cases. But that does not mean performance cannot be improved!

Continue reading

Beyond the Vaccine: AI’s Expanding Role in LNP mediated Drug Delivery

Lipid nanoparticles (LNPs) have evolved from a specialised drug-delivery technology into a cornerstone of modern medicine. Their success became evident during the COVID-19 pandemic, enabling the delivery of the mRNA used in the Pfizer-BioNTech and Moderna vaccines. However, their applications extend far beyond vaccination, as LNPs can also deliver siRNA, plasmid DNA and gene-editing machinery, creating opportunities to treat genetic diseases, cancer and many other conditions.

Yet a fundamental challenge remains: how do we design an LNP that delivers the right payload, to the right cells, in the right place?

Continue reading

Claude now watermarks its text: here’s what is actually happening under the hood

On 11 August Anthropic announced that every Claude model released after 2 August 2026 will watermark the text it produces. Not just in the chat app: the watermark lives at the model level, so it is there whether the text came through the API, Claude Code, Cowork, or anything else built on top. A follow-up post a few days later explained the mechanism, and the short version is that it’s a version of DeepMind’s SynthID-Text, which in turn descends from a scheme Scott Aaronson sketched out while at OpenAI in 2022.

Predictably it has caused a lot of confusion, especially in the academic circles (“is it zero-width characters?”, “can I strip it with a regex?”, “does this mean Turnitin finally works?”), most of which comes from not knowing how an LLM actually turns a probability distribution into words. So this post starts there, builds the watermark up from the sampler, and then tries to be honest about what it does and doesn’t mean for people who write papers and mark essays for a living.

Nothing here is hidden: the watermark is not metadata, not invisible Unicode, not a hidden token. It is a statistical pattern in which words were chosen, which is exactly why it survives copy-paste and exactly why it fades when you rewrite.

Continue reading