AI can remove typing. It cannot remove thinking.​

Students new to programming may wonder why they even need to learn to code any more. They can ask the reasonable question: “Why can’t we just use AI for all our programming?” It’s a fair question. The Genie is out of the bottle. We aren’t putting it back. It absolutely does have its uses. But it also can’t be trusted.

For a new programmer, the important skill is no longer remembering syntax, nor is it learning how to make an AI produce code. It’s learning how to solve problems in a systematic way that they understand well enough to trust.

When getting started with AI, use it to help think through the problem, break the work into manageable pieces, elaborate how the pieces interact and only then start to code.

There’s the old adage “Garbage in, garbage out”. If you aren’t able to articulate your problem in an AI-friendly way, we’re well into garbage out territory. The other part of the problem is whilst AI will certainly generate you some code, if you don’t understand what it’s giving back, you can’t trust it to be correct. If you carry on just accepting what it’s saying, it’s like going on a date in a foreign country and putting all your faith in your translation app. Sooner or later you’re unwittingly going to say something that will earn you a slap.

So, instead of just telling it what you want, try:

Continue reading →

Finally solving drug discovery with a Fly

As I was sitting down this morning, trying to enjoy my rather excellent breakfast, I was suddenly attacked by a fly. Unfortunately for me, my lazy hand swats proved no match for the nimble little thing, who was evidently pretty hungry too. And as I watched the acrobatics this tiny creature could effortlessly pull off in pursuit of my sandwich, I was left wondering: if only we could put these skills to better use.

Open-sourcing a fly

Luckily, I’m not the only one who has been staring at a fly and wondering. A couple of weeks ago, HHMI Janelia, the University of Cambridge and Google Research published a paper in the Cell describing the complete wiring diagram of a male fruit fly’s central nervous system: the brain and the nerve cord, the whole thing, down to individual synapses. To be fair, this isn’t the first time someone has done this. The FlyWire project gave us a complete female fly brain back in 2024, and Janelia had already released half a female brain in 2020. But this one comes with the body wiring attached, and, for whatever reason, this is the one the internet decided to notice.

What the internet did with it

Now, the video games. My question at the top was whether these skills could be put to better use, so here they are:

Continue reading →

One Year Down: Wisdom from a Former OPIG Newbie

With a new cohort of PhD students arriving imminently and my first year coming to a close, I am about to achieve an important milestone: I will no longer be the newest full-time OPIGlet in the group!

Naturally, this makes me extremely qualified to impart wisdom.

In all seriousness, the first year of a PhD has been a strange and wonderful adjustment. I have learned a lot about nanobodies, of course, but I have also learned quite a bit about how to actually do a PhD. So, for the incoming OPIGlets—or anyone else embarking on their first year—here are a few things I wish I had fully appreciated when I started.

Give some structure to all that freedom

One of the strangest things about starting a PhD is suddenly having an almost completely self-guided schedule. It is not something most of us have had to deal with before. Undergrad comes with lectures, tutorials, deadlines, and exams. Most jobs come with meetings, working hours, and someone telling you what needs to be done.

Continue reading →

Is the race for AGI a scam? Are the tech bros con men or visionaries?

To find out, we formally introduce the Stinkometer

In research, it can be tough even for experts to tell fact from fiction. This is especially true for more speculative fields like AI. It can feel like every model is “State-of-the-Art” and that AGI (whatever that means) will be “6 months away” for at least the next 6 months. How do we know we can trust these people?

What is a visionary? What is a con man?

Visionaries sell a fringe idea of the future, usually with strong self-belief and charisma. Con men are visionaries who don’t believe what they’re saying. It’s hard to tell the difference because we can’t read minds and can’t see the future.

I solved this by adapting the insightful “New Political Compass” from Harper O’Connor, an American political YouTuber. Harper found a similar problem. He’s active in local politics and has to decide who he should build alliances with to tackle certain issues. Interestingly, the “left-right” axis wasn’t a helpful guide. Most people he canvassed were reasonable but didn’t follow politics closely enough to have a robust ideology. He realised that someone’s psychology can be more important than their stated ideology.

Therefore, it was more useful to ask: Am I talking to a reasonable person? Here, I expand his framework into the Belief Compass (or the “Stinkometer”). We can answer our question by answering three simpler ones:

Continue reading →

Predicting ADME Properties with Machine Learning: 82% of the Performance From Two Descriptors

A Nature survey of ~1500 scientists reported that more than 70% had failed to reproduce another scientist’s experiments, and 50% had failed to replicate their own (Baker, 2016). This study wasn’t specific to machine learning, but the crisis has its own flavour in computational drug discovery.

One aspect of this is that available datasets have known quality issues, and results built on them can be fragile. MoleculeNet and the Therapeutic Data Commons (TDC) opened drug discovery to a wider ML community, but they’re no longer sufficient for driving further advances (Wognum et al., 2024).

Amongst a number of issues, 71% of molecules in one MoleculeNet dataset (BACE) contained at least one undefined stereocenter (a point where the same atoms can sit in two different 3D arrangements), making it unclear what chemical entity is actually being modelled (Li et al., 2026).

The stakes are high. In 2017, a research group found their cancer-target inhibitor inactive from one vendor and highly active from another. Eventually they traced this to vendors selling different mixtures of the compound’s two 3D forms where only one was active on the target and mechanism they were investigating (Baker, 2017).

Most approved pharmaceuticals are relatively small chemical molecules, typically weighing under 900 g/mol, with most of the rest being biologics. Whilst in Oxford in OPIG over the summer as a UNIQ+ student, I focused on ADME, an aspect of early stage drug discovery, where assays are measured in vitro as a stand-in to predict how a compound will behave in vivo (in the human body) before clinical trials. Absorption (does it enter the body), Distribution (where does it go), Metabolism (how quickly it’s transformed) and Excretion (how quickly it’s eliminated) are processes critical to whether a candidate succeeds. 

Continue reading →

Project Cybersyn: Bayesian Socialist Algocracy in 1970s Chile

Palantir and Peter Thiel have found themselves in the headlines a lot this year, very often in the same sentence as the word “surveillance”. The idea behind Palantir is a simple one: give one company, or one control room, a real-time algorithmic view of how a government or an economy is actually functioning, and let them help run things better.

That idea is not new. In 1971, Salvador Allende’s socialist government in Chile built almost exactly that system, out of telex machines and a room full of fibreglass chairs, on practically no budget. There’s evidence it worked, until a coup destroyed it before anyone found out whether it would have worked at scale. This was known as “Project Cybersyn”.

A modern reconstruction of Project Cybersyn’s hexagonal operations room, with seven chairs arranged in a ring and display screens on the walls.
The space-age themed Opsroom, where decisions were made in response to Bayesian forecasting.
Continue reading →

The cost of a better pose: balancing GNINA sampling and runtime

If you have ever set up a docking experiment or tuned a docking workflow for GNINA, you may have found yourself asking:

What are the “best” settings for running GNINA?

Unfortunately, there is no single objectively correct answer. The optimal settings will depend on the input data, the goal of the docking experiment, and the computational resources available. One reasonable approach is simply to use the standard GNINA settings, dock each molecule once using a single conformer, and leave it at that. The default settings already perform well in many cases. But that does not mean performance cannot be improved!

Continue reading →

Beyond the Vaccine: AI’s Expanding Role in LNP mediated Drug Delivery

Lipid nanoparticles (LNPs) have evolved from a specialised drug-delivery technology into a cornerstone of modern medicine. Their success became evident during the COVID-19 pandemic, enabling the delivery of the mRNA used in the Pfizer-BioNTech and Moderna vaccines. However, their applications extend far beyond vaccination, as LNPs can also deliver siRNA, plasmid DNA and gene-editing machinery, creating opportunities to treat genetic diseases, cancer and many other conditions.

Yet a fundamental challenge remains: how do we design an LNP that delivers the right payload, to the right cells, in the right place?

Continue reading →

Claude now watermarks its text: here’s what is actually happening under the hood

On 11 August Anthropic announced that every Claude model released after 2 August 2026 will watermark the text it produces. Not just in the chat app: the watermark lives at the model level, so it is there whether the text came through the API, Claude Code, Cowork, or anything else built on top. A follow-up post a few days later explained the mechanism, and the short version is that it’s a version of DeepMind’s SynthID-Text, which in turn descends from a scheme Scott Aaronson sketched out while at OpenAI in 2022.

Predictably it has caused a lot of confusion, especially in the academic circles (“is it zero-width characters?”, “can I strip it with a regex?”, “does this mean Turnitin finally works?”), most of which comes from not knowing how an LLM actually turns a probability distribution into words. So this post starts there, builds the watermark up from the sampler, and then tries to be honest about what it does and doesn’t mean for people who write papers and mark essays for a living.

Nothing here is hidden: the watermark is not metadata, not invisible Unicode, not a hidden token. It is a statistical pattern in which words were chosen, which is exactly why it survives copy-paste and exactly why it fades when you rewrite.

Continue reading →