Monthly Archives: July 2026

Diffusion Models Can’t Give You a Likelihood… So How Do We Score Inverse-Folded Sequences?

I’ve been working with inverse folding models for sequence design for a while now, and one question kept coming up. How would you score a sequence’s likelihood under these models? It turns out the answer depends heavily on whether the model is autoregressive or diffusion-based.

Autoregressive models, like GPT-style language models, make this straightforward. You can calculate the probability of a sequence directly, by breaking it into one prediction per position and multiplying them together. Diffusion models define a probability distribution over outputs too, but getting the likelihood of a particular output means accounting for many unobserved intermediate states, and that turns out to be much harder. This is why diffusion models typically rely on the Evidence Lower Bound, or ELBO, a computable estimate, rather than the exact likelihood itself.

The easy case: autoregressive models

An autoregressive model breaks the probability of a sequence into a product of conditional probabilities:

p(x1:n)=i=1np(xi|x<i)p(x_{1:n}) = \prod_{i=1}^{n} p(x_i \mid x_{<i})
Continue reading

The IMGT just became a little FAIRer

The IMGT have adopted a more permissive licence for their data. That is a really good thing. These days, they also have an API. But the service is somewhat hamstrung by usability issues. Also, detailed and digestible documentation for such a sprawling suite of databases and tools never just appears overnight, so this post includes a demo to help you get started.

Continue reading

Work Hard, Play Hard: Balancing Sport and Studies at Oxford

I have always been involved in sport, having played football for Bristol Rovers and at county level when I was younger. That continued after arriving at Oxford, although in several different forms. In my first year, I ran the London Marathon for Alzheimer’s Research UK. In my second, I rowed with my college, and more recently, I took up boxing, competing against Cambridge to pick up the illusive ‘Blue’ and accompanying blazer.

Trying to pursue sport alongside a PhD has not always been easy, but it has been one of the most enjoyable parts of my time at Oxford. It has helped me step away from work, manage stress and maintain a competitive outlet outside academia. At the same time, Oxford sport can become extremely demanding, so finding a balance is sometimes challenging!

There is always more work to do

One of the difficulties of doing a PhD is that the work never feels completely finished. There is always another paper to read, experiment to run, result to analyse or paragraph to improve. If you wait until everything is done before exercising, you may never leave your desk (which is sometimes the case).

Sport creates boundaries that research rarely creates for itself. Training begins at a fixed time, teammates expect you to be there and competitions cannot be rearranged around your workload. I have often gone to training feeling that I should have stayed at my desk, only to return less stressed and able to concentrate more clearly. Taking a break can feel unproductive, but spending another hour staring at the same problem is not always the best way to solve it.

What sport adds to academic life

Sport develops many of the same qualities that academic work demands. Training consistently requires discipline, progress is rarely linear and poor performances teach you to keep going when things are not working as you hoped. It also provides a competitive outlet and a clearer sense of progress than research often does, which can be valuable when your PhD feels slow or uncertain.

Just as importantly, sport introduces you to people beyond your department and gives you something meaningful outside your work. During a PhD, it is easy to let academic progress determine how you feel about yourself. Having another community and another part of your life that matters helps keep a failed experiment or unproductive week in perspective.

Of course, sport does not need to be justified entirely by the lessons it teaches. At the end of the day, it is something I enjoy doing.

When play becomes more work

Although having another ambitious goal alongside your studies can also be extremely rewarding, the difficulty is that sport at Oxford can quickly become a serious commitment. Chasing a varsity place can demand a great deal of training, recovery and focus. Sometimes sport may genuinely matter more than your studies (in the lead-up to a competition, for instance), and that is fine, provided it is a deliberate choice and that things are still kept in perspective, rather than something you have been swept into without considering the trade-offs.

Finding the balance

Balance does not mean dividing every day equally between work and sport, nor does it mean always putting your studies first. There will be periods when training deserves more attention and others when academic work must take priority.

Moderation does not mean lacking ambition. You can train seriously, chase difficult goals and care about winning while still maintaining perspective. Equally, working hard does not require allowing your PhD to consume every other part of your life.

Sport has made my time at Oxford busier and occasionally harder to manage, but it has also made it far more enjoyable. It has given me friendships, challenges and experiences that I would never have found through academic work alone. It may even have made me better at my PhD, but that is not the only reason it was worth doing.

App-ready databases from Python with SQLModel and FastAPI

Python is ubiquitous for those of us coming into programming from academic research backgrounds, both for generating scientific data and for analysing it. Of course, modern data-driven research only thrives when data is FAIR (findable, accessible, interoperable, and reusable) and traditional relational databases and APIs for interacting with them are a tried-and-true method of ensuring FAIRness. (And in the age of machine learning, the machines need things to be FAIR arguably even more than we do…)

“But I’m a data scientist/computational chemist/bioinformatician!” you protest. “I just (ab)use databases using Python, I don’t make them for someone else to use!” Understandable, but the barrier to entry may be much lower than you think, thanks to two related Python libraries: SQLModel and FastAPI. These powerful tools by developer Sebastián Ramírez (@tiangolo on GitHub) work in concert to make relational databases and APIs much easier to implement for Python natives. Let’s see how they can help…

Continue reading

Same science, different stories: writing papers vs writing grants

As a PhD student, you will write a lot of papers, and at a certain point, you will also start writing grants. It took me a moment to realise that these are not the same skill. While they draw on the same science and sometimes the same project, they are different genres with different rules. Treating a grant proposal like a mini-paper is one of the most common and avoidable ways in which people damage their own applications. Here’s what I’ve learnt so far, mostly through trial and error.

Step one: know the landscape before you write a word

Before you open your text editor, you need to research the funding opportunity you’re interested in. Find out when the deadline is, which documents are required, what makes you eligible or ineligible, and what are the mandatory requirements versus those desirable. It may sound obvious, but it’s surprisingly easy to miss a deadline or, even worse, to base your entire career plan on a fellowship that you don’t qualify for or don’t have the right documents for.

Not all funders play by the same rules. Read each charter carefully before assuming that your project fits.
Continue reading

OPIGlet’s First Conference: PEGS Boston 2026

Hello everyone!

For my very first blog post, I’m excited to share my experience attending and speaking at my first conference.

I was invited to present at the Protein & Antibody Engineering Summit (PEGS) Boston 2026 on behalf of OPIG. The conference ran from May 10–15, and like many antibodies-stream OPIGlets before me, I delivered the OPIG tools short course. This formed part of a three-hour session titled In silico and Machine Learning Tools for Antibody Design and Developability Predictions. The short courses took place on the Sunday, before the main conference officially began on Monday, so somewhat terrifyingly, I gave a conference talk before I had ever attended one.

I arrived in Boston on Saturday after a long day of travel and was pleasantly surprised by a free bus into downtown (thanks, Boston MBTA). After a much-needed sleep, I headed to the conference centre, checked into my hotel (thank you, PEGS team), and tried not to explode from nerves. Thankfully, the talk went well; there were plenty of questions, and I quickly settled into speaking in front of the group. With an audience of around 50 people—only slightly larger than an OPIG group meeting—it felt like familiar territory.

Continue reading

OPIG Retreat, 2026: Heatwave Edition

Last week, a sizable fraction of OPIG headed to “The Plough” near Bradford on Avon, Wiltshire, for our OPIG Retreat (a.k.a. “OPIGtreat”). Some of us travelled on Monday by train, encountering a biblical deluge, a darkness resembling a train tunnel but filled with pelting rain and bath-tubs of water washing down the train’s windows — and lightning strikes every 30s while changing in Bath Spa…

The red monster storm with black eye is what greeted us as we were arriving in Bath Spa. (Credit: UK Thunderstorm Updates: “This is what the satellite imagery is currently showing from the storm affecting Somerset/ Wiltshire. Cloud tops on this monster have exceeded 40,000ft….”)

Undeterred, and either shuttled by the fabulous Anita from Bradford on Avon station, or walking on foot, we arrived at our lovely destination and home for the week.

Continue reading