Han Tang, Alexander Tong, Wouter Boomsma
01.10.2026 · Gubra Meeting
Background

Diffusion models learn to generate structures by first adding noise to examples, then learning to reconstruct them.
Method
Forward. We play the animation backwards to show the side chain breaking down into noise, losing its structure and atom identities.
Reverse. Starting from noise, the model builds a side chain by generating atom types, positions and bonds. In this example, it reconstructs serine.
Method
The backbone and surrounding structure stay fixed while UNAAGI removes the side chain beyond Cα and rebuilds it from noise.
The model generates atom types, positions and bonds together. Because it works directly with atoms, it can build both canonical and non-canonical side chains without choosing from a predefined list of amino acids.
Benchmark


The study tested a set of 40 amino acids, replacing one residue at a time and measuring the change in binding (ΔΔG). Negative values mean stronger binding than the original peptide.
Panels A and F from Rogers JM, Passioura T, Suga H. “Nonproteinogenic deep mutational scanning of linear and cyclic peptides.” PNAS 2018;115(43):10959–10964. Open access, CC BY-NC-ND.
Benchmark


Red means stronger binding. Blue means weaker binding. Values show the change from the original peptide (ΔΔG in kcal/mol).
We count substitutions with ΔΔG < 0.5 kcal/mol as tolerated, allowing a small loss in binding. We test whether UNAAGI proposes these substitutions without access to the assay results.
Panels E and G, same source.
Sampling
PUMA bound to MCL-1. The animation zooms in on Leu141, shown in green.
We remove the side chain beyond Cα, keeping the backbone and surrounding structure fixed.
UNAAGI uses the surrounding atoms to generate a replacement side chain. It receives no peptide sequence, assay results or original residue identity.
Sampling · PUMA 144
Sampling · CP2 Thr13 and PUMA Ala139
Restrictive positions
Restrictive positions · and where it falls short
Glu153 tolerates 11 substitutions and the model produced 9 of them, missing only 2-aminooctanoic acid and tert-butylalanine — both of which it does build elsewhere.
Tyr152 tolerates 7, and the model produced exactly one — tryptophan, which happens to be the best of the seven.
The gap is a specific one. Five of the six it misses at Tyr152 are bulky aromatic non-canonicals — benzothienyl-alanine, 2- and 3-thienyl-alanine, O-methyl-tyrosine and 2-naphthyl-alanine; the sixth is 2-aminoheptanoic acid, which the assay only just tolerates (+0.47). The unconditional prior rarely proposes the aromatic class anywhere. It is the clearest target for the conditional model, and the reason we are building one.
Validity
Across every residue shown in this deck we checked the generated molecule rather than the label: element-by-element bond lengths against ideal values, ring planarity, and non-bonded clashes — then the bond orders, which the graph-isomorphism labeller ignores.
Aromatic rings come out planar and correctly conjugated; aliphatic chains come out saturated; amide and carboxylate groups carry their double bonds in the right place.
| Check | Result |
|---|---|
| Side-chain bond orders | 51 / 53 |
| Every bond within 10% of ideal | 45 / 53 |
| No non-bonded contact < 2.45 Å | 51 / 53 |
| Aromatic rings planar and conjugated | 13 / 13 |
| Saturated rings puckered | 0 / 2 |
The one real failure is worth naming. Cyclopentyl-alanine came out with two double bonds in its five-membered ring — a flat cyclopentadiene rather than a puckered cyclopentane. The labeller matched it anyway, because it compares skeletons and not bond orders. Both draws of it, at two separate positions, failed the same way. Saturated rings are a known weak point and we report it as one.
Result
Higher bars mean closer agreement between each model’s ranking and the measured effects on binding.
UNAAGI leads for non-canonical substitutions in both peptides. Across all substitutions, it leads on PUMA and performs similarly to AF3 iPTM on CP2.
Sanity check
In progress
We are developing a model that estimates how well a chosen amino acid fits its local atomic environment. This will support scoring both canonical and non-canonical substitutions.
We plan to use assay results to favour substitutions with better measured binding. New data will guide model updates and the next round of candidate generation.