UNAAGI: Atom-Level Diffusion
for Generating Non-Canonical
Amino Acid Substitutions

Han Tang, Alexander Tong, Wouter Boomsma

01.10.2026  ·  Gubra Meeting

Background

Diffusion learns to build structures from noise

Forward and reverse SDE schematic

Diffusion models learn to generate structures by first adding noise to examples, then learning to reconstruct them.

Background

Method

UNAAGI illustration in motion

Position 7 of the CP2 macrocycle bound to KDM4A (PDB 5LY1); native residue serine. Backbone and pocket are held fixed throughout.

Forward. We play the animation backwards to show the side chain breaking down into noise, losing its structure and atom identities.

Reverse. Starting from noise, the model builds a side chain by generating atom types, positions and bonds. In this example, it reconstructs serine.

Method

Method

Generating side chains to fit their local environment

UNAAGI method schematic

The backbone and surrounding structure stay fixed while UNAAGI removes the side chain beyond Cα and rebuilds it from noise.

The model generates atom types, positions and bonds together. Because it works directly with atoms, it can build both canonical and non-canonical side chains without choosing from a predefined list of amino acids.

Method

Benchmark

Testing amino acid substitutions in two peptides

The 21 non-proteinogenic amino acids used
The 21 non-canonical amino acids tested include longer side chains, N-methylated residues and D-alanine.
CP2 macrocycle bound to KDM4A
CP2 is a cyclic peptide that binds KDM4A. PUMA is a helical peptide that binds MCL-1.

The study tested a set of 40 amino acids, replacing one residue at a time and measuring the change in binding (ΔΔG). Negative values mean stronger binding than the original peptide.

Panels A and F from Rogers JM, Passioura T, Suga H. “Nonproteinogenic deep mutational scanning of linear and cyclic peptides.” PNAS 2018;115(43):10959–10964. Open access, CC BY-NC-ND.

Benchmark

Benchmark

How substitutions affect binding

PUMA deep mutational scan heatmap
PUMA binding to MCL-1: 34 positions × 39 substitutions.
CP2 deep mutational scan heatmap
CP2 binding to KDM4A: 12 positions × 39 substitutions.

Red means stronger binding. Blue means weaker binding. Values show the change from the original peptide (ΔΔG in kcal/mol).

We count substitutions with ΔΔG < 0.5 kcal/mol as tolerated, allowing a small loss in binding. We test whether UNAAGI proposes these substitutions without access to the assay results.

Panels E and G, same source.

Benchmark

Sampling

Take one side chain out, and ask the model to rebuild it

PUMA bound to MCL-1. The animation zooms in on Leu141, shown in green.

We remove the side chain beyond Cα, keeping the backbone and surrounding structure fixed.

UNAAGI uses the surrounding atoms to generate a replacement side chain. It receives no peptide sequence, assay results or original residue identity.

Sampling

Sampling · PUMA 144

At a hydrophobic site, UNAAGI samples amino acids that improve binding, including non-canonical residues

Ala144 — wild typethe site
isoleucine−1.87
norleucine−1.86
phenylalanine−1.68
leucine−1.64
tryptophan−1.08
Sampling · PUMA 144

Sampling · CP2 Thr13 and PUMA Ala139

Several stabilising non-canonicals at a single position

CP2 Thr13the site
phenylalanine−0.96
histidine−0.47
norvaline−0.20
PUMA Ala139the site
leucine−0.80
tyrosine−0.20
norleucine−0.04
Sampling · two more positions

Restrictive positions

UNAAGI finds tolerated substitutions at restrictive sites

Ala145the site
Leu148the site
Leu141the site
→ glycine−1.71
→ tert-butylalanine−0.04
→ tert-butylalanine+0.02
Restrictive positions

Restrictive positions · and where it falls short

Nine of eleven at Glu153; one of seven at Tyr152

Glu153the site
→ 2-aminobutyric acid−0.24
→ norvaline−0.12

Glu153 tolerates 11 substitutions and the model produced 9 of them, missing only 2-aminooctanoic acid and tert-butylalanine — both of which it does build elsewhere.

Tyr152the site
→ tryptophan−1.57

Tyr152 tolerates 7, and the model produced exactly one — tryptophan, which happens to be the best of the seven.

The gap is a specific one. Five of the six it misses at Tyr152 are bulky aromatic non-canonicals — benzothienyl-alanine, 2- and 3-thienyl-alanine, O-methyl-tyrosine and 2-naphthyl-alanine; the sixth is 2-aminoheptanoic acid, which the assay only just tolerates (+0.47). The unconditional prior rarely proposes the aromatic class anywhere. It is the clearest target for the conditional model, and the reason we are building one.

Restrictive positions

Validity

Are the molecules it draws actually right?

51 / 53correct bond orders
45 / 53every bond within 10%

Across every residue shown in this deck we checked the generated molecule rather than the label: element-by-element bond lengths against ideal values, ring planarity, and non-bonded clashes — then the bond orders, which the graph-isomorphism labeller ignores.

Aromatic rings come out planar and correctly conjugated; aliphatic chains come out saturated; amide and carboxylate groups carry their double bonds in the right place.

CheckResult
Side-chain bond orders51 / 53
Every bond within 10% of ideal45 / 53
No non-bonded contact < 2.45 Å51 / 53
Aromatic rings planar and conjugated13 / 13
Saturated rings puckered0 / 2

The one real failure is worth naming. Cyclopentyl-alanine came out with two double bonds in its five-membered ring — a flat cyclopentadiene rather than a puckered cyclopentane. The labeller matched it anyway, because it compares skeletons and not bond orders. Both draws of it, at two separate positions, failed the same way. Saturated rings are a known weak point and we report it as one.

Validity

Result

UNAAGI has the highest correlation for non-canonical substitutions

Protein fitness paired barplot

Higher bars mean closer agreement between each model’s ranking and the measured effects on binding.

UNAAGI leads for non-canonical substitutions in both peptides. Across all substitutions, it leads on PUMA and performs similarly to AF3 iPTM on CP2.

Result

Sanity check

Benchmarking canonical substitutions across 25 ProteinGym assays

Spearman correlation on 25 ProteinGym DMS assays, UNAAGI against ten baselines
Sanity check · ProteinGym

In progress

Next steps: scoring specific amino acids and learning from assay data

1 · Conditional diffusion

Scoring a chosen amino acid

local environment target residue UNAAGI model score

We are developing a model that estimates how well a chosen amino acid fits its local atomic environment. This will support scoring both canonical and non-canonical substitutions.

2 · Preference optimisation

Learning from experimental results

generate measure update model

We plan to use assay results to favour substitutions with better measured binding. New data will guide model updates and the next round of candidate generation.

In progress