Design ligand- and multimer-aware sequences with LigandMPNN

LigandMPNN (Dauparas et al., 2023) extends ProteinMPNN's structure-conditioned design approach to protein-ligand and multimer systems, released under the MIT license.

Status: beta. Install and run both work today. The registry keeps this at beta (rather than available) until execution is verified across more environments.

Install

$ moldesk install ligandmpnn

See CLI usage for reinstalling, uninstalling, and other commands.

Example

Download a sample protein-ligand structure — trypsin bound to benzamidine (3PTB) from the RCSB PDB — and design new sequences around it:

$ curl -O https://files.rcsb.org/download/3PTB.pdb
$ moldesk run ligandmpnn 3PTB.pdb --output ./results

MoleculeDesk prints the run's status and the path to each output file as it completes. --output ./results also copies the designed sequences into ./results as .fa (FASTA) files you can open directly.

What LigandMPNN is used for

LigandMPNN is designed for cases where ProteinMPNN's protein-only view isn't enough:

  • Protein-ligand interface design — designing sequences that account for a bound small molecule or cofactor.
  • Multimer design — designing sequences across multi-chain assemblies rather than a single chain.

Requirements

Model version0.1.2
LicenseMIT
Python3.11, via uv
Key dependenciestorch 2.2.1, numpy 1.23.5, prody 2.4.1, biopython 1.79, scipy 1.12.0
Platformsdarwin-arm64, linux-x64
Input format.pdb
Outputsequences (seqs/*.fa)

MoleculeDesk pins LigandMPNN to a specific upstream commit and verifies its model checkpoint (~10.5 MB) by SHA-256 before use.

LigandMPNN benchmarks: speed, VRAM and cost per design batch

On an NVIDIA GeForce RTX 3090, LigandMPNN takes about 10 s per design batch once warm (15 s on the first run), roughly 360 design batchs per GPU-hour, or about <$0.01 per design batch.

LigandMPNN on NVIDIA GeForce RTX 3090 (24 GiB), measured with MoleculeDesk 57776a4
MetricNVIDIA GeForce RTX 3090
Time per run, first (cold)15 s
Time per run, warm10 s
Design batchs per GPU-hour360
Cost per run, first (cold)<$0.01
Cost per run, warm<$0.01
Peak GPU memory0.4 GiB
Peak GPU utilization10 %
Peak GPU power108 W
Install time59 s
Disk per install7 GiB

Workload: One 20-residue chain (examples/diffdock/protein.pdb), default sampling, 1 sequence. Median of 3 runs.

Machine: RunPod GPU pod, Linux x64, 125 GiB RAM; AMD EPYC 7H12 (32 vCPU allocated); driver 580.126.20.

Cost: at $0.50/GPU-hour (RunPod Secure Cloud, EU-CZ-1, compute only; storage adds about $0.02/hr); excludes storage and idle time.

Warm time is the mean of runs 2 and 3 (9.7 s and 9.9 s). Install time is with PyTorch already in the package cache; expect several minutes on a first install.

FAQ

Can I run LigandMPNN through MoleculeDesk today? Yes — both moldesk install ligandmpnn and moldesk run ligandmpnn <input>.pdb work today (see the example above). See Model compatibility & registry for what beta means.

How is this different from ProteinMPNN? LigandMPNN adds awareness of bound ligands and multi-chain (multimer) context; use ProteinMPNN for simpler single-chain, protein-only design.

What license is LigandMPNN under? MIT.