Design ligand- and multimer-aware sequences with LigandMPNN
LigandMPNN (Dauparas et al., 2023) extends ProteinMPNN's structure-conditioned design approach to protein-ligand and multimer systems, released under the MIT license.
Status: beta. Install and run both work today. The registry keeps this at beta (rather than available) until execution is verified across more environments.
Install
$ moldesk install ligandmpnnSee CLI usage for reinstalling, uninstalling, and other commands.
Example
Download a sample protein-ligand structure — trypsin bound to benzamidine (3PTB) from the RCSB PDB — and design new sequences around it:
$ curl -O https://files.rcsb.org/download/3PTB.pdb$ moldesk run ligandmpnn 3PTB.pdb --output ./resultsMoleculeDesk prints the run's status and the path to each output file as it completes. --output ./results also copies the designed sequences into ./results as .fa (FASTA) files you can open directly.
What LigandMPNN is used for
LigandMPNN is designed for cases where ProteinMPNN's protein-only view isn't enough:
- Protein-ligand interface design — designing sequences that account for a bound small molecule or cofactor.
- Multimer design — designing sequences across multi-chain assemblies rather than a single chain.
Requirements
| Model version | 0.1.2 |
| License | MIT |
| Python | 3.11, via uv |
| Key dependencies | torch 2.2.1, numpy 1.23.5, prody 2.4.1, biopython 1.79, scipy 1.12.0 |
| Platforms | darwin-arm64, linux-x64 |
| Input format | .pdb |
| Output | sequences (seqs/*.fa) |
MoleculeDesk pins LigandMPNN to a specific upstream commit and verifies its model checkpoint (~10.5 MB) by SHA-256 before use.
LigandMPNN benchmarks: speed, VRAM and cost per design batch
On an NVIDIA GeForce RTX 3090, LigandMPNN takes about 10 s per design batch once warm (15 s on the first run), roughly 360 design batchs per GPU-hour, or about <$0.01 per design batch.
| Metric | NVIDIA GeForce RTX 3090 |
|---|---|
| Time per run, first (cold) | 15 s |
| Time per run, warm | 10 s |
| Design batchs per GPU-hour | 360 |
| Cost per run, first (cold) | <$0.01 |
| Cost per run, warm | <$0.01 |
| Peak GPU memory | 0.4 GiB |
| Peak GPU utilization | 10 % |
| Peak GPU power | 108 W |
| Install time | 59 s |
| Disk per install | 7 GiB |
Workload: One 20-residue chain (examples/diffdock/protein.pdb), default sampling, 1 sequence. Median of 3 runs.
Machine: RunPod GPU pod, Linux x64, 125 GiB RAM; AMD EPYC 7H12 (32 vCPU allocated); driver 580.126.20.
Cost: at $0.50/GPU-hour (RunPod Secure Cloud, EU-CZ-1, compute only; storage adds about $0.02/hr); excludes storage and idle time.
Warm time is the mean of runs 2 and 3 (9.7 s and 9.9 s). Install time is with PyTorch already in the package cache; expect several minutes on a first install.
FAQ
Can I run LigandMPNN through MoleculeDesk today?
Yes — both moldesk install ligandmpnn and moldesk run ligandmpnn <input>.pdb work today (see the example above). See Model compatibility & registry for what beta means.
How is this different from ProteinMPNN? LigandMPNN adds awareness of bound ligands and multi-chain (multimer) context; use ProteinMPNN for simpler single-chain, protein-only design.
What license is LigandMPNN under? MIT.