Design protein sequences with ProteinMPNN

ProteinMPNN (Dauparas et al., Science 2022) is a structure-conditioned protein sequence design model, released under the MIT license. Given a protein backbone structure, it designs amino-acid sequences predicted to fold into that structure.

Status: beta. Install and run both work today. The registry keeps this at beta (rather than available) until execution is verified across more environments.

Install

$ moldesk install proteinmpnn

See CLI usage for reinstalling, uninstalling, and other commands.

Example

Download a sample structure — ubiquitin (1UBQ) from the RCSB PDB — and design new sequences for it:

$ curl -O https://files.rcsb.org/download/1UBQ.pdb
$ moldesk run proteinmpnn 1UBQ.pdb --output ./results

MoleculeDesk prints the run's status and the path to each output file as it completes. --output ./results also copies the designed sequences into ./results as .fa (FASTA) files you can open directly.

What ProteinMPNN is used for

ProteinMPNN is widely used upstream for:

  • Fixed-backbone sequence design — generating new sequences for a known structure.
  • Enzyme redesign — designing sequence variants around a target active site or fold.
  • Binder design — as a component in pipelines that design proteins to bind a specific target.

Requirements

Model versionv_48_020 (default upstream checkpoint)
LicenseMIT
Python3.11, via uv
Dependenciestorch 2.2.1, numpy 1.26.4
Platformsdarwin-arm64, linux-x64
Input format.pdb
Outputsequences (seqs/*.fa)

MoleculeDesk pins ProteinMPNN to a specific upstream commit and verifies the model checkpoint (~6.7 MB) by SHA-256 before use, so every install is reproducible.

ProteinMPNN benchmarks: speed, VRAM and cost per design batch

On an NVIDIA GeForce RTX 3090, ProteinMPNN takes about 8 s per design batch once warm (16 s on the first run), roughly 450 design batchs per GPU-hour, or about <$0.01 per design batch.

ProteinMPNN on NVIDIA GeForce RTX 3090 (24 GiB), measured with MoleculeDesk 57776a4
MetricNVIDIA GeForce RTX 3090
Time per run, first (cold)16 s
Time per run, warm8 s
Design batchs per GPU-hour450
Cost per run, first (cold)<$0.01
Cost per run, warm<$0.01
Peak GPU memory0.4 GiB
Peak GPU utilization9 %
Peak GPU power108 W
Install time6.1 min
Disk per install6 GiB

Workload: One 20-residue chain (examples/diffdock/protein.pdb), default sampling, 1 sequence. Median of 3 runs.

Machine: RunPod GPU pod, Linux x64, 125 GiB RAM; AMD EPYC 7H12 (32 vCPU allocated); driver 580.126.20.

Cost: at $0.50/GPU-hour (RunPod Secure Cloud, EU-CZ-1, compute only; storage adds about $0.02/hr); excludes storage and idle time.

Warm time is the mean of runs 2 and 3 (7.9 s and 8.1 s). Install time includes downloading PyTorch with an empty package cache. Peak GPU power is near the idle-to-light-load range because the job is tiny.

FAQ

Can I run ProteinMPNN through MoleculeDesk today? Yes — both moldesk install proteinmpnn and moldesk run proteinmpnn <input>.pdb work today (see the example above). See Model compatibility & registry for what beta means.

Does this require a GPU? No specific GPU requirement is declared for ProteinMPNN's manifest, though torch will use one if available. Run moldesk doctor to check your machine.

What license is ProteinMPNN under? MIT.