FPBench introduces application-oriented error decomposition through metrics that resolve performance according to the physically consequential quantities and configurations governing computational tasks, including the fractions of highly accurate and large-force-error atoms, far-from-equilibrium atoms, relative phase-stability and convex-hull agreement, and along-path errors in ion migration. These metrics from error decompositions identify where FP errors arise within specific computational tasks, providing targeted guidance for model development.
FPBench evaluates foundation potentials (FPs), also known as universal machine-learning interatomic potentials (MLIPs), across three computational materials workflows: force prediction, phase stability and elemental ordering, and ion migration by NEB calculations.
Force-error metrics for FPs on MatPES-PBE, MatPES-r2SCAN, and OMat24 rattled-1000, covering highly accurate force predictions, joint magnitude-angle accuracy, large-force-error atoms, and far-from-equilibrium force errors.
View Force Prediction →Convex-hull/relative-phase-stability, elemental-ordering energy-ranking, and structural-relaxation (RMSD) metrics for FPs, evaluated on the FPBench convex-hull and elemental-ordering benchmark subsets.
View Phase Stability & Ordering →Migration-barrier and migration-pathway metrics for FPs on 154 real Li- and Na-ion migration pathways using the nudged elastic band (NEB) method.
View Ion Migration (NEB) →Benchmark data: Phase Stability & Ordering dataset (Figshare) · Ion Migration dataset (Figshare)
Exact checkpoints and versions used in this study. Model size and training-dataset size are the values used in this study's own code/records; official source links point to each architecture's primary release page.
| FP | Model & version | Model size | Training dataset | Approx. training-dataset size | Official source |
|---|---|---|---|---|---|
| MACE | >=v0.3.10 (MACE-MPA-0, medium) | 9.06M | MPtrj + sAlex | ~3.5M | MACE foundation models |
| CHGNet | v0.3.0 | 412.5K | MPtrj | ~1.58M | CHGNet (GitHub) |
| M3GNet | MP-2021.2.8-PES | 288.2K | MP-2021.2.8 | ~176.6K | MatGL |
| UMA | s-1p1 | 146.5M | OC20 + ODAC23 + OMat24 + OMC25 + OMol25 | ~500M | FAIR Chemistry |
| M3GNet-MatPES | v2025.1 | 664.2K | MatPES-PBE | ~435K | MatGL · MatPES |
| TensorNet-MatPES | v2025.1 | 837.9K | MatPES-PBE | ~435K | MatGL · MatPES |
| MACE-MatPES | >=v0.3.10 | 9.06M | Fine-tuned on MatPES-PBE* | ~435K | MACE (GitHub) |
| M3GNet-MatPES-r2SCAN | v2025.1 | 664.2K | MatPES-r2SCAN | ~388K | MatGL · MatPES |
| TensorNet-MatPES-r2SCAN | v2025.1 | 837.9K | MatPES-r2SCAN | ~388K | MatGL · MatPES |
| MACE-MatPES-r2SCAN | >=v0.3.10 | 9.06M | Fine-tuned on MatPES-r2SCAN* | ~388K | MACE (GitHub) |
| Orb | orb-v3-conservative-inf-omat-20250404 | 25.5M | OMat24, AIMD subset only | ~55M‡ | Rhodes et al. 2025 |
| SevenNet | 7net-mf-ompa (modal mpa) |
25.7M | MPtrj + sAlex + OMat24, multi-fidelity; mpa selects the MPtrj + sAlex task |
not reported for this checkpoint | SevenNet pretrained models · Kim et al. 2025 |
| MatterSim | MatterSim-v1.0.0-5M | 4.55M | Nonpublic MatterSim active-learning dataset, GGA-PBE(+U)† | 6M | Model card · Yang et al. 2024 |
| Nequix | nequix-oam-1 | 707.6K | OMat24 + sAlex + MPtrj, DFT (PBE+U)§ | not reported for this checkpoint | nequix repository · Koker et al. 2025 |
| GPTFF | gptff_v2 | 502.5K | Atomly‖ | ~37.6M configurations | Xie et al. 2024 |
| ALIGNN | alignnff_wt10 | 4.03M | JARVIS-DFT¶ | ~307.1K | Choudhary et al. 2023 |
| NEP89 | nep89_20250409 | 976.3K | OMat24, MPtrj, SPICE, ANI-1xnr, SSE-ABACUS, SSE-VASP, Protein, UNEP-v1, CH, CHONPS, Water; mixed QM levels†† | 537,641 configurations | GPUMD potentials · NEP89 paper |
| DPA4 | DPA4-Plus-OMat24-v20260805 | 8.85M | OMat24, DFT / DFT+U# | ~100.6M frames | DPA4-OMat24 model card |
| GRACE | GRACE-3L-OMAT-large-ft-AM | 42.1M‡‡ | OMat24 pretraining → sAlex + MPtrj fine-tuning | not reported for this checkpoint | GRACE foundation models |
| eqV2 | eqV2_31M_omat_mp_salex.pt | 31.2M | OMat24 pretraining → MPtrj + sAlex fine-tuning, DFT / DFT+U# | not reported for this checkpoint | Meta OMat24 models |
| eSEN | esen_30m_oam.pt | 30.2M | OMat24 pretraining → MPtrj + sAlex fine-tuning, DFT / DFT+U# | not reported for this checkpoint | Meta OMat24 models |
* Pre-trained on MACE-OMAT-0, then fine-tuned on the matched MatPES functional.
† The MatterSim model card reports "Training Data Size: 6M", "Model Parameters: 4.5M", and training on "a specific variant of Density Functional Theory (PBE)". The GGA-PBE(+U) labelling, with Hubbard U applied to selected materials following Materials Project settings, is described in Yang et al. rather than on the model card. The dataset itself is not released, so the OOD label reflects the absence of documented training exposure rather than a verified composition.
‡ Rhodes et al. state that "all orb-v3-*-omat models are only trained on the AIMD subset of OMat24", and that the OMat24 dataset "contains ~55 million AIMD-sampled structures".
§ The FPBench checkpoint is nequix-oam-1. The official nequix repository documents this checkpoint as trained on OMat24, sAlex and MPtrj at the DFT (PBE+U) level. The Nequix paper describes an MPtrj-trained model and does not report a training-set size or parameter count for this released OAM checkpoint; the 707,569-parameter model size shown here was measured from the loaded checkpoint by FPBench.
‖ GPTFF’s Atomly labels were computed with VASP at the GGA-PBE level with a 520 eV plane-wave cutoff and version 5.4 PAW pseudopotentials. The functional matches the FPBench reference; the cutoff and pseudopotential version do not necessarily, so small systematic offsets relative to MatPES-PBE are expected independently of model quality.
¶ ALIGNN is trained on JARVIS-DFT at the OptB88vdW level, not PBE. Its training set is 307,113 points, split 90:5:5. It is the only FP here whose training reference functional differs from the evaluation reference, so its scores measure agreement with the FPBench MatPES-PBE reference rather than a functional-matched fitting error.
†† NEP89 is trained on eleven datasets computed at different quantum-mechanical levels rather than a single reference functional. Its FPBench MatPES-PBE scores therefore measure practical agreement with the PBE benchmark reference rather than agreement with a single matched reference functional. The eleven datasets and the 537,641-configuration total are reported in the NEP89 paper; the 976,331 trainable parameters were counted from the released model by FPBench.
‡‡ FPBench counts 42.1M parameters (42,112,223) for the GRACE potential used in inference. The released SavedModel additionally contains a 23.4M-parameter GMM uncertainty head that is disabled in FPBench runs and is therefore not included in the reported model size. The GRACE documentation does not publish a parameter count; 42.1M was measured from the released model by FPBench.
# The immediate model sources describe these labels as DFT and DFT+U total-energy labels. OMat24 itself uses PBE/PBE+U reference calculations; see the OMat24 dataset reference.
The potentials listed after the r2SCAN entries were evaluated after submission of the manuscript, which reports the ten FPs above them. Their model sizes are parameter counts measured from the loaded checkpoints.
This is the current set of FPs evaluated. As additional FPs are tested, they will be added to this table; existing benchmark results remain tied to their original model versions and are not silently replaced.
Which FPs were evaluated on which FPBench component/dataset. A checkmark means the FP was evaluated there; the parenthetical states whether that dataset falls inside (training) or outside (OOD, out-of-distribution) the FP's own training data, whether the FP is a fine-tune of an OOD base model onto that dataset (fine-tuned), or whether that dataset was used to pretrain the model before it was fine-tuned on other data (pretraining).
| FP | MatPES-PBE (Force Prediction) | OMat24 rattled-1000 (Force Prediction, SI) | Phase Stability/Ordering & Ion Migration (NEB) |
|---|---|---|---|
| MACE | ✓ (OOD) | ✓ (OOD) | ✓ (OOD) |
| CHGNet | ✓ (OOD) | ✓ (OOD) | ✓ (OOD) |
| M3GNet | ✓ (OOD) | ✓ (OOD) | ✓ (OOD) |
| UMA | ✓ (OOD) | ✓ (training) | ✓ (OOD) |
| M3GNet-MatPES | ✓ (training) | ✓ (OOD) | ✓ (OOD) |
| TensorNet-MatPES | ✓ (training) | ✓ (OOD) | ✓ (OOD) |
| MACE-MatPES | ✓ (training) | ✓ (pretraining)* | ✓ (OOD) |
| Orb | ✓ (OOD) | ✓ (training) | not evaluated |
| SevenNet | ✓ (OOD) | ✓ (training) | not evaluated |
| MatterSim | ✓ (OOD)† | ✓ (OOD) | not evaluated |
| Nequix | ✓ (OOD)§ | ✓ (training) | not evaluated |
| GPTFF | ✓ (OOD) | ✓ (OOD) | not evaluated |
| ALIGNN | ✓ (OOD)¶ | not evaluated | not evaluated |
| NEP89 | ✓ (OOD)†† | ✓ (OOD) | not evaluated |
| DPA4 | ✓ (OOD) | ✓ (training) | not evaluated |
| GRACE | ✓ (OOD) | ✓ (pretraining) | not evaluated |
| eqV2 | ✓ (OOD) | ✓ (pretraining) | not evaluated |
| eSEN | ✓ (OOD) | ✓ (pretraining) | not evaluated |
* Pre-trained on MACE-OMAT-0, fine-tuned on MatPES-PBE.
The three r2SCAN-trained FPs (M3GNet-MatPES-r2SCAN, TensorNet-MatPES-r2SCAN, MACE-MatPES-r2SCAN) are evaluated on the MatPES-r2SCAN dataset (training), in addition to the models listed above.
The "Phase Stability/Ordering & Ion Migration (NEB)" column reflects evaluation in the manuscript. Public code and results for both are now available -- see Phase Stability & Ordering and Ion Migration (NEB).
Interested in evaluating a new foundation potential, or having it considered for inclusion in FPBench? See our Adding a Potential guide to integrate and evaluate a new model with FPBench. For inclusion in the public leaderboard, please contact Prof. Yifei Mo at yfmo@umd.edu with the model name, version/checkpoint, and a link to the official implementation or model weights.
If you use FPBench, please cite:
Kiyan Amirian, Ramanuja Srinivasan Saravanan, Felix Adams, Charles E Schwarz and Yifei Mo, “FPBench: Application-Oriented Error Decomposition for Foundation Potentials”, arXiv:2609.05714 (2026). arxiv.org/abs/2609.05714