FPBench

Application-Oriented Error Decomposition for Foundation Potentials

FPBench introduces application-oriented error decomposition through metrics that resolve performance according to the physically consequential quantities and configurations governing computational tasks, including the fractions of highly accurate and large-force-error atoms, far-from-equilibrium atoms, relative phase-stability and convex-hull agreement, and along-path errors in ion migration. These metrics from error decompositions identify where FP errors arise within specific computational tasks, providing targeted guidance for model development.

FPBench evaluates foundation potentials (FPs), also known as universal machine-learning interatomic potentials (MLIPs), across three computational materials workflows: force prediction, phase stability and elemental ordering, and ion migration by NEB calculations.

Paper (arXiv) GitHub

Components

Available

Force Prediction

Force-error metrics for FPs on MatPES-PBE, MatPES-r2SCAN, and OMat24 rattled-1000, covering highly accurate force predictions, joint magnitude-angle accuracy, large-force-error atoms, and far-from-equilibrium force errors.

View Force Prediction →
Available

Phase Stability and Elemental Ordering

Convex-hull/relative-phase-stability, elemental-ordering energy-ranking, and structural-relaxation (RMSD) metrics for FPs, evaluated on the FPBench convex-hull and elemental-ordering benchmark subsets.

View Phase Stability & Ordering →
Available

Ion Migration by NEB

Migration-barrier and migration-pathway metrics for FPs on 154 real Li- and Na-ion migration pathways using the nudged elastic band (NEB) method.

View Ion Migration (NEB) →

Benchmark data: Phase Stability & Ordering dataset (Figshare) · Ion Migration dataset (Figshare)

Foundation Potentials Evaluated

Exact checkpoints and versions used in this study. Model size and training-dataset size are the values used in this study's own code/records; official source links point to each architecture's primary release page.

FP Model & version Model size Training dataset Approx. training-dataset size Official source
MACE >=v0.3.10 (MACE-MPA-0, medium) 9.06M MPtrj + sAlex ~3.5M MACE foundation models
CHGNet v0.3.0 412.5K MPtrj ~1.58M CHGNet (GitHub)
M3GNet MP-2021.2.8-PES 288.2K MP-2021.2.8 ~176.6K MatGL
UMA s-1p1 146.5M OC20 + ODAC23 + OMat24 + OMC25 + OMol25 ~500M FAIR Chemistry
M3GNet-MatPES v2025.1 664.2K MatPES-PBE ~435K MatGL · MatPES
TensorNet-MatPES v2025.1 837.9K MatPES-PBE ~435K MatGL · MatPES
MACE-MatPES >=v0.3.10 9.06M Fine-tuned on MatPES-PBE* ~435K MACE (GitHub)
M3GNet-MatPES-r2SCAN v2025.1 664.2K MatPES-r2SCAN ~388K MatGL · MatPES
TensorNet-MatPES-r2SCAN v2025.1 837.9K MatPES-r2SCAN ~388K MatGL · MatPES
MACE-MatPES-r2SCAN >=v0.3.10 9.06M Fine-tuned on MatPES-r2SCAN* ~388K MACE (GitHub)
Orb orb-v3-conservative-inf-omat-20250404 25.5M OMat24, AIMD subset only ~55M‡ Rhodes et al. 2025
SevenNet 7net-mf-ompa (modal mpa) 25.7M MPtrj + sAlex + OMat24, multi-fidelity; mpa selects the MPtrj + sAlex task not reported for this checkpoint SevenNet pretrained models · Kim et al. 2025
MatterSim MatterSim-v1.0.0-5M 4.55M Nonpublic MatterSim active-learning dataset, GGA-PBE(+U)† 6M Model card · Yang et al. 2024
Nequix nequix-oam-1 707.6K OMat24 + sAlex + MPtrj, DFT (PBE+U)§ not reported for this checkpoint nequix repository · Koker et al. 2025
GPTFF gptff_v2 502.5K Atomly‖ ~37.6M configurations Xie et al. 2024
ALIGNN alignnff_wt10 4.03M JARVIS-DFT¶ ~307.1K Choudhary et al. 2023
NEP89 nep89_20250409 976.3K OMat24, MPtrj, SPICE, ANI-1xnr, SSE-ABACUS, SSE-VASP, Protein, UNEP-v1, CH, CHONPS, Water; mixed QM levels†† 537,641 configurations GPUMD potentials · NEP89 paper
DPA4 DPA4-Plus-OMat24-v20260805 8.85M OMat24, DFT / DFT+U# ~100.6M frames DPA4-OMat24 model card
GRACE GRACE-3L-OMAT-large-ft-AM 42.1M‡‡ OMat24 pretraining → sAlex + MPtrj fine-tuning not reported for this checkpoint GRACE foundation models
eqV2 eqV2_31M_omat_mp_salex.pt 31.2M OMat24 pretraining → MPtrj + sAlex fine-tuning, DFT / DFT+U# not reported for this checkpoint Meta OMat24 models
eSEN esen_30m_oam.pt 30.2M OMat24 pretraining → MPtrj + sAlex fine-tuning, DFT / DFT+U# not reported for this checkpoint Meta OMat24 models

* Pre-trained on MACE-OMAT-0, then fine-tuned on the matched MatPES functional.

† The MatterSim model card reports "Training Data Size: 6M", "Model Parameters: 4.5M", and training on "a specific variant of Density Functional Theory (PBE)". The GGA-PBE(+U) labelling, with Hubbard U applied to selected materials following Materials Project settings, is described in Yang et al. rather than on the model card. The dataset itself is not released, so the OOD label reflects the absence of documented training exposure rather than a verified composition.

‡ Rhodes et al. state that "all orb-v3-*-omat models are only trained on the AIMD subset of OMat24", and that the OMat24 dataset "contains ~55 million AIMD-sampled structures".

§ The FPBench checkpoint is nequix-oam-1. The official nequix repository documents this checkpoint as trained on OMat24, sAlex and MPtrj at the DFT (PBE+U) level. The Nequix paper describes an MPtrj-trained model and does not report a training-set size or parameter count for this released OAM checkpoint; the 707,569-parameter model size shown here was measured from the loaded checkpoint by FPBench.

‖ GPTFF’s Atomly labels were computed with VASP at the GGA-PBE level with a 520 eV plane-wave cutoff and version 5.4 PAW pseudopotentials. The functional matches the FPBench reference; the cutoff and pseudopotential version do not necessarily, so small systematic offsets relative to MatPES-PBE are expected independently of model quality.

¶ ALIGNN is trained on JARVIS-DFT at the OptB88vdW level, not PBE. Its training set is 307,113 points, split 90:5:5. It is the only FP here whose training reference functional differs from the evaluation reference, so its scores measure agreement with the FPBench MatPES-PBE reference rather than a functional-matched fitting error.

†† NEP89 is trained on eleven datasets computed at different quantum-mechanical levels rather than a single reference functional. Its FPBench MatPES-PBE scores therefore measure practical agreement with the PBE benchmark reference rather than agreement with a single matched reference functional. The eleven datasets and the 537,641-configuration total are reported in the NEP89 paper; the 976,331 trainable parameters were counted from the released model by FPBench.

‡‡ FPBench counts 42.1M parameters (42,112,223) for the GRACE potential used in inference. The released SavedModel additionally contains a 23.4M-parameter GMM uncertainty head that is disabled in FPBench runs and is therefore not included in the reported model size. The GRACE documentation does not publish a parameter count; 42.1M was measured from the released model by FPBench.

# The immediate model sources describe these labels as DFT and DFT+U total-energy labels. OMat24 itself uses PBE/PBE+U reference calculations; see the OMat24 dataset reference.

The potentials listed after the r2SCAN entries were evaluated after submission of the manuscript, which reports the ten FPs above them. Their model sizes are parameter counts measured from the loaded checkpoints.

This is the current set of FPs evaluated. As additional FPs are tested, they will be added to this table; existing benchmark results remain tied to their original model versions and are not silently replaced.

Evaluation Matrix

Which FPs were evaluated on which FPBench component/dataset. A checkmark means the FP was evaluated there; the parenthetical states whether that dataset falls inside (training) or outside (OOD, out-of-distribution) the FP's own training data, whether the FP is a fine-tune of an OOD base model onto that dataset (fine-tuned), or whether that dataset was used to pretrain the model before it was fine-tuned on other data (pretraining).

FPMatPES-PBE (Force Prediction)OMat24 rattled-1000 (Force Prediction, SI)Phase Stability/Ordering & Ion Migration (NEB)
MACE✓ (OOD)✓ (OOD)✓ (OOD)
CHGNet✓ (OOD)✓ (OOD)✓ (OOD)
M3GNet✓ (OOD)✓ (OOD)✓ (OOD)
UMA✓ (OOD)✓ (training)✓ (OOD)
M3GNet-MatPES✓ (training)✓ (OOD)✓ (OOD)
TensorNet-MatPES✓ (training)✓ (OOD)✓ (OOD)
MACE-MatPES✓ (training)✓ (pretraining)*✓ (OOD)
Orb✓ (OOD)✓ (training)not evaluated
SevenNet✓ (OOD)✓ (training)not evaluated
MatterSim✓ (OOD)†✓ (OOD)not evaluated
Nequix✓ (OOD)§✓ (training)not evaluated
GPTFF✓ (OOD)✓ (OOD)not evaluated
ALIGNN✓ (OOD)¶not evaluatednot evaluated
NEP89✓ (OOD)††✓ (OOD)not evaluated
DPA4✓ (OOD)✓ (training)not evaluated
GRACE✓ (OOD)✓ (pretraining)not evaluated
eqV2✓ (OOD)✓ (pretraining)not evaluated
eSEN✓ (OOD)✓ (pretraining)not evaluated

* Pre-trained on MACE-OMAT-0, fine-tuned on MatPES-PBE.

The three r2SCAN-trained FPs (M3GNet-MatPES-r2SCAN, TensorNet-MatPES-r2SCAN, MACE-MatPES-r2SCAN) are evaluated on the MatPES-r2SCAN dataset (training), in addition to the models listed above.

The "Phase Stability/Ordering & Ion Migration (NEB)" column reflects evaluation in the manuscript. Public code and results for both are now available -- see Phase Stability & Ordering and Ion Migration (NEB).

Contribute

Interested in evaluating a new foundation potential, or having it considered for inclusion in FPBench? See our Adding a Potential guide to integrate and evaluate a new model with FPBench. For inclusion in the public leaderboard, please contact Prof. Yifei Mo at yfmo@umd.edu with the model name, version/checkpoint, and a link to the official implementation or model weights.

Citation

If you use FPBench, please cite:

Kiyan Amirian, Ramanuja Srinivasan Saravanan, Felix Adams, Charles E Schwarz and Yifei Mo, “FPBench: Application-Oriented Error Decomposition for Foundation Potentials”, arXiv:2609.05714 (2026). arxiv.org/abs/2609.05714

@misc{amirian2026fpbench, author = {Kiyan Amirian and Ramanuja Srinivasan Saravanan and Felix Adams and Charles E Schwarz and Yifei Mo}, title = {FPBench: Application-Oriented Error Decomposition for Foundation Potentials}, year = {2026}, eprint = {2609.05714}, archivePrefix = {arXiv}, primaryClass = {cond-mat.mtrl-sci}, url = {https://arxiv.org/abs/2609.05714} }