FPBench · Phase Stability and Elemental Ordering

FPBench: Phase Stability and Elemental Ordering

Convex-hull/relative-phase-stability, elemental-ordering energy-ranking, and structural-relaxation (RMSD) metrics for foundation potentials (FPs), evaluated on the FPBench convex-hull and elemental-ordering benchmark subsets.

View Leaderboard Use FPBench Dataset (Figshare) Paper (arXiv) GitHub
lower is better higher is better green = best in column cell shading: green = better, gold = middle, red = worse identifiers are never colored

Convex-hull metrics are evaluated across 22 ternary chalcogenide tie-line systems for both full FP relaxation and static FP evaluations on the DFT-relaxed structures. Ground-state agreement measures whether the FP identifies the same lowest-energy phase at each composition as DFT, while within-phase hull-minimum agreement and hull-minimum agreement compare the predicted hull minima. See Metrics for full definitions.

The per-column best value (green) updates when you click a column header. The Version (training) column gives the exact checkpoint evaluated and the dataset it was trained on.

Full FP relaxation

FP Average energy error
(meV/atom) ↓
Ground-state agreement
(%) ↑
Within-phase
hull-minimum
agreement (%) ↑
Hull-minimum
agreement
(%) ↑
Version (training)
MACE5368.362.040.9>=v0.3.10 (MACE-MPA-0, medium) (MPtrj + sAlex)
CHGNet24253.440.518.2v0.3.0 (MPtrj)
M3GNet4955.347.122.7MP-2021.2.8-PES (MP-2021.2.8)
UMA8362.760.336.4s-1p1 (OC20+ODAC23+OMat24+OMC25+OMol25)
M3GNet-MatPES11750.355.422.7v2025.1 (MatPES-PBE)
TensorNet-MatPES8952.858.713.6v2025.1 (MatPES-PBE)
MACE-MatPES8949.759.518.2>=v0.3.10 (MatPES-PBE (fine-tuned))

Static FP evaluations on the DFT-relaxed structures

FP Average energy error
(meV/atom) ↓
Ground-state agreement
(%) ↑
Within-phase
hull-minimum
agreement (%) ↑
Hull-minimum
agreement
(%) ↑
MACE1570.273.640.9
CHGNet24655.948.822.7
M3GNet3355.959.522.7
UMA4374.578.540.9
M3GNet-MatPES10949.157.022.7
TensorNet-MatPES6649.163.613.6
MACE-MatPES5553.471.918.2

Elemental-ordering metrics are evaluated across 305 ordering groups of 20 candidates each, for both full FP relaxation and static FP evaluations on the DFT-relaxed structures. Top-1 accuracy measures recovery of the DFT-preferred ordering, Recall@k whether it is recovered among the top-ranked candidates, and Spearman ρ overall ranking agreement. See Metrics for full definitions.

The per-column best value (green) updates when you click a column header. The Version (training) column gives the exact checkpoint evaluated and the dataset it was trained on.

Full FP relaxation

FP Average energy error
(meV/atom) ↓
Top-1 accuracy (%) ↑ Recall@3 (%) ↑ Recall@10 (%) ↑ Spearman ρ ↑ Rate of ranking
errors
(%) ↓
Mean/max ΔEDFT
(meV/atom) ↓
Version (training)
MACE4323.632.860.90.2838.9>=v0.3.10 (MACE-MPA-0, medium) (MPtrj + sAlex)
CHGNet2728.220.152.80.0946.8v0.3.0 (MPtrj)
M3GNet567.218.651.30.0548.1MP-2021.2.8-PES (MP-2021.2.8)
UMA6524.334.661.90.3137.6s-1p1 (OC20+ODAC23+OMat24+OMC25+OMol25)
M3GNet-MatPES9510.221.054.80.1245.7v2025.1 (MatPES-PBE)
TensorNet-MatPES6912.825.456.10.1743.6v2025.1 (MatPES-PBE)
MACE-MatPES6921.332.760.50.2839.1>=v0.3.10 (MatPES-PBE (fine-tuned))

Static FP evaluations on the DFT-relaxed structures

FP Average energy error
(meV/atom) ↓
Top-1 accuracy (%) ↑ Recall@3 (%) ↑ Recall@10 (%) ↑ Spearman ρ ↑ Rate of ranking
errors
(%) ↓
Mean/max ΔEDFT
(meV/atom) ↓
MACE1446.654.572.70.5626.6
CHGNet27020.329.759.00.2440.6
M3GNet4122.328.758.70.2341.0
UMA3555.765.479.90.6919.7
M3GNet-MatPES9419.729.860.00.2640.2
TensorNet-MatPES5227.941.465.30.3834.8
MACE-MatPES4341.048.070.00.5029.6

RMSD applies only to full FP relaxation -- static evaluation does not produce an FP-relaxed structure to compare against DFT. Computed via pymatgen StructureMatcher, converted to Å via the geometric-mean cell volume. See Metrics for full definitions.

The per-column best value (green) updates when you click a column header. The Version (training) column gives the exact checkpoint evaluated and the dataset it was trained on.

FP Map success (%) ↑ Mean/max RMSD (Å) ↓ RMSD < 0.05 Å (%) ↑ RMSD < 0.10 Å (%) ↑ RMSD < 0.20 Å (%) ↑ Version (training)
MACE89.924.044.565.5>=v0.3.10 (MACE-MPA-0, medium) (MPtrj + sAlex)
CHGNet87.311.926.549.5v0.3.0 (MPtrj)
M3GNet88.810.221.139.4MP-2021.2.8-PES (MP-2021.2.8)
UMA89.924.445.466.1s-1p1 (OC20+ODAC23+OMat24+OMC25+OMol25)
M3GNet-MatPES89.812.326.351.3v2025.1 (MatPES-PBE)
TensorNet-MatPES89.915.636.763.1v2025.1 (MatPES-PBE)
MACE-MatPES89.623.644.564.5>=v0.3.10 (MatPES-PBE (fine-tuned))

Metrics

Evaluated for both full FP relaxation and static FP evaluation on the DFT-relaxed structure.

Convex hull and relative phase stability

NameMetrics
Average energy errorEnergy error relative to DFT, reported as the MAE over all structures in the convex-hull benchmark.
Ground-state agreement: lowest-energy phase at each composition.For each composition, the FP-predicted ground-state phase is compared with DFT. The reported metric is the fraction of compositions, across all tie-line systems, for which the two agree.
Within-phase hull-minimum agreement: lowest-energy composition within each phase.For each phase, the FP-predicted minimum-energy composition is compared with DFT. The reported metric is the fraction of phases, across all tie-line systems, for which the two agree.
Hull-minimum agreement: global convex-hull minimum across the tie-line.For each tie-line system, the FP-predicted convex-hull minimum is compared with DFT. A system with no structure below the tie line has no stable intermediate. The reported metric is the fraction of systems for which the two agree: they identify the same stable intermediate, or both find none and their tie lines are anchored on the same endpoint structures.

Energy rankings of elemental orderings

NameMetrics
Average energy errorEnergy error relative to DFT, reported as the MAE over all structures in the ordering benchmark.
Top-1 accuracyFraction of ordering groups for which the FP identifies the same lowest-energy configuration (LEC) as DFT.
Recall@kFraction of the k lowest-energy DFT orderings retained among the top k FP-ranked orderings within each ordering group. The reported metric is the mean over all ordering groups.
Spearman's ρRank correlation between FP and DFT orderings within each ordering group. The reported metric is the mean over all ordering groups.
Rate of ranking errorsFraction of configuration pairs whose FP and DFT rankings disagree, within each ordering group. The reported metric is the mean over all ordering groups.
ΔEDFT of misranked pairsDFT energy difference between configurations that the FP misranks. For each ordering group, the mean and maximum ΔEDFT over its misranked pairs are computed and then averaged across all ordering groups.

Structural-relaxation effects

Applies to full FP relaxation only (static evaluation does not produce an FP-relaxed structure to compare).

NameMetrics
Structure relaxation errorFor each structure, the RMSD between the FP-relaxed and DFT-relaxed structures is computed. The reported metric is the mean/maximum over all structures in the dataset.
Mean/max RMSDGeometric deviation after alignment (Å).
Map successFraction of cases where FP and DFT structures remain sufficiently similar for meaningful RMSD evaluation; failures correspond to large structural deviation.
RMSD thresholdsIncreasingly strict measures of fidelity.

These metrics are complementary and should be interpreted together rather than combined into a single overall ranking. Click any column heading to sort by that metric.

Benchmark Dataset

Convex hull
Tie-line systems: 22
22 ternary chalcogenide tie-line systems relevant to phase-change materials (PCMs), composed of Ge, Sn, Sb, Bi, Si, Ga, In, and Ti combined with Se or Te.
Elemental ordering
Ordering groups: 305
Candidates per group: 20
One group per tie-line composition selected for ordering analysis; every group has exactly 20 distinct atomic-ordering candidates.

Data availability. The FPBench convex-hull and elemental-ordering benchmark subsets are available on Figshare. The underlying ternary chalcogenide DFT data are available from Adams et al. on Dryad.

All DFT energies are raw, uncorrected VASP-PBE energies -- no Materials Project MP2020 compatibility correction is applied anywhere. Model versions and official sources for the seven FPs evaluated here are documented once on the FPBench home page. See the Phase_stability_ordering README for the full standardized-file schema.

Use FPBench

Provided FPBench reference + new FP calculations
                         or
User DFT reference + user FP results
                          ↓
           standardized reference/results
                          ↓
                validation and analysis
                          ↓
          hull, ordering, and RMSD tables

Evaluate a new FP on the provided benchmark. Use generation/convexhull_ordering_run_generator.ipynb to generate per-FP jobs and submission scripts against the shipped DFT reference, run them on your cluster, merge the results, then load the merged file directly in the analysis notebook -- your FP appears in every table above.

Apply the analysis functions to another dataset. Call build_phase_stability_ordering_results(...) and the table builders in scripts/convexhull_analysis_utils.py directly with your own DFT reference and FP results.

git clone https://github.com/mogroupumd/FPBench.git
cd FPBench/Phase_stability_ordering
pip install -r requirements.txt
pip install jupyterlab
jupyter lab analysis/convexhull_ordering_analysis_all_models.ipynb

See the Phase_stability_ordering README for the full quick start, the generator and analysis notebooks for complete operational detail, and examples/ for a small runnable slice of real data.

Contribute

Interested in evaluating a new foundation potential, or having it considered for inclusion in FPBench? See our Adding a Potential guide to integrate and evaluate a new model with FPBench. For inclusion in the public leaderboard, please contact Prof. Yifei Mo at yfmo@umd.edu with the model name, version/checkpoint, and a link to the official implementation or model weights.