Convex-hull/relative-phase-stability, elemental-ordering energy-ranking, and structural-relaxation (RMSD) metrics for foundation potentials (FPs), evaluated on the FPBench convex-hull and elemental-ordering benchmark subsets.
Convex-hull metrics are evaluated across 22 ternary chalcogenide tie-line systems for both full FP relaxation and static FP evaluations on the DFT-relaxed structures. Ground-state agreement measures whether the FP identifies the same lowest-energy phase at each composition as DFT, while within-phase hull-minimum agreement and hull-minimum agreement compare the predicted hull minima. See Metrics for full definitions.
The per-column best value (green) updates when you click a column header. The Version (training) column gives the exact checkpoint evaluated and the dataset it was trained on.
| FP | Average energy error (meV/atom) ↓ |
Ground-state agreement (%) ↑ |
Within-phase hull-minimum agreement (%) ↑ |
Hull-minimum agreement (%) ↑ |
Version (training) |
|---|---|---|---|---|---|
| MACE | 53 | 68.3 | 62.0 | 40.9 | >=v0.3.10 (MACE-MPA-0, medium) (MPtrj + sAlex) |
| CHGNet | 242 | 53.4 | 40.5 | 18.2 | v0.3.0 (MPtrj) |
| M3GNet | 49 | 55.3 | 47.1 | 22.7 | MP-2021.2.8-PES (MP-2021.2.8) |
| UMA | 83 | 62.7 | 60.3 | 36.4 | s-1p1 (OC20+ |
| M3GNet-MatPES | 117 | 50.3 | 55.4 | 22.7 | v2025.1 (MatPES-PBE) |
| TensorNet-MatPES | 89 | 52.8 | 58.7 | 13.6 | v2025.1 (MatPES-PBE) |
| MACE-MatPES | 89 | 49.7 | 59.5 | 18.2 | >=v0.3.10 (MatPES-PBE (fine-tuned)) |
| FP | Average energy error (meV/atom) ↓ |
Ground-state agreement (%) ↑ |
Within-phase hull-minimum agreement (%) ↑ |
Hull-minimum agreement (%) ↑ |
|---|---|---|---|---|
| MACE | 15 | 70.2 | 73.6 | 40.9 |
| CHGNet | 246 | 55.9 | 48.8 | 22.7 |
| M3GNet | 33 | 55.9 | 59.5 | 22.7 |
| UMA | 43 | 74.5 | 78.5 | 40.9 |
| M3GNet-MatPES | 109 | 49.1 | 57.0 | 22.7 |
| TensorNet-MatPES | 66 | 49.1 | 63.6 | 13.6 |
| MACE-MatPES | 55 | 53.4 | 71.9 | 18.2 |
Elemental-ordering metrics are evaluated across 305 ordering groups of 20 candidates each, for both full FP relaxation and static FP evaluations on the DFT-relaxed structures. Top-1 accuracy measures recovery of the DFT-preferred ordering, Recall@k whether it is recovered among the top-ranked candidates, and Spearman ρ overall ranking agreement. See Metrics for full definitions.
The per-column best value (green) updates when you click a column header. The Version (training) column gives the exact checkpoint evaluated and the dataset it was trained on.
| FP | Average energy error (meV/atom) ↓ |
Top-1 accuracy (%) ↑ | Recall@3 (%) ↑ | Recall@10 (%) ↑ | Spearman ρ ↑ | Rate of ranking errors (%) ↓ |
Mean/max ΔEDFT (meV/atom) ↓ |
Version (training) |
|---|---|---|---|---|---|---|---|---|
| MACE | 43 | 23.6 | 32.8 | 60.9 | 0.28 | 38.9 | >=v0.3.10 (MACE-MPA-0, medium) (MPtrj + sAlex) | |
| CHGNet | 272 | 8.2 | 20.1 | 52.8 | 0.09 | 46.8 | v0.3.0 (MPtrj) | |
| M3GNet | 56 | 7.2 | 18.6 | 51.3 | 0.05 | 48.1 | MP-2021.2.8-PES (MP-2021.2.8) | |
| UMA | 65 | 24.3 | 34.6 | 61.9 | 0.31 | 37.6 | s-1p1 (OC20+ | |
| M3GNet-MatPES | 95 | 10.2 | 21.0 | 54.8 | 0.12 | 45.7 | v2025.1 (MatPES-PBE) | |
| TensorNet-MatPES | 69 | 12.8 | 25.4 | 56.1 | 0.17 | 43.6 | v2025.1 (MatPES-PBE) | |
| MACE-MatPES | 69 | 21.3 | 32.7 | 60.5 | 0.28 | 39.1 | >=v0.3.10 (MatPES-PBE (fine-tuned)) |
| FP | Average energy error (meV/atom) ↓ |
Top-1 accuracy (%) ↑ | Recall@3 (%) ↑ | Recall@10 (%) ↑ | Spearman ρ ↑ | Rate of ranking errors (%) ↓ |
Mean/max ΔEDFT (meV/atom) ↓ |
|---|---|---|---|---|---|---|---|
| MACE | 14 | 46.6 | 54.5 | 72.7 | 0.56 | 26.6 | |
| CHGNet | 270 | 20.3 | 29.7 | 59.0 | 0.24 | 40.6 | |
| M3GNet | 41 | 22.3 | 28.7 | 58.7 | 0.23 | 41.0 | |
| UMA | 35 | 55.7 | 65.4 | 79.9 | 0.69 | 19.7 | |
| M3GNet-MatPES | 94 | 19.7 | 29.8 | 60.0 | 0.26 | 40.2 | |
| TensorNet-MatPES | 52 | 27.9 | 41.4 | 65.3 | 0.38 | 34.8 | |
| MACE-MatPES | 43 | 41.0 | 48.0 | 70.0 | 0.50 | 29.6 |
RMSD applies only to full FP relaxation --
static evaluation does not produce an FP-relaxed structure to compare against DFT. Computed
via pymatgen StructureMatcher, converted to Å via the geometric-mean cell
volume. See Metrics for full definitions.
The per-column best value (green) updates when you click a column header. The Version (training) column gives the exact checkpoint evaluated and the dataset it was trained on.
| FP | Map success (%) ↑ | Mean/max RMSD (Å) ↓ | RMSD < 0.05 Å (%) ↑ | RMSD < 0.10 Å (%) ↑ | RMSD < 0.20 Å (%) ↑ | Version (training) |
|---|---|---|---|---|---|---|
| MACE | 89.9 | 24.0 | 44.5 | 65.5 | >=v0.3.10 (MACE-MPA-0, medium) (MPtrj + sAlex) | |
| CHGNet | 87.3 | 11.9 | 26.5 | 49.5 | v0.3.0 (MPtrj) | |
| M3GNet | 88.8 | 10.2 | 21.1 | 39.4 | MP-2021.2.8-PES (MP-2021.2.8) | |
| UMA | 89.9 | 24.4 | 45.4 | 66.1 | s-1p1 (OC20+ | |
| M3GNet-MatPES | 89.8 | 12.3 | 26.3 | 51.3 | v2025.1 (MatPES-PBE) | |
| TensorNet-MatPES | 89.9 | 15.6 | 36.7 | 63.1 | v2025.1 (MatPES-PBE) | |
| MACE-MatPES | 89.6 | 23.6 | 44.5 | 64.5 | >=v0.3.10 (MatPES-PBE (fine-tuned)) |
Evaluated for both full FP relaxation and static FP evaluation on the DFT-relaxed structure.
| Name | Metrics |
|---|---|
| Average energy error | Energy error relative to DFT, reported as the MAE over all structures in the convex-hull benchmark. |
| Ground-state agreement: lowest-energy phase at each composition. | For each composition, the FP-predicted ground-state phase is compared with DFT. The reported metric is the fraction of compositions, across all tie-line systems, for which the two agree. |
| Within-phase hull-minimum agreement: lowest-energy composition within each phase. | For each phase, the FP-predicted minimum-energy composition is compared with DFT. The reported metric is the fraction of phases, across all tie-line systems, for which the two agree. |
| Hull-minimum agreement: global convex-hull minimum across the tie-line. | For each tie-line system, the FP-predicted convex-hull minimum is compared with DFT. A system with no structure below the tie line has no stable intermediate. The reported metric is the fraction of systems for which the two agree: they identify the same stable intermediate, or both find none and their tie lines are anchored on the same endpoint structures. |
| Name | Metrics |
|---|---|
| Average energy error | Energy error relative to DFT, reported as the MAE over all structures in the ordering benchmark. |
| Top-1 accuracy | Fraction of ordering groups for which the FP identifies the same lowest-energy configuration (LEC) as DFT. |
| Recall@k | Fraction of the k lowest-energy DFT orderings retained among the top k FP-ranked orderings within each ordering group. The reported metric is the mean over all ordering groups. |
| Spearman's ρ | Rank correlation between FP and DFT orderings within each ordering group. The reported metric is the mean over all ordering groups. |
| Rate of ranking errors | Fraction of configuration pairs whose FP and DFT rankings disagree, within each ordering group. The reported metric is the mean over all ordering groups. |
| ΔEDFT of misranked pairs | DFT energy difference between configurations that the FP misranks. For each ordering group, the mean and maximum ΔEDFT over its misranked pairs are computed and then averaged across all ordering groups. |
Applies to full FP relaxation only (static evaluation does not produce an FP-relaxed structure to compare).
| Name | Metrics |
|---|---|
| Structure relaxation error | For each structure, the RMSD between the FP-relaxed and DFT-relaxed structures is computed. The reported metric is the mean/maximum over all structures in the dataset. |
| Mean/max RMSD | Geometric deviation after alignment (Å). |
| Map success | Fraction of cases where FP and DFT structures remain sufficiently similar for meaningful RMSD evaluation; failures correspond to large structural deviation. |
| RMSD thresholds | Increasingly strict measures of fidelity. |
These metrics are complementary and should be interpreted together rather than combined into a single overall ranking. Click any column heading to sort by that metric.
Data availability. The FPBench convex-hull and elemental-ordering benchmark subsets are available on Figshare. The underlying ternary chalcogenide DFT data are available from Adams et al. on Dryad.
All DFT energies are raw, uncorrected VASP-PBE energies -- no Materials Project MP2020 compatibility correction is applied anywhere. Model versions and official sources for the seven FPs evaluated here are documented once on the FPBench home page. See the Phase_stability_ordering README for the full standardized-file schema.
Provided FPBench reference + new FP calculations
or
User DFT reference + user FP results
↓
standardized reference/results
↓
validation and analysis
↓
hull, ordering, and RMSD tables
Evaluate a new FP on the provided benchmark. Use generation/convexhull_ordering_run_generator.ipynb to generate per-FP jobs and submission scripts against the shipped DFT reference, run them on your cluster, merge the results, then load the merged file directly in the analysis notebook -- your FP appears in every table above.
Apply the analysis functions to another dataset. Call build_phase_stability_ordering_results(...) and the table builders in scripts/convexhull_analysis_utils.py directly with your own DFT reference and FP results.
git clone https://github.com/mogroupumd/FPBench.git
cd FPBench/Phase_stability_ordering
pip install -r requirements.txt
pip install jupyterlab
jupyter lab analysis/convexhull_ordering_analysis_all_models.ipynb
See the Phase_stability_ordering README for the full quick start, the generator and analysis notebooks for complete operational detail, and examples/ for a small runnable slice of real data.
Interested in evaluating a new foundation potential, or having it considered for inclusion in FPBench? See our Adding a Potential guide to integrate and evaluate a new model with FPBench. For inclusion in the public leaderboard, please contact Prof. Yifei Mo at yfmo@umd.edu with the model name, version/checkpoint, and a link to the official implementation or model weights.