FPBench · Force Prediction

FPBench: Force Prediction

Application-oriented force-error metrics for foundation potentials (FPs), evaluated on MatPES-PBE, MatPES-r2SCAN, and OMat24 rattled-1000.

View Leaderboard Use FPBench Paper (arXiv) GitHub
lower is better higher is better green = best in column cell shading: green = better, gold = middle, red = worse identifiers, badges, and thresholds are never colored

The summary leaderboard highlights average force error together with application-oriented metrics for highly accurate force predictions, large-force-error atoms, and far-from-equilibrium (FE) atoms. Highly accurate force predictions have force-magnitude errors |Δ|F|| < 0.01 eV/Å, large-force-error atoms have |Δ|F|| > 1 eV/Å, and FE atoms are defined by |FDFT| > 1 eV/Å. Full definitions are in Metrics, below the leaderboard.

The per-column best value (green) updates when you click a column header. FP badges show whether that FP's own training data included this dataset's functional/source (training), included it only during pretraining before the released checkpoint was fine-tuned onto a different dataset (pretraining), did not include it (OOD), or is a fine-tune of an OOD base model onto this dataset (fine-tuned).

Leaderboard

Combined summary
FP Average errorΔ|F|MAE / RMSE(eV/Å) ↓ Large-force-error atomsFrac(|Δ|F|| > 1 eV/Å)(%) ↓ FE atomsΔ|F|MAE / RMSE(eV/Å) ↓ FE atomsΔθMAE / RMSE(°) ↓ Highly accurate force predictionsFrac(|Δ|F|| < 0.01 eV/Å)(%) ↑ Joint force magnitude-angle accuracyFrac(<0.01 eV/Å& <1°/20°)(%) ↑ Average errorΔθMAE / RMSE(°) ↓ Excluding large-error atomsΔ|F|MAE / RMSE(eV/Å), on |Δ|F|| < 1 eV/Å Version (training)
Orb OOD0.17 / 1.122.000.31 / 2.0912 / 2412.61.4 / 8.930 / 520.14 / 0.22orb-v3-conservative-inf-omat-20250404 (OMat24)
SevenNet OOD0.19 / 0.912.630.39 / 1.7015 / 2811.21.6 / 7.433 / 550.15 / 0.247net-mf-ompa (modal mpa) (OMat24 + sAlex + MPtrj)
MatterSim OOD0.21 / 1.292.650.40 / 2.4116 / 287.91.00 / 4.537 / 590.17 / 0.26MatterSim-v1.0.0-5M (MatterSim dataset (PBE))
Nequix OOD0.21 / 1.522.830.43 / 2.8615 / 2710.71.6 / 6.933 / 550.16 / 0.25nequix-oam-1 (OMat24 + sAlex + MPtrj)
MACE OOD0.23 / 1.933.440.49 / 3.6518 / 298.341.07 / 4.5137 / 580.17 / 0.27>=v0.3.10 (MACE-MPA-0, medium) (MPtrj + sAlex)
CHGNet OOD0.31 / 2.195.710.70 / 4.1424 / 365.610.66 / 2.1645 / 660.22 / 0.32v0.3.0 (MPtrj)
GPTFF OOD0.42 / 1.899.821.00 / 3.5530 / 434.70.59 / 1.450 / 700.26 / 0.36gptff_v2 (Atomly (PBE))
M3GNet OOD0.44 / 1.4910.790.98 / 2.7640 / 564.560.54 / 1.4556 / 760.26 / 0.36MP-2021.2.8-PES (MP-2021.2.8)
UMA OOD0.56 / 7.932.691.13 / 11.8912 / 2515.132.79 / 11.4428 / 510.12 / 0.21s-1p1 (OC20+ODAC23+OMat24+OMC25+OMol25)
ALIGNN OOD functional mismatch2.08 / 24.8024.235.51 / 46.9561 / 772.50.20 / 0.5570 / 870.31 / 0.40alignnff_wt10 (JARVIS-DFT)OptB88vdW → PBE
M3GNet-MatPES training0.21 / 1.242.350.43 / 2.3418 / 277.140.92 / 3.7534 / 530.17 / 0.25v2025.1 (MatPES-PBE)
TensorNet-MatPES training0.17 / 1.391.200.33 / 2.6312 / 187.660.98 / 4.6529 / 490.15 / 0.22v2025.1 (MatPES-PBE)
MACE-MatPES training0.08 / 1.100.250.16 / 2.096 / 1019.023.88 / 15.3516 / 320.07 / 0.13>=v0.3.10 (MatPES-PBE (fine-tuned))
NEP89 OOD0.40 / 3.707.600.74 / 6.9931 / 443.20.31 / 1.352 / 730.27 / 0.36nep89_20250409 (mixed QM levels)
DPA4 OOD0.15 / 2.161.640.27 / 3.7110 / 2219.74.7 / 1625 / 480.11 / 0.20DPA4-Plus-OMat24-v20260805 (OMat24)
GRACE OOD0.16 / 1.492.080.30 / 2.8112 / 2517.83.7 / 1427 / 500.12 / 0.21GRACE-3L-OMAT-large-ft-AM (OMat24 → sAlex + MPtrj)
eqV2 OOD0.16 / 1.271.980.28 / 2.3911 / 2514.62.5 / 1228 / 510.12 / 0.21eqV2_31M_omat_mp_salex (OMat24 → MPtrj + sAlex)
eSEN OOD0.15 / 1.381.930.28 / 2.6011 / 2518.64.0 / 1527 / 510.12 / 0.21esen_30m_oam (OMat24 → MPtrj + sAlex)
Notes

Hover or click the model markers (* §) in the table for model-specific details.

* MatterSim’s training dataset is not publicly documented in detail. The model card states only that the released MatterSim-v1.0.0-5M checkpoint was trained on a PBE-based dataset of approximately 6M structures; the dataset is neither named nor released. Its OOD label therefore reflects the absence of any documented MatPES or OMat24 training exposure, not a verified dataset composition.

The FPBench checkpoint is nequix-oam-1. The official nequix repository documents this checkpoint as trained on OMat24, sAlex and MPtrj at the DFT (PBE+U) level. The Nequix paper (Koker, Kotak & Smidt, arXiv:2508.16067) describes a model trained on MPtrj and does not report a training-set size or parameter count for this released OAM checkpoint; the 707,569-parameter model size shown here was measured from the loaded checkpoint by FPBench.

§ NEP89 (nep89_20250409) is trained on eleven datasets computed at different quantum-mechanical levels rather than a single reference functional. Its scores on this board therefore measure practical agreement with the MatPES-PBE reference rather than agreement with a single matched reference functional. The eleven datasets and the 537,641-configuration training total are reported in the NEP89 paper.

ALIGNN (alignnff_wt10) is trained on JARVIS-DFT at the OptB88vdW level, not PBE. It is the only potential on this board whose training reference functional differs from the evaluation reference, so its scores measure agreement with the FPBench MatPES-PBE reference rather than a functional-matched fitting error.

Leaderboard

Combined summary
FP Average errorΔ|F|MAE / RMSE(eV/Å) ↓ Large-force-error atomsFrac(|Δ|F|| > 1 eV/Å)(%) ↓ FE atomsΔ|F|MAE / RMSE(eV/Å) ↓ FE atomsΔθMAE / RMSE(°) ↓ Highly accurate force predictionsFrac(|Δ|F|| < 0.01 eV/Å)(%) ↑ Joint force magnitude-angle accuracyFrac(<0.01 eV/Å& <1°/20°)(%) ↑ Average errorΔθMAE / RMSE(°) ↓ Excluding large-error atomsΔ|F|MAE / RMSE(eV/Å), on |Δ|F|| < 1 eV/Å Version (training)
M3GNet-MatPES-r2SCAN training0.23 / 0.652.870.41 / 1.0718 / 275.930.80 / 3.4830 / 480.19 / 0.27v2025.1 (MatPES-r2SCAN)
TensorNet-MatPES-r2SCAN training0.19 / 0.711.640.33 / 1.1713 / 196.390.79 / 4.1926 / 450.17 / 0.24v2025.1 (MatPES-r2SCAN)
MACE-MatPES-r2SCAN training0.12 / 1.340.540.20 / 2.278 / 1212.822.33 / 10.3517 / 330.10 / 0.17>=v0.3.10 (MatPES-r2SCAN (fine-tuned))

Leaderboard

Combined summary
FP Average errorΔ|F|MAE / RMSE(eV/Å) ↓ Large-force-error atomsFrac(|Δ|F|| > 1 eV/Å)(%) ↓ FE atomsΔ|F|MAE / RMSE(eV/Å) ↓ FE atomsΔθMAE / RMSE(°) ↓ Highly accurate force predictionsFrac(|Δ|F|| < 0.01 eV/Å)(%) ↑ Joint force magnitude-angle accuracyFrac(<0.01 eV/Å& <1°/20°)(%) ↑ Average errorΔθMAE / RMSE(°) ↓ Excluding large-error atomsΔ|F|MAE / RMSE(eV/Å), on |Δ|F|| < 1 eV/Å Version (training)
MACE OOD0.34 / 2.835.720.40 / 3.166 / 94.080.27 / 3.718 / 140.21 / 0.30>=v0.3.10 (MACE-MPA-0, medium) (MPtrj + sAlex)
CHGNet OOD0.76 / 4.0018.500.90 / 4.469 / 151.810.03 / 1.3413 / 220.34 / 0.42v0.3.0 (MPtrj)
M3GNet OOD0.91 / 1.9627.661.06 / 2.1813 / 211.300.01 / 0.8119 / 300.37 / 0.46MP-2021.2.8-PES (MP-2021.2.8)
UMA training0.09 / 0.220.510.10 / 0.242 / 513.144.76 / 12.974 / 80.08 / 0.14s-1p1 (OC20+ODAC23+OMat24+OMC25+OMol25)
M3GNet-MatPES OOD0.55 / 2.1812.430.64 / 2.439 / 142.300.06 / 1.7913 / 220.29 / 0.38v2025.1 (MatPES-PBE)
TensorNet-MatPES OOD0.60 / 4.0412.190.70 / 4.528 / 122.760.11 / 2.3011 / 190.27 / 0.36v2025.1 (MatPES-PBE)
MACE-MatPES pretraining0.43 / 2.258.940.51 / 2.516 / 94.120.31 / 3.748 / 150.22 / 0.31>=v0.3.10 (OMat24 → MatPES-PBE)
Orb training0.15 / 0.631.470.17 / 0.713 / 69.072.13 / 8.845 / 100.12 / 0.19orb-v3-conservative-inf-omat-20250404 (OMat24)
SevenNet training0.16 / 0.371.670.19 / 0.413 / 67.681.41 / 7.445 / 100.14 / 0.217net-mf-ompa (modal mpa) (OMat24 + sAlex + MPtrj)
MatterSim OOD0.26 / 0.893.170.30 / 0.985 / 94.210.34 / 3.818 / 150.20 / 0.27MatterSim-v1.0.0-5M (MatterSim dataset (PBE))
Nequix training0.25 / 2.333.210.29 / 2.614 / 75.880.69 / 5.616 / 110.17 / 0.25nequix-oam-1 (OMat24 + sAlex + MPtrj)
GPTFF OOD0.99 / 3.0027.601.18 / 3.3512 / 191.180.01 / 0.7217 / 270.39 / 0.47gptff_v2 (Atomly (PBE))
NEP89 OOD0.70 / 2.7114.290.83 / 3.0210 / 162.240.05 / 1.7114 / 240.30 / 0.38nep89_20250409 (mixed QM levels)
DPA4 training0.11 / 3.560.460.12 / 3.982 / 516.647.50 / 16.513 / 70.07 / 0.12DPA4-Plus-OMat24-v20260805 (OMat24)
GRACE pretraining0.17 / 1.671.980.20 / 1.873 / 610.312.83 / 10.144 / 90.11 / 0.19GRACE-3L-OMAT-large-ft-AM (OMat24 → sAlex + MPtrj)
eqV2 pretraining0.17 / 1.812.130.20 / 2.033 / 610.102.66 / 9.944 / 90.12 / 0.19eqV2_31M_omat_mp_salex (OMat24 → MPtrj + sAlex)
eSEN pretraining0.17 / 2.511.970.20 / 2.802 / 611.523.73 / 11.374 / 80.11 / 0.18esen_30m_oam (OMat24 → MPtrj + sAlex)

Metrics

Evaluated for atoms with |FDFT| > 0.01 eV/Å (except the all-atom average-error analysis, which uses every atom); far-from-equilibrium (FE) atoms are |FDFT| > 1 eV/Å.

NameMetrics
Average errorForce-magnitude error Δ|F| and force-angle error Δθ, reported as MAE/RMSE over all atoms or a selected subset.
Cumulative distribution functions (CDFs) of force errorsCDFs of |Δ|F||, Δθ, and norm of the force-vector error evec, over all atoms or a selected subset.
Highly accurate force predictions (small-force-error atoms)Fraction of atoms with very small values of |Δ|F|| and Δθ below threshold (e.g. |Δ|F|| < 0.01 eV/Å).
Joint force magnitude-angle accuracyFraction of atoms with simultaneously small values of |Δ|F|| and Δθ (e.g. |Δ|F|| < 0.01 eV/Å and Δθ < 1° or 20°).
Force-magnitude error excluding large-error atomsMAE/RMSE evaluated after excluding atoms with force-magnitude errors > 1 eV/Å, showing model accuracy outside the large-error tail. Reported on the summary leaderboard as Δ|F| MAE/RMSE on |Δ|F|| < 1 eV/Å.
Large-force-error atomsFraction of atoms with high values of |Δ|F|| and Δθ (e.g. |Δ|F|| > 0.5 eV/Å).
Force errors on far-from-equilibrium (FE) atomsMAE/RMSE evaluated for Δ|F|, Δθ over the FE atoms selected as |FDFT| > 1 eV/Å. Fraction of FE atoms with relative force-magnitude error rF below increasing thresholds.
Notation: Δ|F| = |FFP| − |FDFT| (force-magnitude error); |Δ|F|| for CDFs, thresholds, and fraction metrics; Δθ, the angle between the FP and DFT force vectors; evec = ||FFPFDFT|| (force-vector error); rF = |Δ|F|| / |FDFT| (relative force-magnitude error, FE atoms only). Atoms with zero FP or DFT force are excluded from every metric, since the force angle is undefined in that case.

These metrics are complementary and should be interpreted together rather than combined into a single overall ranking. Click any column heading to sort by that metric.

Detailed results

Per-dataset tables for the metrics defined above. These follow the dataset tab selected with the leaderboard, currently MatPES-PBE.

Highly accurate force predictions
Percentage of evaluated atoms satisfying |Δ|F|| < threshold. Higher is better.
FP< 0.01 eV/Å< 0.02 eV/Å< 0.05 eV/Å< 0.07 eV/Å< 0.1 eV/Å< 0.2 eV/Å< 0.5 eV/Å
Orb12.5822.4842.1050.5759.7376.4892.33
SevenNet11.2220.5039.5547.9657.1574.2291.05
MatterSim7.9115.0831.8540.1649.9269.4789.81
Nequix10.7419.6037.7145.7554.6471.7789.99
MACE8.3415.9233.1141.2150.3168.0388.05
CHGNet5.6110.9123.9330.7939.0657.2181.66
GPTFF4.669.0119.7225.3132.1648.0673.62
M3GNet4.568.8819.5425.3032.4248.8173.45
UMA15.1326.4046.8655.1363.7278.7392.22
ALIGNN2.555.0411.6615.5620.8734.8558.48
M3GNet-MatPES7.1413.8730.4238.9949.2069.7690.64
TensorNet-MatPES7.6614.8132.5341.8052.8674.7493.92
MACE-MatPES19.0233.5058.9268.4777.6690.9798.47
NEP893.206.2414.6419.8026.9646.3277.00
DPA419.6732.4453.4461.3669.3682.3893.87
GRACE17.7729.1648.8056.6764.9279.3192.65
eqV214.5726.6748.2456.6565.2479.6192.70
eSEN18.6330.9151.5059.3967.4080.7993.04
Joint force magnitude-angle accuracy
Percentage of evaluated atoms satisfying both |Δ|F|| < threshold and Δθ < 1° or 20°. Each cell reports the Δθ < 1° fraction (left) and the Δθ < 20° fraction (right). Higher is better.
FP< 0.01 eV/Å< 0.02 eV/Å< 0.05 eV/Å< 0.07 eV/Å< 0.1 eV/Å< 0.2 eV/Å< 0.5 eV/Å
Orb
SevenNet
MatterSim
Nequix
MACE
CHGNet
GPTFF
M3GNet
UMA
ALIGNN
M3GNet-MatPES
TensorNet-MatPES
MACE-MatPES
NEP89
DPA4
GRACE
eqV2
eSEN
Large-force-error atoms
Percentage of evaluated atoms with |Δ|F|| exceeding each threshold. Lower is better.
FP> 0.5 eV/Å> 1 eV/Å> 2 eV/Å> 3 eV/Å> 4 eV/Å> 5 eV/Å> 7 eV/Å> 10 eV/Å
Orb7.6722.0010.3070.1180.0630.0410.0240.015
SevenNet8.9482.6320.5300.2230.1190.0750.0410.025
MatterSim10.1912.6530.3770.1420.0850.0590.0370.024
Nequix10.0072.8260.4640.1580.0780.0500.0300.019
MACE11.9543.4380.5900.1940.0830.0460.0250.018
CHGNet18.3385.7070.9450.2900.1220.0690.0360.023
GPTFF26.3819.8212.1540.7270.3180.1690.0700.036
M3GNet26.55310.7892.7571.0560.4930.2610.1030.035
UMA7.7852.6861.0550.8080.7120.6540.5830.520
ALIGNN41.51924.23112.2368.0836.0274.7933.3662.304
M3GNet-MatPES9.3632.3550.4520.1590.0730.0400.0200.013
TensorNet-MatPES6.0801.1960.1750.0650.0430.0330.0240.018
MACE-MatPES1.5350.2490.0390.0210.0170.0140.0110.010
NEP8922.9977.6031.7300.6330.3180.2050.1270.085
DPA46.1301.6360.2580.1190.0810.0640.0460.034
GRACE7.3492.0760.2840.1070.0610.0420.0270.019
eqV27.3011.9810.2190.0760.0420.0320.0220.018
eSEN6.9591.9280.2490.1010.0620.0460.0320.022
Δ|F| MAE/RMSE across DFT force-magnitude subsets
Mean absolute error and root-mean-square error of Δ|F|, restricted to atoms with |FDFT| above each threshold. The top row shows the average percentage of atoms (across FPs) above each threshold. Each cell reports MAE (left) and RMSE (right). MAE and RMSE use separate color scales; lower is better for both. FE atoms correspond to |FDFT| > 1 eV/Å.
FPAll atoms> 0.01 eV/Å> 0.05 eV/Å> 0.1 eV/Å> 0.2 eV/Å> 0.5 eV/Å> 0.7 eV/Å> 1 eV/Å> 2 eV/Å
Atoms with |FDFT| > threshold100.0%97.0%86.0%76.7%66.8%48.7%38.7%26.8%8.3%
Orb
SevenNet
MatterSim
Nequix
MACE
CHGNet
GPTFF
M3GNet
UMA
ALIGNN
M3GNet-MatPES
TensorNet-MatPES
MACE-MatPES
NEP89
DPA4
GRACE
eqV2
eSEN
FE relative force-magnitude error (rF) thresholds
FE atoms correspond to |FDFT| > 1 eV/Å. rF = |Δ|F|| / |FDFT|. Percentage of FE atoms with rF below each threshold. Higher is better.
FP<0.01<0.05<0.1<0.2<0.3<0.4<0.5<1<2
Orb8.4835.2355.6676.1885.2989.8892.5199.5299.91
SevenNet6.3628.7048.3970.6581.6187.3490.6099.2099.80
MatterSim5.3824.7443.3466.7079.3286.3790.4399.4599.91
Nequix5.0023.1540.9464.3578.0085.7890.1199.5199.91
MACE4.0018.9034.5657.1871.8281.2687.2299.2999.89
CHGNet1.919.4318.4035.3650.8264.4475.6599.4799.93
GPTFF0.763.857.7816.1025.9037.5050.6399.5999.92
M3GNet1.075.3510.5420.9731.6542.9254.6798.3499.63
UMA11.0742.0462.0078.9385.6389.1191.2597.9898.59
ALIGNN0.582.845.6711.5717.9224.9832.9981.9187.76
M3GNet-MatPES3.9419.0635.2559.0474.6984.7791.2599.8399.99
TensorNet-MatPES5.2224.8745.1871.8585.7892.8896.4199.92100.00
MACE-MatPES13.9452.5475.3091.7596.6598.4099.1299.99100.00
NEP892.1210.4420.5839.3055.1768.1178.2398.0899.69
DPA415.8652.0270.5983.7088.7391.5093.2599.5599.89
GRACE10.6941.7462.5380.0687.0190.5192.5999.6499.94
eqV211.7145.2665.8481.7387.6790.7092.5699.7099.97
eSEN13.3748.1667.7382.3787.8890.7692.5599.6399.94
Highly accurate force predictions
Percentage of evaluated atoms satisfying |Δ|F|| < threshold. Higher is better.
FP< 0.01 eV/Å< 0.02 eV/Å< 0.05 eV/Å< 0.07 eV/Å< 0.1 eV/Å< 0.2 eV/Å< 0.5 eV/Å
M3GNet-MatPES-r2SCAN5.9311.6126.3434.4244.4265.7188.81
TensorNet-MatPES-r2SCAN6.3912.4528.1736.7847.5370.1692.17
MACE-MatPES-r2SCAN12.8223.8146.8356.7867.2284.6196.91
Joint force magnitude-angle accuracy
Percentage of evaluated atoms satisfying both |Δ|F|| < threshold and Δθ < 1° or 20°. Each cell reports the Δθ < 1° fraction (left) and the Δθ < 20° fraction (right). Higher is better.
FP< 0.01 eV/Å< 0.02 eV/Å< 0.05 eV/Å< 0.07 eV/Å< 0.1 eV/Å< 0.2 eV/Å< 0.5 eV/Å
M3GNet-MatPES-r2SCAN
TensorNet-MatPES-r2SCAN
MACE-MatPES-r2SCAN
Large-force-error atoms
Percentage of evaluated atoms with |Δ|F|| exceeding each threshold. Lower is better.
FP> 0.5 eV/Å> 1 eV/Å> 2 eV/Å> 3 eV/Å> 4 eV/Å> 5 eV/Å> 7 eV/Å> 10 eV/Å
M3GNet-MatPES-r2SCAN11.1892.8660.5380.1890.0850.0430.0180.009
TensorNet-MatPES-r2SCAN7.8341.6360.2540.0970.0580.0410.0240.014
MACE-MatPES-r2SCAN3.0910.5430.0870.0420.0280.0210.0120.007
Δ|F| MAE/RMSE across DFT force-magnitude subsets
Mean absolute error and root-mean-square error of Δ|F|, restricted to atoms with |FDFT| above each threshold. The top row shows the average percentage of atoms (across FPs) above each threshold. Each cell reports MAE (left) and RMSE (right). MAE and RMSE use separate color scales; lower is better for both. FE atoms correspond to |FDFT| > 1 eV/Å.
FPAll atoms> 0.01 eV/Å> 0.05 eV/Å> 0.1 eV/Å> 0.2 eV/Å> 0.5 eV/Å> 0.7 eV/Å> 1 eV/Å> 2 eV/Å
Atoms with |FDFT| > threshold100.0%96.4%89.0%84.1%76.6%57.8%46.6%33.3%10.8%
M3GNet-MatPES-r2SCAN
TensorNet-MatPES-r2SCAN
MACE-MatPES-r2SCAN
FE relative force-magnitude error (rF) thresholds
FE atoms correspond to |FDFT| > 1 eV/Å. rF = |Δ|F|| / |FDFT|. Percentage of FE atoms with rF below each threshold. Higher is better.
FP<0.01<0.05<0.1<0.2<0.3<0.4<0.5<1<2
M3GNet-MatPES-r2SCAN3.9819.2035.5059.7975.3985.4291.7299.8399.99
TensorNet-MatPES-r2SCAN4.9923.9843.9670.7085.0792.4896.2299.8999.99
MACE-MatPES-r2SCAN10.1642.4265.9686.9294.4797.4498.7299.97100.00
Highly accurate force predictions
Percentage of evaluated atoms satisfying |Δ|F|| < threshold. Higher is better.
FP< 0.01 eV/Å< 0.02 eV/Å< 0.05 eV/Å< 0.07 eV/Å< 0.1 eV/Å< 0.2 eV/Å< 0.5 eV/Å
MACE4.088.1519.7526.7736.0958.1683.88
CHGNet1.813.608.8912.3217.2331.7360.55
M3GNet1.302.606.499.0512.8524.4950.02
UMA13.1424.9651.4363.0574.5790.2998.05
M3GNet-MatPES2.304.5911.3015.6321.8739.7670.26
TensorNet-MatPES2.765.5213.5718.6825.7844.6572.94
MACE-MatPES4.128.1919.6226.3435.0755.2279.65
Orb9.0717.6538.8549.4261.2881.4795.12
SevenNet7.6815.0834.1944.2956.0977.5693.83
MatterSim4.218.3620.4027.9038.0062.4189.04
Nequix5.8811.6427.2536.0346.9169.2990.02
GPTFF1.182.335.858.1511.6122.6448.84
NEP892.244.4711.0715.3321.4338.8668.60
DPA416.6431.0060.2371.5481.6993.4498.47
GRACE10.3119.9142.9253.8765.3882.8194.40
eqV210.1019.4741.9252.6664.2482.0094.03
eSEN11.5222.0145.8656.7067.7783.7894.52
Joint force magnitude-angle accuracy
Percentage of evaluated atoms satisfying both |Δ|F|| < threshold and Δθ < 1° or 20°. Each cell reports the Δθ < 1° fraction (left) and the Δθ < 20° fraction (right). Higher is better.
FP< 0.01 eV/Å< 0.02 eV/Å< 0.05 eV/Å< 0.07 eV/Å< 0.1 eV/Å< 0.2 eV/Å< 0.5 eV/Å
MACE
CHGNet
M3GNet
UMA
M3GNet-MatPES
TensorNet-MatPES
MACE-MatPES
Orb
SevenNet
MatterSim
Nequix
GPTFF
NEP89
DPA4
GRACE
eqV2
eSEN
Large-force-error atoms
Percentage of evaluated atoms with |Δ|F|| exceeding each threshold. Lower is better.
FP> 0.5 eV/Å> 1 eV/Å> 2 eV/Å> 3 eV/Å> 4 eV/Å> 5 eV/Å> 7 eV/Å> 10 eV/Å
MACE16.1155.7171.5950.7160.4070.2660.1450.082
CHGNet39.44718.5026.8213.5452.1701.4610.7750.376
M3GNet49.97727.65610.8475.1112.7361.6650.7290.279
UMA1.9470.5140.1210.0460.0240.0140.0070.003
M3GNet-MatPES29.73712.4324.0421.9271.1170.7190.3740.191
TensorNet-MatPES27.06312.1924.6892.5611.6151.1140.6310.337
MACE-MatPES20.3518.9393.1911.5820.9380.6200.3280.172
Orb4.8791.4720.4150.1980.1220.0830.0410.021
SevenNet6.1681.6670.3460.1240.0650.0390.0200.010
MatterSim10.9653.1710.8760.4230.2620.1800.1060.062
Nequix9.9813.2130.8800.4180.2560.1780.1020.058
GPTFF51.15827.59911.1815.8393.4822.2521.0980.475
NEP8931.40314.2905.8793.5582.4891.8681.1680.664
DPA41.5250.4560.1300.0640.0450.0350.0270.021
GRACE5.5961.9850.6400.3270.2110.1480.0870.053
eqV25.9732.1310.6740.3430.2190.1570.0920.054
eSEN5.4831.9740.6310.3210.2040.1430.0840.049
Δ|F| MAE/RMSE across DFT force-magnitude subsets
Mean absolute error and root-mean-square error of Δ|F|, restricted to atoms with |FDFT| above each threshold. The top row shows the average percentage of atoms (across FPs) above each threshold. Each cell reports MAE (left) and RMSE (right). MAE and RMSE use separate color scales; lower is better for both. FE atoms correspond to |FDFT| > 1 eV/Å.
FPAll atoms> 0.01 eV/Å> 0.05 eV/Å> 0.1 eV/Å> 0.2 eV/Å> 0.5 eV/Å> 0.7 eV/Å> 1 eV/Å> 2 eV/Å
Atoms with |FDFT| > threshold100.0%100.0%100.0%99.8%99.0%93.4%88.2%80.0%56.6%
MACE
CHGNet
M3GNet
UMA
M3GNet-MatPES
TensorNet-MatPES
MACE-MatPES
Orb
SevenNet
MatterSim
Nequix
GPTFF
NEP89
DPA4
GRACE
eqV2
eSEN
FE relative force-magnitude error (rF) thresholds
FE atoms correspond to |FDFT| > 1 eV/Å. rF = |Δ|F|| / |FDFT|. Percentage of FE atoms with rF below each threshold. Higher is better.
FP<0.01<0.05<0.1<0.2<0.3<0.4<0.5<1<2
MACE9.7944.2770.7690.4896.1298.1499.0399.93100.00
CHGNet3.8118.7236.1363.6080.6389.8894.6999.6599.98
M3GNet3.0414.9729.0352.1168.1778.7885.6097.1899.70
UMA36.8783.7894.4798.3899.2299.5799.7499.97100.00
M3GNet-MatPES5.6527.0549.0175.2487.4593.5796.7799.8699.99
TensorNet-MatPES6.2629.9453.0878.3689.2394.4297.1299.8799.98
MACE-MatPES9.0240.8065.3985.6192.9496.4398.2499.97100.00
Orb24.2273.8990.2797.3498.8499.3999.6599.97100.00
SevenNet19.6069.3088.5796.9498.7399.3699.6399.97100.00
MatterSim12.4851.4175.8992.4897.1198.6999.3399.95100.00
Nequix13.7457.7082.5895.3098.1599.1199.5299.97100.00
GPTFF2.3511.8523.7446.3865.2378.4586.9899.7699.97
NEP895.3625.5146.0071.7884.8991.7195.4399.7799.98
DPA444.6888.2695.9998.6499.3199.6099.7699.98100.00
GRACE25.5976.0491.7197.5398.8399.3599.6199.97100.00
eqV223.9675.4091.5997.4898.7899.3299.5999.97100.00
eSEN27.7778.0792.5197.6498.8499.3599.6199.97100.00

Benchmark Datasets

MatPES-PBE
DFT functional: PBE
Equilibrium and off-equilibrium structures selected from finite-temperature MD trajectories, augmented with Materials Project ground-state structures.
MatPES-r2SCAN
DFT functional: r2SCAN
The r2SCAN release of MatPES, evaluated against the three r2SCAN-trained FPs.
OMat24 rattled-1000
DFT functional: PBE
The OMat24 rattled-1000 validation subset (approximately 117,000 structures rattled at a sampling temperature of 1000 K), evaluated against the FPs included in the FPBench leaderboard.

Model versions and official sources for the evaluated FPs are documented once on the FPBench home page.

Use FPBench

Paired Cartesian DFT and FP forces
                  or
Full generator calculations on a dataset
                   ↓
        standardized force_results
                   ↓
          validation and analysis
                   ↓
       FPBench force-error tables

Build standardized force results directly from paired Cartesian DFT and FP forces. Call build_force_results(dft_forces, fp_forces, structure_ids=None) from scripts/force_results.py with your own data.

Use a generator notebook as a template for full dataset/cluster generation. Register your FP's checkpoint and calculator setup in one of the three generator notebooks' POTENTIAL_REGISTRY, run the jobs on your cluster (the notebooks never submit jobs by themselves), then run the matching analysis notebook -- your FP appears in every table above.

git clone https://github.com/mogroupumd/FPBench.git
cd FPBench/Force_error
pip install -r requirements.txt
pip install jupyterlab
jupyter lab analysis/force_error_analysis_matpes_pbe.ipynb

See the Force_error README for the full quick start, the generator and analysis notebooks for complete operational detail, and the included Cartesian-force example for a small runnable slice of real data.

Contribute

Interested in evaluating a new foundation potential, or having it considered for inclusion in FPBench? See our Adding a Potential guide to integrate and evaluate a new model with FPBench. For inclusion in the public leaderboard, please contact Prof. Yifei Mo at yfmo@umd.edu with the model name, version/checkpoint, and a link to the official implementation or model weights.