Products & Services

Matlantis PFP v9 Ties for First Overall on the MLIP Arena Benchmark

Version 9 of Matlantis PFP, the universal machine learning interatomic potential (MLIP) behind Matlantis™, is tied for first place overall on MLIP Arena, an open benchmark for machine learning interatomic potentials. Full results were published on the Tech Blog of Preferred Networks (PFN), which develops Matlantis PFP.

MLIP Arena (Chiang et al., NeurIPS 2025) looks past energy and force prediction errors, asking instead whether a model reproduces physically meaningful behavior in practical simulation settings. It covers five tasks: homonuclear diatomics, bulk equation of state, energy-volume scans, stability, and combustion.

The PFP development team ran all five tasks with v9 and compared the results against MLIP Arena leaderboard values as of March 1, 2026. PFP v9 is tied for first overall with MACE-MPA, and ranks first in three individual tasks: diatomics, equation of state, and stability. Leading three of the five tasks means PFP v9 takes more tasks than any single opponent can, so no model in the arena currently comes out ahead of it one on one.

Overall ranking on MLIP Arena. Matlantis PFP v9 is tied for first place with MACE-MPA. Values for other models are from the public MLIP Arena leaderboard as of March 1, 2026. (Source: Table 1, “Introducing Matlantis-PFP v9,” PFN Tech Blog)

The evaluation used the PBE calculation mode, since no public r2SCAN-based benchmark exists yet. The r2SCAN mode, which improves agreement with experimental values, is where v9 improved most, now covering all 96 elements from hydrogen to curium. In an additional hydrogen-combustion benchmark, switching from PBE to r2SCAN cut the root-mean-square error of reaction energies by 56%.

As universal MLIPs move into practical use, benchmarks have to cover long-time simulation behavior and agreement with experiment, not regression accuracy alone. MLIP Arena is a pioneering effort in that direction, and it aligns with how Matlantis and PFN are building out the evaluation basis for PFP: avoiding over-optimization to any single DFT functional and prioritizing comparison against experimental values.

Read the full articles here: