Optimizing Reference Site Selection for Machine Learning in Heterogeneous Catalysis
Faculty Mentor Information
Dr. Oliviero Andreussi, Boise State University; and Dr. Jakob Filser, Boise State University
Presentation Date
7-16-2026
Abstract
Computational screening is widely used for discovering new catalysts by predicting adsorption properties before experimental synthesis. Identifying representative adsorption sites on disordered materials, however, remains a major computational challenge because of the large number of chemically distinct local environments that must be evaluated. MapSy, a newly developed Python library that combines symmetry-function descriptors with machine learning, is being used to characterize catalytic surfaces and accelerate the screening of complex materials. Using hydrogen adsorption on amorphous cobalt phosphide (CoP), a promising catalyst for the hydrogen evolution reaction (HER), as a benchmark system, we evaluated strategies for selecting representative adsorption sites for machine learning predictions. Specifically, centroid-based sampling of chemically similar adsorption-site clusters was compared with random sampling while varying the number of reference sites. Prediction accuracy was assessed by comparing reconstructed adsorption energy maps with reference calculations. We found that increasing the number of sampled sites consistently improves prediction accuracy, while centroid-based sampling more effectively captures the diversity of local chemical environments than random selection. These results provide practical guidelines for reducing the computational cost of HER catalyst screening with MapSy and establish a foundation for future adaptive sampling strategies that further improve efficiency without sacrificing predictive accuracy.
Optimizing Reference Site Selection for Machine Learning in Heterogeneous Catalysis
Computational screening is widely used for discovering new catalysts by predicting adsorption properties before experimental synthesis. Identifying representative adsorption sites on disordered materials, however, remains a major computational challenge because of the large number of chemically distinct local environments that must be evaluated. MapSy, a newly developed Python library that combines symmetry-function descriptors with machine learning, is being used to characterize catalytic surfaces and accelerate the screening of complex materials. Using hydrogen adsorption on amorphous cobalt phosphide (CoP), a promising catalyst for the hydrogen evolution reaction (HER), as a benchmark system, we evaluated strategies for selecting representative adsorption sites for machine learning predictions. Specifically, centroid-based sampling of chemically similar adsorption-site clusters was compared with random sampling while varying the number of reference sites. Prediction accuracy was assessed by comparing reconstructed adsorption energy maps with reference calculations. We found that increasing the number of sampled sites consistently improves prediction accuracy, while centroid-based sampling more effectively captures the diversity of local chemical environments than random selection. These results provide practical guidelines for reducing the computational cost of HER catalyst screening with MapSy and establish a foundation for future adaptive sampling strategies that further improve efficiency without sacrificing predictive accuracy.