Optimizing Reference Site Selection for Machine Learning in Heterogeneous Catalysis

Faculty Mentor Information

Dr. Oliviero Andreussi, Boise State University; and Dr. Jakob Filser, Boise State University

Presentation Date

7-16-2026

Abstract

Computational screening is widely used for discovering new catalysts by predicting adsorption properties before experimental synthesis. Identifying representative adsorption sites on disordered materials, however, remains a major computational challenge because of the large number of chemically distinct local environments that must be evaluated. MapSy, a newly developed Python library that combines symmetry-function descriptors with machine learning, is being used to characterize catalytic surfaces and accelerate the screening of complex materials. Using hydrogen adsorption on amorphous cobalt phosphide (CoP), a promising catalyst for the hydrogen evolution reaction (HER), as a benchmark system, we evaluated strategies for selecting representative adsorption sites for machine learning predictions. Specifically, centroid-based sampling of chemically similar adsorption-site clusters was compared with random sampling while varying the number of reference sites. Prediction accuracy was assessed by comparing reconstructed adsorption energy maps with reference calculations. We found that increasing the number of sampled sites consistently improves prediction accuracy, while centroid-based sampling more effectively captures the diversity of local chemical environments than random selection. These results provide practical guidelines for reducing the computational cost of HER catalyst screening with MapSy and establish a foundation for future adaptive sampling strategies that further improve efficiency without sacrificing predictive accuracy.

This document is currently not available here.

Share

COinS
 

Optimizing Reference Site Selection for Machine Learning in Heterogeneous Catalysis

Computational screening is widely used for discovering new catalysts by predicting adsorption properties before experimental synthesis. Identifying representative adsorption sites on disordered materials, however, remains a major computational challenge because of the large number of chemically distinct local environments that must be evaluated. MapSy, a newly developed Python library that combines symmetry-function descriptors with machine learning, is being used to characterize catalytic surfaces and accelerate the screening of complex materials. Using hydrogen adsorption on amorphous cobalt phosphide (CoP), a promising catalyst for the hydrogen evolution reaction (HER), as a benchmark system, we evaluated strategies for selecting representative adsorption sites for machine learning predictions. Specifically, centroid-based sampling of chemically similar adsorption-site clusters was compared with random sampling while varying the number of reference sites. Prediction accuracy was assessed by comparing reconstructed adsorption energy maps with reference calculations. We found that increasing the number of sampled sites consistently improves prediction accuracy, while centroid-based sampling more effectively captures the diversity of local chemical environments than random selection. These results provide practical guidelines for reducing the computational cost of HER catalyst screening with MapSy and establish a foundation for future adaptive sampling strategies that further improve efficiency without sacrificing predictive accuracy.