Machine learning (ML) and artificial intelligence (AI) are changing how catalysis science is performed, with the prospect of accelerating the discovery of catalysis knowledge. Despite the substantial efforts and resources invested in the development and application of AI/ML approaches, their performance still exhibits clear limitations due to the lack of usable data. Although decades of work have produced an abundance of catalysis data, much of it was generated without sufficient understanding of AI/ML or community-wide consensus on terminology and reporting standards. As a result, not all of this data is suitable for training models. It is therefore essential to define what constitutes AI-ready data and understand the efforts required to obtain it, in order to prepare for the era of AI-driven catalysis. In this perspective, we define the essential properties of AI-ready data that enable the construction of robust catalysis AI/ML models. Then, we discuss current practices in computational and experimental catalysis through the lens of data readiness for AI/ML applications. Lastly, we discuss community-level efforts needed to establish large-scale datasets ready for AI-driven catalysis research.
Data as the Backbone of Artificial Intelligence: Insights from Heterogeneous Catalysis
Year of publication
2026
Journal
ACS Catalysis
Issue
16
Volume
16
Starting page
15446
Ending page
15464
Research Areas
SUNCAT People