A call for greater transparency in artificial intelligence models used for protein research has been issued by scientists at the Centre for Genomic Regulation (CRG). In a perspective paper published on May 12, 2026, in Nature Machine Intelligence, researchers analyzed the application of explainable artificial intelligence (XAI) techniques to protein language models (pLMs). These AI tools are capable of designing novel protein structures with potential applications ranging from carbon capture to industrial catalysis.
However, the CRG researchers noted that pLMs largely function as black boxes. This lack of transparency makes it difficult to understand how these models arrive at their predictions, raising concerns about their reliability, potential biases, and safety for real-world applications. Dr. Noelia Ferruz, a Group Leader at CRG and corresponding author of the paper, stated that while protein language models are advancing rapidly, the understanding of fundamental biological processes has not kept pace. She added that in some ways, the transparency of older, physics-based models has been lost.
The paper, titled "Towards the Explainability of Protein Language Models," surveys existing applications of XAI in protein research. The authors examined how XAI tools are currently used with pLMs, reviewing scientific literature and dozens of studies. They identified four key areas within the protein AI modeling workflow where explainability can be applied: the training data, user inputs, the internal model architecture, and the relationships between inputs and outputs.
The researchers categorized the potential roles of XAI in protein research into five types: Evaluator, Multitasker, Engineer, Coach, and Teacher. Currently, the Evaluator role is the most widely adopted, where XAI is used to verify and support existing findings. However, the study suggests that a limited number of studies are beginning to use XAI insights in more active roles, such as an "Engineer" or "Coach," to refine model architectures and guide the generation of protein sequences with desired traits.
The paper emphasizes that explainability should not be an afterthought but an integral part of the development process if pLMs are to become trustworthy partners in scientific discovery and design. Andrea Hunklinger, the paper's first author, stressed that without better methods to explain what these models learn and how they make decisions, there is a risk of developing powerful tools that cannot be fully trusted.
The study also points out that while AI can predict protein structure and function with high accuracy, the underlying mechanisms remain unclear. Understanding these internal workings could help researchers select more appropriate models for specific tasks and streamline the process of identifying new drugs or vaccine targets.
The CRG researchers are calling for the research community to prioritize transparency and develop more robust XAI tools specifically tailored for protein language models. They envision a future where AI-driven protein research not only provides answers but also inspires new hypotheses, supported by explainable frameworks accessible to scientists across different computational expertise levels. The ultimate validation of any AI-derived insight must occur in the laboratory, translating computational patterns into experimentally confirmed biological knowledge.
