Study uses alternative splicing data to predict protein function

A single human gene can produce multiple protein isoforms through alternative splicing, greatly expanding the functional diversity of the proteome. Yet despite decades of research, scientists still struggle to determine the specific biological functions of individual isoforms, many of which differ by only subtle sequence changes but perform remarkably different roles in cells.

A new study published in Computational Biomedicine presents a computational framework that leverages the biological information embedded in alternative splicing events to improve protein isoform function prediction, offering fresh insights into one of molecular biology's longstanding challenges.

Alternative splicing affects more than 95% of human multi-exon genes and plays essential roles in development, tissue specialization, and disease. Aberrant splicing has been linked to numerous disorders, including cancer, neurodegenerative diseases, and inherited genetic conditions. However, while sequencing technologies have revealed millions of transcript isoforms, experimentally characterizing the function of each variant remains impractical due to the enormous time and cost involved.

Existing computational approaches typically rely on protein sequence similarity or gene-level annotations. Although useful, these methods often overlook the biological significance of individual splicing events, limiting their ability to distinguish closely related isoforms with distinct functions.

To address this challenge, the researchers developed SpliceEM, a computational framework that integrates alternative splicing information with protein sequences, functional annotations, and molecular interaction data. Rather than treating all transcript variants equally, the framework explicitly models how different types of splicing events contribute to functional divergence between isoforms.

The study demonstrated that incorporating alternative splicing information substantially improved the prediction of protein isoform functions across multiple benchmark datasets. The framework consistently outperformed existing computational approaches, particularly for biological functions with limited experimental annotations, where accurate prediction is often most difficult.

Beyond improved predictive performance, the study also provides biological insights into how alternative splicing shapes protein function.

By examining the model's learned representations, the researchers found that skipped exons (SE) and alternative first exons (AF) contributed disproportionately to functional divergence among protein isoforms. These splicing events were strongly associated with signaling pathways involved in cancer, including the MAPK and JAK–STAT pathways, suggesting that localized RNA splicing changes may have widespread consequences for cellular regulation and disease development.

The framework further demonstrated its ability to distinguish functional differences among isoforms originating from the same gene. Case analyses revealed that individual transcript variants can participate in distinct biological processes depending on their splicing patterns, highlighting the importance of studying proteins at the isoform level rather than relying solely on gene-level analyses.

According to the researchers, these findings support the growing view that alternative splicing is not simply a mechanism for generating transcript diversity, but also an important source of functional information that can improve computational annotation of the proteome.

As large-scale transcriptomic and single-cell sequencing datasets continue to expand, approaches capable of interpreting isoform-specific biology will become increasingly important. The authors suggest that incorporating biologically meaningful splicing information into computational analyses could accelerate studies of disease mechanisms, functional genomics, and biomarker discovery, while providing a more refined understanding of how transcript diversity contributes to human health and disease.

Although additional experimental validation will be needed for newly predicted isoform functions, the study establishes a biologically informed framework for exploring one of the least understood dimensions of gene regulation.

Source:
Journal reference:

Gu, T., & Wang, J. (2026). Isoform function prediction via knowledge distillation from alternative splicing. Computational Biomedicine. DOI: 10.70401/cbm.2026.0019. https://www.sciexplor.com/cbm/articles/cbm.2026.0019

Posted in: Genomics

Comments

The opinions expressed here are the views of the writer and do not necessarily reflect the views and opinions of News Medical.
Post a new comment
Post

While we only use edited and approved content for Azthena answers, it may on occasions provide incorrect responses. Please confirm any data provided with the related suppliers or authors. We do not provide medical advice, if you search for medical information you must always consult a medical professional before acting on any information provided.

Your questions, but not your email details will be shared with OpenAI and retained for 30 days in accordance with their privacy principles.

Please do not ask questions that use sensitive or confidential information.

Read the full Terms & Conditions.

You might also like...
Researchers discover smart molecular glue that activates inside oxidative cancer cells