AI Models & Platforms
Illumina Launches SpliceAI2 Model for Splice Variant Interpretation

Illumina introduced SpliceAI2 on October 8, 2026, a genomic AI model from its BioInsight AI Lab that predicts how genetic variants alter RNA splicing. Illumina said the model identified 17% more disease-relevant variants than other splicing models when applied to a rare disease research dataset.
According to the company’s announcement, SpliceAI2 is designed to help researchers identify disease-relevant splice variants that might otherwise go unnoticed, with splice effect prediction key to resolving variants of uncertain significance in rare disease research, hereditary cancer testing research, and drug discovery. The model joins PromoterAI and PrimateAI-3D in a suite of genomic AI models covering splice, promoter, and missense variant effect types, and Illumina said the three collectively now enable researchers to identify up to twice as many variants with predicted biological impact.
Rami Mehio, senior vice president and general manager of BioInsight, described variant effect prediction tools such as SpliceAI2 as one of the BioInsight AI Lab’s key focus areas; the lab works on generating genomic and multiomic data, developing genomic AI models for variant effect prediction and prioritization, and training biological foundation models. Kyle Farh, vice president of the BioInsight AI Lab, said Illumina is advancing AI “to systematically shrink the portion of the genome that remains uninterpretable.”
SpliceAI2 succeeds the original SpliceAI, which the lab released in 2019. Illumina said the first-generation model has been cited in more than 3,400 publications and is incorporated into splice variant interpretation recommendations from ClinGen, a clinical research body that sets standards for clinical genomics. The new model was trained on a dataset more than 100 times larger than its predecessor’s.
Model Design and Training Data
Where the original SpliceAI predicted whether a cell would splice at a given location, SpliceAI2 addresses three questions, according to Illumina’s technical article: which positions in a gene are used as splice sites and how frequently, which splice sites connect through splice junctions, and which full-length RNA transcript isoforms are produced. The model requires only a DNA sequence as input, which Illumina said enables transcript-level analysis without RNA data from difficult-to-obtain tissues.
A manuscript detailing the work describes a model that processes 196,608 base pairs of genomic sequence with roughly 13 million trainable parameters, integrating splice site and junction predictions into a splice graph from which complete transcripts and their usage are inferred. SpliceAI2 was trained end-to-end on 314,745 RNA sequencing samples spanning human and nine additional mammalian species, covering more than 46 million observed splice junctions after filtering, together with 330 long-read RNA sequencing samples from the public ENCODE project. The authors report that the model exactly reconstructed the most abundant transcript for 82% of held-out genes, compared with 78% when trained without long-read data.
The manuscript also describes a fine-tuning framework that conditions predictions on the expression levels of 147 RNA binding proteins, capturing tissue-specific splicing differences across 48 Genotype-Tissue Expression (GTEx) tissues in nearly 15 million differential splice site usage measurements. The authors report that the model independently learned sequence motifs recognized by real splicing regulators without being explicitly taught those relationships, and that the framework was adapted to disease states including SF3B1-mutant tumors and myotonic dystrophy.
Benchmarks and Rare Disease Findings
Across three independent benchmarks, the manuscript reports that SpliceAI2 outperformed the original SpliceAI, Pangolin, and Google DeepMind’s AlphaGenome, with the AlphaGenome comparisons run independently by collaborators at the University of Oxford. SpliceAI2 scored an auPRC of 0.77 versus 0.66 for the next best model on GTEx cryptic splice variant detection, a Spearman correlation of 0.63 versus 0.47 on splice site usage quantification, an auROC of 0.76 versus 0.73 on splicing quantitative trait locus classification, and a Spearman correlation of 0.59 versus 0.55 on the OpenSplice massively parallel reporter assay benchmark. Illumina’s announcement states that the model improved quantification of splice site usage by 34% compared with the next best model.
In an analysis of 7,504 probands from the Genomics England 100,000 Genomes Project, the manuscript reports that variants prioritized by SpliceAI2 were significantly enriched in phenotype-matched disease genes, identifying 17% more disease-associated variants than any other tested splicing model at matched confidence thresholds and recovering 133 excess variants versus 114 for the next best model at a fixed odds ratio of 2. Illumina reported that SpliceAI2 found 33% more disease-relevant splice variants than the legacy SpliceAI at a 2X confidence interval and 66% more at a 4X interval. Roughly half of the cryptic splice variants identified were located deep within intronic regions, and the manuscript states that variants more than 50 base pairs into introns would typically be missed by exome sequencing. Predicted splice-altering variants accounted for 15% of the excess genetic burden in the cohort, according to the manuscript.
Population-scale validation reported in the manuscript drew on more than 627,000 genomes from gnomAD, TOPMed, and UK Biobank, where variants assigned the highest SpliceAI2 scores were strongly depleted, approaching the depletion observed for protein-truncating loss-of-function mutations. UK Biobank proteomic data from 36,764 participants showed that carriers of higher-scoring variants had lower plasma protein levels, with a Pearson correlation of -0.50, the strongest among the models evaluated. RNA sequencing data from 5,435 Genomics England participants validated predictions in individual cases, including a PEX1 donor-loss variant associated with rod-cone dystrophy and a PKD1 cryptic donor variant associated with cystic kidney disease.
Availability and Licensing
Illumina said SpliceAI2 is accessible through its BioInsight Platform applications, including DRAGEN Annotation, Emedgene, and Illumina Connected Insights. Source code, trained model weights, and precomputed predictions for all possible single-nucleotide variants within human gene bodies (4 billion) and indels observed in human populations (150 million) are available through the SpliceAI2 GitHub repository and Hugging Face for academic and non-commercial research use, with a separate contact for commercial licensing. The software installs through PyPI and requires a CUDA-capable GPU.
According to the repository documentation, the spliceai2_summary_score ranges from 0 to 1, with recommended thresholds of 0.1 for high recall, 0.25 for balanced precision and recall, and 0.5 for high precision, corresponding to SpliceAI equivalents of 0.2, 0.5, and 0.8. The SpliceAI2 work was led by Kishore Jaganathan and Kyle Farh, with collaborators at the University of Oxford, UCSF, and the New York Genome Center.












