RNA-seq-Derived Molecular Phenotypes
Summary¶
A single RNA-seq experiment can be reduced to many different molecular phenotypes, and the choice is consequential rather than cosmetic. Intron excision ratios, productive-versus-unproductive junction usage, haplotype-resolved read counts, and raw coverage are all legitimate readouts of the same reads, but they define different QTLs, support different mechanistic claims, and constitute different benchmarks. Sequence-based variant-effect models are scored against whichever phenotype the benchmark authors chose — so "state of the art on sQTL prediction" is a statement about agreement with one quantification scheme, not about splicing in general.
The phenotype menu¶
| Phenotype | Tool | What it measures | What it cannot see |
|---|---|---|---|
| Intron excision ratio within a cluster | LeafCutter | relative usage of overlapping introns sharing a donor or acceptor, annotation-free [1] | alternative TSS and alternative polyadenylation — neither is captured by intron excision [1] |
| Productive vs unproductive junction usage | LeafCutter2 | whether a junction can participate in a complete ORF, or produces a PTC and triggers NMD [2] | events where NMD depletion is total in polyA data |
| Haplotypic expression / allelic imbalance | phASER | unique reads assigned per haplotype, within an individual [3] | population-level average effects; requires heterozygosity |
| Splice donor/acceptor probability | SpliceAI (prediction, GENCODE-trained) | per-base probability that a position is a splice site [4] | competition between sites; junction identity; usage quantity |
| Splice site usage, per tissue | Pangolin (prediction, SpliSER-trained) | probability and fraction of transcripts using each site, in 4 tissues [7] | intron-level clustering; unannotated complex events |
| RNA-seq coverage itself | Borzoi (prediction) | the raw shape, from which expression, splicing and polyadenylation statistics are derived [5] | phenotypes confounded by 3′-end and GC coverage bias [5] |
Why the choice changes the answer¶
Yield. LeafCutter found 5,774 sQTLs at 5% FDR in 372 GEUVADIS LCLs where the original transcript-ratio analysis found 620 — a ninefold difference from re-defining the phenotype on the same data, and 1.4–2.1× more than Cufflinks2 or Altrans under identical downstream processing.[1] Phenotype definition is not a preprocessing detail; it is a power decision.
Apparent tissue sharing. Using intron-excision ratios, 75–93% of sQTLs replicate across GTEx tissue pairs, against 9–48% previously reported for the same data.[1] A claim about how tissue-specific splicing regulation is turns out to depend heavily on how splicing was quantified.
Mechanistic interpretability. A p-sQTL and a u-sQTL are the same statistical object with entirely different biology. Alleles increasing unproductive splicing have strong negative effects on host-gene expression across the majority of 49 GTEx tissues, whereas productive-splicing QTL effects show no such correlation in any tissue.[2] Splitting one phenotype into two recovered a mechanism that the merged phenotype averages away.
What is invisible. 31.5% of alternatively excised introns detected across 14 GTEx tissues are absent from GENCODE, Ensembl and UCSC — 48.5% in testis.[1] Any predictor whose output space is keyed to an annotated transcript model cannot represent these events at all, regardless of how good its sequence understanding is.
The benchmark consequence¶
Synthesis: the sQTL benchmarks used to rank AlphaGenome, Borzoi, Pangolin and SpliceAI are built on fine-mapped intron-excision-ratio QTLs [1,5]. This creates a mismatch that none of the cited papers discusses directly: a splice-site probability model is being evaluated on an intron-usage phenotype, so part of what looks like a modelling deficit may be a units mismatch. It also explains a reported pattern — Borzoi, which predicts coverage and therefore reads out usage naturally, beats Pangolin within 200 bp of a junction, while Pangolin, which scores site probability, wins on distant de novo splice-gain variants where a new site simply appears.[5] Each model is strongest where the benchmark phenotype most resembles its own output space. This interpretation is inference, not a stated conclusion of either paper.
Two practical implications follow:
- Read what the benchmark measured before believing a ranking. "auPRC on sQTLs" from two papers may not be the same task if the QTL sets differ in phenotype definition, distance thresholds, or negative-set matching.
- Ensembles beat individuals for a structural reason. Averaging Borzoi and Pangolin ranks beat either alone (ΔAUPRC > 0.02) [5]; combining Borzoi with CADD beat either alone on gnomAD singleton-versus-common discrimination (AUROC 0.57 vs 0.55 each) [5]. When models have different output spaces, their errors are not the same errors.
The gaps this framing exposes¶
Open question: AlphaGenome predicts splice junctions explicitly — the closest any sequence model has come to LeafCutter's phenotype — but it does not classify junctions as productive or NMD-inducing. Since unproductive splicing is the branch that carries the expression effect [2], a model that predicts junction usage without productivity status predicts the splicing change but not its consequence for protein output. No cited study evaluates a sequence model against u-sQTLs.
Open question: all the expression benchmarks above are population-level. phASER supplies an individual-level, allele-resolved expression phenotype [3], and the AlphaGenome authors name personal-genome prediction as an unbenchmarked weakness of the whole model class [6]. Haplotypic expression is the obvious phenotype for that missing benchmark, and it has not been used for it.
Practical guidance¶
Lab interpretation: when designing a molecular-QTL analysis intended to interpret GWAS loci, pick the phenotype from the mechanism you want to be able to claim, not from convention.
- To ask whether splicing changes → intron excision ratios (LeafCutter).
- To ask whether splicing changes protein output → productive/unproductive classification (LeafCutter2); the negative control that u-sQTLs show no H3K9ac haQTL enrichment while showing strong eQTL/pQTL enrichment [2] is a useful template for demonstrating post-transcriptional action.
- To ask about this individual's alleles, or to resolve compound heterozygotes → haplotype phasing and haplotypic expression (phASER), noting that stop-gain variants are ~2.9× more likely to be mis-phased by population phasing than other variants [3].
- To interpret a sequence model's score → check which of these phenotypes it was benchmarked on before treating its output as evidence about yours.
See Also¶
-
Mechanism-Anchored Partitioned Polygenic Scores (MAP-PGS) — proposes unproductive-splicing-mediated risk as a separable polygenic axis.
-
Sequence-to-Function Genomic Models — the predictors evaluated against these phenotypes.
- Cryptic Splice Variants — the variant class where phenotype definition most changes the call.
- Direction-of-Effect Assignment at GWAS Loci — where allele-resolved and population-level expression evidence meet.
- Statistical Colocalization — consumes these phenotypes as the molecular trait.
Citations¶
[1] Li, Y. I., Knowles, D. A., Humphrey, J., Barbeira, A. N., Dickinson, S. P., Im, H. K. & Pritchard, J. K. (2018). Annotation-free quantification of RNA splicing using LeafCutter. Nature Genetics 50:151–158. Supports: intron-excision clustering and its ATSS/APA blind spot; the 5,774-vs-620 and 1.36–2.06× sQTL yield figures; 75–93% tissue replication; the 31.5% and 48.5% unannotated-intron figures. Location: Results ("Overview of LeafCutter"; "De novo identification of RNA splicing in mammalian organs"; "Mapping splicing QTLs using LeafCutter"); Table 1. Source copy: s41588-017-0004-9.md
[2] Buen Abad Najar, C. F. et al. (2025). Genetic and functional analysis of unproductive splicing using LeafCutter2. bioRxiv (preprint). Supports: the productive/unproductive junction classification; the u-sQTL-versus-p-sQTL expression correlation across 49 tissues; eQTL/pQTL enrichment and the haQTL null result. Location: Results ("Overview of LeafCutter2"; "Unproductive splicing mediates the effect of genetic variants…"; "Genetic basis of unproductive splicing…"); Figs. 4–5. Source paper: 2025.04.06.646893v1.full.pdf
[3] Castel, S. E., Mohammadi, P., Chung, W. K., Shen, Y. & Lappalainen, T. (2016). Rare variant phasing and haplotypic expression from RNA sequencing with phASER. Nature Communications 7:12817. Supports: haplotypic expression as a phenotype; the 2.9× stop-gain mis-phasing figure; the allelic-imbalance false-positive reduction. Location: Results ("Application of phASER to genetic studies"; "Application of phASER to allelic expression studies"). Source paper: ncomms12817.pdf
[4] Jaganathan, K. et al. (2019). Predicting Splicing from Primary Sequence with Deep Learning. Cell 176(3):535–548.e24. Supports: SpliceAI's output space as per-base donor/acceptor probability trained on GENCODE annotations. Location: Results ("Accurate Prediction of Splicing from a Primary Sequence Using Deep Learning"). Source copy: j.cell.2018.12.015.md
[5] Linder, J. et al. (2025). Predicting RNA-seq coverage from DNA sequence as a unifying model of gene regulation. Nature Genetics 57:949–961. Supports: coverage as the predicted object and the derived statistics; the eQTL Catalogue fine-mapped sQTL benchmark construction; the Pangolin-vs-Borzoi distance-dependent result and the ensemble gain; the coverage-bias confound; the Borzoi+CADD gnomAD result. Location: Results ("Functional splicing variant interpretation"; "Borzoi prioritizes genetic variants that influence expression"); Discussion; Figs. 5, 7. Source paper: s41588-024-02053-6.pdf
[6] Avsec, Ž. et al. (2026). Advancing regulatory variant effect prediction with AlphaGenome. Nature 649:1206–1217. Supports: splice junction prediction without productivity classification; personal-genome prediction named as an unbenchmarked class weakness. Location: "Improved splicing variant predictions"; Discussion (limitations). Source paper: s41586-025-10014-0.pdf
[7] Zeng, T. & Li, Y. I. (2022). Predicting RNA splicing from DNA sequence using Pangolin. Genome Biology 23:103. Supports: Pangolin's dual probability/usage output across four tissues, and its SpliSER-derived quantitative training targets. Location: Main text; Methods ("Deep neural network architecture"; "Generating training and test sets"). Source paper: s13059-022-02664-4.pdf