Detectability planning for top/middle-down proteomics

MS strategy

Top-down analyzes the intact proteoform. Middle-down simulates a limited (partial) protease digestion first, then runs the same MS1/MS2 analysis on the resulting large peptides instead of the intact protein.

Default 10-220 kDa covers the 2.5th-97.5th percentile (~95%) of the reviewed human proteome's intact monoisotopic mass, excluding both small fragments and the long tail of very large proteins (titin, dystrophin, etc.) that are impractical top-down MS targets. Isoforms outside this range are hidden from selection below, and the confounding-protein search pool is limited to it too -- adjust freely for your instrument's real usable mass range.

Lys-C, Lys-N, and Glu-C are also used in bottom-up proteomics; under a limited (short-time) middle-down digestion they leave missed-cleavage sites, which this mass window is simulating rather than modeling digestion kinetics directly.

MS resolution parameters

These feed every downstream scoring step (resolving power, envelope crowding, confounder search) regardless of which input path below you use.

Input selection


Gene / isoform / proteoform selection



Isoform catalog (real Ensembl transcripts for this gene)

Result proteoform table

Every checked isoform (plus its parsed PTM combinations) is included in the comparison below. Pick one as the confounder-search target:


Middle-down peptide candidates

1. Relevant-proteoform comparison
(run analysis to populate)

2. Confounding-protein search (single target)
(run analysis to populate)

FASTA sequence input

Provide a single spliced mRNA/cDNA (or CDS) nucleotide sequence -- not raw genomic DNA with introns. Two independent steps run: minimap2 spliced-aligns it against GRCh38 to identify which known gene/isoforms it belongs to (this does not depend on translation at all), and TransDecoder separately finds candidate open reading frames. Pick a known isoform to compare its exon structure against, pick which ORF candidate to use as the translation, add PTMs, then send it into the same comparison view as Option 1.



rMATS alternative-splicing results

rMATS only reports the differential exon(s) and their immediate flanking exons, not the rest of the transcript -- so a full-length proteoform can't be computed from the event alone. Instead, ProteoformTracker looks up which already-annotated transcripts of the gene (in the same precomputed exon index Option 1 uses) structurally match each arm of the event (exon-inclusion vs. exon-skipping for SE; 1st-exon vs. 2nd-exon for MXE; intron-retained vs. -spliced for RI; long- vs. short-exon form for A5SS/A3SS), so you get real, full-length proteoforms rather than just the local differential region. Every arm also gets a "constructed" synthetic isoform (a user-pickable backbone transcript with the local region replaced by rMATS' own reported exons), so an arm with no real annotated match still has something usable.