RNA-Seq: Ultrafast gene expression analysis. Now with ambient shipping for cells and purified RNA. Learn More
Back
Technical Documentation

RNA-Seq

https://plasmid-saurus.transforms.svdcdn.com/production/Resource-Center/Tech-Docs/PS-0006-E-RNA-Seq/RNA-Seq-Overview.png?w=4000&h=1096&auto=compress%2Cformat&fit=crop&dm=1789685300&s=8983af5d5b8afe3d8dc22ca72a6ce27b

How it works

Plasmidsaurus RNA-Seq uses short-read Illumina sequencing and a 3′ end counting workflow optimized for gene-level expression analysis. It is designed to produce accurate gene-level molecule counts for gene-level differential expression analysis. It is not intended for full-length transcript reconstruction or isoform analysis.

https://plasmid-saurus.transforms.svdcdn.com/production/Resource-Center/Tech-Docs/PS-0006-E-RNA-Seq/RNA-Seq-Sample-Submission.png?w=4000&h=1647&auto=compress%2Cformat&fit=crop&dm=1789685312&s=717e591cd504273abf78662be9cf73d4

Customers may submit either cells preserved in Zymo DNA/RNA Shield or purified total RNA on dry ice or in SEQguard Dino Preserve at room temperature. Refer to the Sample Preparation Guidelines for detailed input requirements.

https://plasmid-saurus.transforms.svdcdn.com/production/Resource-Center/Tech-Docs/PS-0006-E-RNA-Seq/RNA-Seq-prep-and-sequence.png?w=4000&h=1577&auto=compress%2Cformat&fit=crop&dm=1790091953&s=8788b4a80e621e66e8e1ee462988f581

For preserved cells, total RNA is isolated using a bead-based extraction method. RNA yield and sequencing performance depend on the number of cells submitted and the integrity of the starting material.

Following extraction, or upon receipt of purified RNA, RNA concentration is measured using a fluorescence-based plate assay. Samples are then normalized before library preparation. Plasmidsaurus does not measure RNA integrity number (RIN) before library preparation or hold samples when RNA concentration is below the recommended input level. Submitting an adequate number of intact cells or sufficient high-quality purified RNA is therefore important for obtaining optimal results.

Messenger RNA is reverse-transcribed into first-strand complementary DNA (cDNA) using a poly(dT)VN primer. During this step, each RNA molecule is tagged with a molecular barcode containing a unique molecular identifier (UMI), which enables sequencing reads derived from the same original molecule to be identified and collapsed during data processing. After second-strand synthesis, the double-stranded cDNA is tagmented and amplified to prepare it for Illumina sequencing. Unique dual indices (UDIs) are incorporated during amplification to support multiplexing and demultiplexing. 

Libraries are sequenced on the Illumina NovaSeq X Plus platform (catalog no. 20084804) using asymmetric paired-end sequencing. The mRNA reads are approximately 90 bases in length and are enriched near the 3′ end of captured transcripts, with most aligning within the final approximately 400 nucleotides. The library is stranded: each read matches the sense strand of the transcript it came from, so reads align to whichever genomic strand the gene is encoded on. Reads are forward-stranded for the purposes of strand-specific counting.

 


Bioinformatics

Each sample is processed through our primary RNA-Seq analysis pipeline before results are delivered. Sequencing reads are demultiplexed, quality filtered and trimmed, aligned to the selected reference genome, deduplicated using unique molecular identifiers (UMIs), and quantified at the gene level.

The resulting annotated gene-count matrix is used by interactive on-platform tools for secondary analysis. These tools enable you to evaluate relationships among samples using principal component analysis (PCA), perform differential gene expression analysis between user-defined sample groups, and assess differentially expressed genes for gene set enrichment.

RNA-Seq analysis workflow

Primary analysis

Step and softwareDescription
FASTQ generation and demultiplexing BCL Convert, fqtkConverts Illumina base call files to FASTQ format and assigns reads to samples using library indexes.
Quality filtering fastpTrims poly-X tails and low-quality bases from the 3′ end. Reads must meet a minimum Phred quality score of 15 and a minimum length of 50 bp after filtering.
Alignment STARAligns reads to the selected reference genome. Noncanonical splice junctions are removed, and unmapped reads are retained in the output.
BAM sorting samtoolsSorts aligned BAM files by genomic coordinate.
Deduplication UMICollapseCollapses reads sharing a unique molecular identifier at the same alignment position so that each original RNA molecule is counted once.
Mapping quality control RustQCCalculates alignment metrics, strand specificity, and read distribution across genomic features.
Aggregate quality-control reporting MultiQCConsolidates sample-level quality-control metrics into a single report.
Gene quantification featureCounts (Subread)Performs strand-specific counting with fractional assignment of multimapping reads. Reads overlapping exons and 3′ UTR features are assigned and grouped by gene_id. Count tables are annotated with gene biotype and other metadata from the reference GTF file.

Secondary analysis

Analysis and methodDescription
Normalization Trimmed mean of M-values (TMM)Normalizes gene counts to account for differences in library size and composition between samples.
Sample correlation Pearson correlationCalculates pairwise sample correlations from normalized counts, with hierarchical clustering, for the sample correlation heatmap.
Principal component analysis Uses normalized gene counts to evaluate the major sources of variation among samples.
Differential expression edgePythonReports gene-level expression differences between defined sample groups.
Functional enrichment GSEApyPerforms gene set enrichment analysis using the MSigDB Hallmark gene sets. Supported for human and mouse samples only.
Batch correction (optional) Method based on RUVSeqAnchor-based correction estimating technical variation between samples processed in separate batches. See the RNA-Seq anchor-based batch correction technical note.

 


Data and Deliverables

Filename key

ORDERCODE

Unique order identifier, such as JPY6QT.

OSID

Order-sample identifier, such as JPY6QT_1, used to trace files back to a specific tube.

SAMPLE

Customer-provided sample name, such as Control_rep_1. Used with the OSID in filenames like <OSID>_<SAMPLE>.

Download Results button <ORDERCODE>_results.zip

Results generated from automated data analysis pipeline.

File or deliverableDescription
<ORDERCODE>_mapping-stats/<OSID>.tsv <ORDERCODE>_mapping-stats/<OSID>_<SAMPLE>.tsvPer sample statistics on number of reads: mapped, unmapped, and deduplicated.
<ORDERCODE>_per-sample-multiqc-report/<OSID>.html <ORDERCODE>_per-sample-multiqc-report/<OSID>_<SAMPLE>.htmlPer sample statistics on the quality of reads sequenced and mapped.
<ORDERCODE>-correlation-heatmap.pngPairwise Pearson correlation coefficients with hierarchical clustering. Also exportable as .png or .svg online.
<ORDERCODE>-expression-matrix.tsvGene count matrix for all the samples in the order, including both total counts and CPM (counts per million).
<ORDERCODE>-gene-biotype-5plus_reads-summary.csvBarplot of genes detected by biotype, only including genes with at least 5 reads.
<ORDERCODE>-gene-biotype-5plus_reads.htmlNumber of genes detected by biotype, only including genes with at least 5 reads.
<ORDERCODE>-mapping-stats-reads.csvOrder-level statistics on number of reads: mapped, unmapped, and deduplicated.
<ORDERCODE>-methods.txtDescription of the bioinformatics methods used to filter, align, and process the raw reads.
<ORDERCODE>-multiqc-report.htmlOrder-level statistics on the quality of reads sequenced and mapped.
<ORDERCODE>-pca-scatter.htmlPrincipal Component Analysis uses normalized gene counts to evaluate major sources of variation among samples. Also exportable as .png or .svg online.

FASTQ button <ORDERCODE>_fastq.zip

File or deliverableDescription
<OSID>_<SAMPLE>.fastq.gzAdapter-trimmed and length-filtered reads, demultiplexed and forward-stranded. Poly(A) sequences are retained, reads are not UMI-deduplicated, and each UMI is carried in the read header.

BAM button <ORDERCODE>.bam.zip

File or deliverableDescription
<OSID>_<SAMPLE>_dedup-mapped-reads.bam.csi <OSID>_<SAMPLE>_dedup-mapped-reads.bam <OSID>_<SAMPLE>_unmapped-reads.bamUMI-deduplicated reads aligned to reference in bam.zip folder.

Online Results Page

Primary analysis

File or deliverableDescription
Interactive gene expression profile plots gene-expression-<ORDERCODE>.png gene-expression-<ORDERCODE>.svg gene-expression-raw_<ORDERCODE>.csvHeatmap or bar plots of gene expression compared across samples. Export figure as .png, .svg. Export all, or selected data as .csv.

Secondary analysis

File or deliverableDescription
Interactive volcano or MA plots dge-volcano-plot-<ORDERCODE>.png dge-volcano-plot-<ORDERCODE>.svg dge-ma-plot-<ORDERCODE>.png dge-ma-plot-<ORDERCODE>.svgPlot showing up- and down-regulated genes between two conditions. Export figure as .png or .svg.
Differentially Expressed Genes table <ORDERCODE>_<SAMPLE>_dge.csvDEG data used to create the volcano or MA plots. Option to select all genes, currently selected genes, or genes that match filtered results.
Interactive Functional Enrichment heatmap and dotplot functional_enrichment.pngHeatmap or dotplot showing pathways identified with gene set enrichment analysis (GSEA) using the Hallmark gene set.
Enriched Pathways table <Pathway Library>_gsea_functional_enrichment.csvResults of functional enrichment analysis, showing enriched pathways and their statistical significance.

Technical notes and related guides

 


Performance and troubleshooting

Samples that meet input requirements typically yield 10 million or more unique, deduplicated reads. Read count is not guaranteed and varies with sample input and RNA quality.

Deduplication is performed after alignment, so only reads that map to the selected reference genome are collapsed and counted. The reported unique, deduplicated read count therefore reflects mapped reads rather than total sequencing output. For samples with a low mapping rate, this count can be substantially lower than the number of reads generated, and a low value alone does not distinguish insufficient sequencing depth from poor alignment. Per-sample mapping statistics reporting mapped, unmapped, and deduplicated read counts are provided with your results.

 


Reruns

If a processing deviation occurs, Plasmidsaurus may rerun the sample when sufficient material is available. Submitting additional volume can help support repeat processing.

For questions about sequencing results or rerun eligibility, contact support@plasmidsaurus.com.

 


Version Information
Document ID: PS-0006-E 
Version: 1.1
Revision date: 9/12/2026