RNA-Seq: Ultrafast gene expression analysis. Now with ambient shipping for cells and purified RNA. Learn More
Back
Technical Documentation

Shotgun Metagenomics

https://plasmid-saurus.transforms.svdcdn.com/production/Resource-Center/Tech-Docs/PS-0005-E-Shotgun-Metagenomics/Shotgun-Metagenomics-Overview.png?w=4000&h=1084&auto=compress%2Cformat&fit=crop&dm=1790094496&s=2d5950604e10c5b2a4b63275a27635c8

How it works 

Plasmidsaurus Shotgun Metagenomics Sequencing uses short-read Illumina sequencing to comprehensively profile all genomic DNA present in a sample, rather than targeting a single marker gene. This captures bacteria, archaea, fungi, protists, and viruses in a single assay and supports both taxonomic and functional profiling of the community.

Samples are submitted as extracted metagenomic DNA. Refer to the Sample Preparation Documentation for detailed input requirements. Libraries are built using transposase-mediated library preparation and sequenced with paired-end chemistry (2 × 150) on an Illumina NovaSeq X Plus (catalog no. 20084804).

https://plasmid-saurus.transforms.svdcdn.com/production/Resource-Center/Tech-Docs/PS-0005-E-Shotgun-Metagenomics/Shotgun-Metagenomics-library-prep-and-sequencing.png?w=4000&h=1087&auto=compress%2Cformat&fit=crop&dm=1790094491&s=13c9cba8e7806e3eb4ca9556e588c455

Bioinformatics

Raw sequencing reads are processed through an automated data analysis pipeline before delivery. The pipeline performs quality filtering, host read removal, and taxonomic and functional profiling. 

https://plasmid-saurus.transforms.svdcdn.com/production/Resource-Center/Tech-Docs/PS-0005-E-Shotgun-Metagenomics/Shotgun-Metagenomics-data-analysis-2.png?w=4000&h=2274&auto=compress%2Cformat&fit=crop&dm=1790094476&s=436617a4b45f508abfd4ded37fbed9a4
Step and softwareDescription
Quality control fastpRaw paired-end reads are filtered at a minimum Phred quality score of 15 and a minimum read length of 50 bp. Adapters are detected automatically, and poly(G) tails are trimmed to remove NovaSeq two-color chemistry artifacts. Samples must yield at least 1 million reads after QC to proceed to taxonomic analysis.
Host removal minimap2Reads are aligned against human and mouse reference genomes. Only unmapped (non-host) read pairs are retained for downstream analysis.
Taxonomic profiling Sylph; sylph-taxNon-host reads are profiled using abundance-corrected MinHash against curated reference databases covering bacteria, archaea (GTDB r226), fungi (RefSeq), protists, and eukaryotic DNA viruses. Taxonomy is assigned via sylph-tax.
Functional profiling HUMAnN; Bowtie 2; DIAMONDNon-host reads are screened against species-specific ChocoPhlAn nucleotide pangenomes using Bowtie 2; reads that do not map are aligned against the UniRef90 protein database using DIAMOND translated search. Gene family abundances are regrouped to EggNOG v5 orthologous groups with human-readable names and aggregated into MetaCyc pathways, reported as both relative abundance and pathway coverage scores. Functional tables are delivered as community totals (unstratified).
Differential analysis On platformOn your results page, select up to two groups of samples to generate differential abundance comparisons of taxonomic and functional profiles. We use centered log-ratio transformation for compositional abundances, feature-level statistical testing with multiple-testing correction, and PERMANOVA on Aitchison distances for community-level comparisons.

Data and Deliverables

Filename key

ORDERCODE

Unique order identifier, such as JPY6QT.

OSID

Order-sample identifier, such as JPY6QT_1, used to trace files back to a specific tube.

SAMPLE

Customer-provided sample name, such as Control_rep_1. Used with the OSID in filenames like <OSID>_<SAMPLE>.

Download Results button <ORDERCODE>_results.zip

Results generated from automated data analysis pipeline.

File or deliverableDescription
<ORDERCODE>_qc-short-reads/<OSID>_<SAMPLE>.html <ORDERCODE>_qc-short-reads/<OSID>_<SAMPLE>.jsonRead quality report for the sample, before host removal.
<ORDERCODE>_taxonomy-sylph/<OSID>_<SAMPLE>.tsvSpecies and genus-level taxonomy calls for the sample.
<ORDERCODE>_host-removal/<OSID>_<SAMPLE>.flagstatRead counts and alignment stats from filtering out host (human/mouse) reads.
<ORDERCODE>_pathabundance-relab-unstratified/<OSID>_<SAMPLE>.tsvRelative abundance of metabolic pathways detected in the sample.
<ORDERCODE>_pathcoverage-unstratified/<OSID>_<SAMPLE>.tsvWhich steps of each metabolic pathway were covered by the sample's reads.
<ORDERCODE>_genefamilies-eggnog-named-unstratified/<OSID>_<SAMPLE>.tsvGene family abundance for the sample, annotated with eggNOG functional categories.
<ORDERCODE>_taxonomy_sylph_genus_relative_abundance.tsvGenus-level abundance table across every sample in the order.
<ORDERCODE>_taxonomy_sylph_species_relative_abundance.tsvSpecies-level abundance table across every sample in the order.
<ORDERCODE>_pathways_humann_relative_abundance.tsvPathway abundance compared across every sample in the order.
<ORDERCODE>_pathways_humann_pathway_completeness.tsvPathway completeness compared across every sample in the order.

FASTQ button <ORDERCODE>_fastq.zip

File or deliverableDescription
<OSID>_<SAMPLE>_R1.fastq.gz <OSID>_<SAMPLE>_R2.fastq.gzQuality-filtered paired-end reads for each sample.
<OSID>_<SAMPLE>_hostremoved_R1.fastq.gz <OSID>_<SAMPLE>_hostremoved_R2.fastq.gzHost-removed paired-end reads for each sample.

Online Results Page

DeliverableDescription
Cross-sample Taxonomic Summary (order-level) taxonomy-summary-<ORDERCODE>.pngStacked bar/area chart comparing taxonomic composition across every sample in the order, switchable between phylum, family, genus, and species.
Per-sample Taxonomic summary (Sankey + Table) metagenomics-taxonomy-sankey-<OSID>-<SAMPLE>.png metagenomics-taxonomy-sankey-<OSID>-<SAMPLE>.csvInteractive breakdown of the taxonomy detected in one sample, from domain down to species.
Per-sample Functional profile (Top 50 Pathways) metagenomics-functional-profile-<OSID>-<SAMPLE>.png metagenomics-pathway-abundance-<OSID>-<SAMPLE>.csvBar chart of the most abundant metabolic pathways detected in one sample.
Observed Species / Shannon Diversity (per sample)Two diversity metrics shown alongside each sample's read stats.
Pathway Heatmap (cross-sample) top-pathways-heatmap-<ORDERCODE>.pngFunctional pathway abundance compared across all samples in the order.
Compare: group differential analysis + PCoA plots taxonomy-pcoa--genus-.png top-differential-taxa--genus-.png pathways-pcoa.png top-differential-pathways.pngStatistical comparison of taxa or pathways between two groups of samples you define, plus PCoA plots for taxonomy and pathways.

Performance and troubleshooting 

Shotgun Metagenomics sequencing status is determined by raw data yield. A sample is considered unsuccessful when it does not produce the minimum raw-data target for the selected service.

Taxonomic results are not used to determine completion status. Metagenomic samples vary widely in microbial composition, total biomass, host or background DNA content, and the relative abundance of each organism in the community. As a result, taxonomic profiles and classification results may vary between sample types, even when the sequencing run meets the raw-data target.

Although the exact cause of an unsuccessful or low-yield result cannot always be determined, the most common causes are sample impurities that interfere with library preparation or sequencing, and DNA submitted below the required concentration.

Issue

Possible cause

Recommended action

Raw-data yield below the minimum target

Sequencing inhibitors or contaminants remaining after extraction

Purify with a spin column (Plasmidsaurus recommends the Zymo OneStep PCR Inhibitor Removal Kit for metagenomic gDNA), AMPure XP beads, or an equivalent method. Elute in 10 mM Tris (pH 8.5) or nuclease-free water.

DNA below the required 10 ng/µL, often from NanoDrop-based quantification

Quantify with a Qubit or equivalent fluorescence-based assay and submit at the required concentration.

To support optimal sequencing results, follow the recommended sample preparation requirements for input amount, concentration, purity, and buffer conditions.

 


Reruns

Shotgun Metagenomics outcomes depend on the quantity, quality, and purity of the submitted sample. For this reason, Plasmidsaurus does not guarantee that every sample will meet the raw-data target for the selected service level.

When a sample provides adequate DNA concentration and quality but does not reach its raw-data target, we review the initial sequencing results to determine whether additional sequencing is likely to improve the outcome. Eligible samples are queued for one complimentary rerun, and data from the initial run and rerun are combined to increase the likelihood of a successful result.

If the raw-data target is not achieved after the complimentary rerun, no further complimentary reruns will be performed. Charges still apply to unsuccessful samples because library preparation, sequencing, analysis, and additional review have already been completed.

To attempt sequencing again, submit a new sample that meets all applicable sample preparation and quality-control requirements and place a new sequencing order.

For questions about sequencing results or rerun eligibility, contact support@plasmidsaurus.com.

 


Version Information
Document ID: PS-0005-E 
Version: 1.1
Revision date: 9/22/2026