Shotgun Metagenomics

How it works
Plasmidsaurus Shotgun Metagenomics Sequencing uses short-read Illumina sequencing to comprehensively profile all genomic DNA present in a sample, rather than targeting a single marker gene. This captures bacteria, archaea, fungi, protists, and viruses in a single assay and supports both taxonomic and functional profiling of the community.
Samples are submitted as extracted metagenomic DNA. Refer to the Sample Preparation Documentation for detailed input requirements. Libraries are built using transposase-mediated library preparation and sequenced with paired-end chemistry (2 × 150) on an Illumina NovaSeq X Plus (catalog no. 20084804).

Bioinformatics
Raw sequencing reads are processed through an automated data analysis pipeline before delivery. The pipeline performs quality filtering, host read removal, and taxonomic and functional profiling.

| Step and software | Description |
|---|---|
| Quality control fastp | Raw paired-end reads are filtered at a minimum Phred quality score of 15 and a minimum read length of 50 bp. Adapters are detected automatically, and poly(G) tails are trimmed to remove NovaSeq two-color chemistry artifacts. Samples must yield at least 1 million reads after QC to proceed to taxonomic analysis. |
| Host removal minimap2 | Reads are aligned against human and mouse reference genomes. Only unmapped (non-host) read pairs are retained for downstream analysis. |
| Taxonomic profiling Sylph; sylph-tax | Non-host reads are profiled using abundance-corrected MinHash against curated reference databases covering bacteria, archaea (GTDB r226), fungi (RefSeq), protists, and eukaryotic DNA viruses. Taxonomy is assigned via sylph-tax. |
| Functional profiling HUMAnN; Bowtie 2; DIAMOND | Non-host reads are screened against species-specific ChocoPhlAn nucleotide pangenomes using Bowtie 2; reads that do not map are aligned against the UniRef90 protein database using DIAMOND translated search. Gene family abundances are regrouped to EggNOG v5 orthologous groups with human-readable names and aggregated into MetaCyc pathways, reported as both relative abundance and pathway coverage scores. Functional tables are delivered as community totals (unstratified). |
| Differential analysis On platform | On your results page, select up to two groups of samples to generate differential abundance comparisons of taxonomic and functional profiles. We use centered log-ratio transformation for compositional abundances, feature-level statistical testing with multiple-testing correction, and PERMANOVA on Aitchison distances for community-level comparisons. |
Data and Deliverables
Filename key
ORDERCODEUnique order identifier, such as
JPY6QT.OSIDOrder-sample identifier, such as
JPY6QT_1, used to trace files back to a specific tube.SAMPLECustomer-provided sample name, such as
Control_rep_1. Used with the OSID in filenames like<OSID>_<SAMPLE>.
Download Results button <ORDERCODE>_results.zip
Results generated from automated data analysis pipeline.
| File or deliverable | Description |
|---|---|
<ORDERCODE>_qc-short-reads/<OSID>_<SAMPLE>.html <ORDERCODE>_qc-short-reads/<OSID>_<SAMPLE>.json | Read quality report for the sample, before host removal. |
<ORDERCODE>_taxonomy-sylph/<OSID>_<SAMPLE>.tsv | Species and genus-level taxonomy calls for the sample. |
<ORDERCODE>_host-removal/<OSID>_<SAMPLE>.flagstat | Read counts and alignment stats from filtering out host (human/mouse) reads. |
<ORDERCODE>_pathabundance-relab-unstratified/<OSID>_<SAMPLE>.tsv | Relative abundance of metabolic pathways detected in the sample. |
<ORDERCODE>_pathcoverage-unstratified/<OSID>_<SAMPLE>.tsv | Which steps of each metabolic pathway were covered by the sample's reads. |
<ORDERCODE>_genefamilies-eggnog-named-unstratified/<OSID>_<SAMPLE>.tsv | Gene family abundance for the sample, annotated with eggNOG functional categories. |
<ORDERCODE>_taxonomy_sylph_genus_relative_abundance.tsv | Genus-level abundance table across every sample in the order. |
<ORDERCODE>_taxonomy_sylph_species_relative_abundance.tsv | Species-level abundance table across every sample in the order. |
<ORDERCODE>_pathways_humann_relative_abundance.tsv | Pathway abundance compared across every sample in the order. |
<ORDERCODE>_pathways_humann_pathway_completeness.tsv | Pathway completeness compared across every sample in the order. |
FASTQ button <ORDERCODE>_fastq.zip
| File or deliverable | Description |
|---|---|
<OSID>_<SAMPLE>_R1.fastq.gz <OSID>_<SAMPLE>_R2.fastq.gz | Quality-filtered paired-end reads for each sample. |
<OSID>_<SAMPLE>_hostremoved_R1.fastq.gz <OSID>_<SAMPLE>_hostremoved_R2.fastq.gz | Host-removed paired-end reads for each sample. |
Online Results Page
| Deliverable | Description |
|---|---|
Cross-sample Taxonomic Summary (order-level) taxonomy-summary-<ORDERCODE>.png | Stacked bar/area chart comparing taxonomic composition across every sample in the order, switchable between phylum, family, genus, and species. |
Per-sample Taxonomic summary (Sankey + Table) metagenomics-taxonomy-sankey-<OSID>-<SAMPLE>.png metagenomics-taxonomy-sankey-<OSID>-<SAMPLE>.csv | Interactive breakdown of the taxonomy detected in one sample, from domain down to species. |
Per-sample Functional profile (Top 50 Pathways) metagenomics-functional-profile-<OSID>-<SAMPLE>.png metagenomics-pathway-abundance-<OSID>-<SAMPLE>.csv | Bar chart of the most abundant metabolic pathways detected in one sample. |
| Observed Species / Shannon Diversity (per sample) | Two diversity metrics shown alongside each sample's read stats. |
Pathway Heatmap (cross-sample) top-pathways-heatmap-<ORDERCODE>.png | Functional pathway abundance compared across all samples in the order. |
Compare: group differential analysis + PCoA plots taxonomy-pcoa--genus-.png top-differential-taxa--genus-.png pathways-pcoa.png top-differential-pathways.png | Statistical comparison of taxa or pathways between two groups of samples you define, plus PCoA plots for taxonomy and pathways. |
Technical notes and related guides
Performance and troubleshooting
Shotgun Metagenomics sequencing status is determined by raw data yield. A sample is considered unsuccessful when it does not produce the minimum raw-data target for the selected service.
Taxonomic results are not used to determine completion status. Metagenomic samples vary widely in microbial composition, total biomass, host or background DNA content, and the relative abundance of each organism in the community. As a result, taxonomic profiles and classification results may vary between sample types, even when the sequencing run meets the raw-data target.
Although the exact cause of an unsuccessful or low-yield result cannot always be determined, the most common causes are sample impurities that interfere with library preparation or sequencing, and DNA submitted below the required concentration.
To support optimal sequencing results, follow the recommended sample preparation requirements for input amount, concentration, purity, and buffer conditions.
Reruns
Shotgun Metagenomics outcomes depend on the quantity, quality, and purity of the submitted sample. For this reason, Plasmidsaurus does not guarantee that every sample will meet the raw-data target for the selected service level.
When a sample provides adequate DNA concentration and quality but does not reach its raw-data target, we review the initial sequencing results to determine whether additional sequencing is likely to improve the outcome. Eligible samples are queued for one complimentary rerun, and data from the initial run and rerun are combined to increase the likelihood of a successful result.
If the raw-data target is not achieved after the complimentary rerun, no further complimentary reruns will be performed. Charges still apply to unsuccessful samples because library preparation, sequencing, analysis, and additional review have already been completed.
To attempt sequencing again, submit a new sample that meets all applicable sample preparation and quality-control requirements and place a new sequencing order.
For questions about sequencing results or rerun eligibility, contact support@plasmidsaurus.com.
Version Information
Document ID: PS-0005-E
Version: 1.1
Revision date: 9/22/2026