Omics field

Metagenomics

Metagenomics is the study of genetic material recovered directly from environmental or clinical samples, without culturing the organisms first. It describes which organisms are present in a community and, with deeper sequencing, what genes and pathways that community carries.

Questions this field can answer

  • Which taxa are present in a sample, and at what relative abundance?
  • How does community composition differ between sites, treatments or time points?
  • What functional potential — enzymes, pathways, resistance genes — does the community carry?
  • Can genomes of uncultured organisms be reconstructed from the sequence data?
  • Is an observed difference plausibly biological, or an artefact of sampling and processing?

Method

Research workflow

Established practice, from framing a question to depositing data others can reuse.

  1. 01

    Study design

    Define the comparison, the unit of replication and the number of biological replicates before sampling. Randomise collection and processing order across groups.

  2. 02

    Sampling

    Standardise collection, transport and storage. Record location, time, temperature and matrix. Always include negative field and extraction controls plus a mock community if possible.

  3. 03

    Assay and data generation

    Choose amplicon (16S/18S/ITS) for cheap composition, or shotgun sequencing for functional potential and metagenome-assembled genomes. Fix primers, region, kit and depth for the whole study.

  4. 04

    Quality control

    Trim and filter reads, remove host and adapter sequence, and screen blanks for contaminant taxa. Report read depth per sample and drop samples below a stated threshold.

  5. 05

    Analysis

    Taxonomic profiling against a named reference release; functional profiling or assembly plus binning for MAGs. Treat counts as compositional — use rarefaction or ratio-based transforms, not raw proportions.

  6. 06

    Interpretation

    Composition is not activity. DNA persists after cell death, and relative abundance shifts when any one taxon changes. State the effect size and the uncertainty.

  7. 07

    Reproducibility

    Deposit raw reads in ENA or SRA with structured sample metadata, and record tool versions, database releases and parameters.

  8. 08

    Deposition

    Submit to MGnify for standardised public analysis, cite the study accession in the manuscript, and share analysis code.

Practice

Samples, technologies and what can go wrong

Method choice sets the ceiling on what an analysis can show. Limitations are part of the method, not an afterthought.

Sample types

  • Soil, sediment and rhizosphere
  • Freshwater, marine and wastewater
  • Gut, stool and other host-associated sites
  • Skin, oral and respiratory swabs
  • Built-environment and air surfaces
  • Low-biomass clinical samples, where contamination dominates most easily

Major technologies

Amplicon sequencing (16S/18S/ITS)
Cheap and robust for community composition; limited taxonomic resolution and no direct functional readout.
Shotgun metagenomics
Sequences all DNA present; supports functional profiling, strain resolution and assembly, but costs more and is sensitive to host DNA.
Metagenome-assembled genomes (MAGs)
Assembly plus binning reconstructs draft genomes of uncultured organisms; report completeness and contamination estimates.
Long-read sequencing
Improves assembly contiguity and repeat resolution at higher per-base error rates than short reads.

Limitations

  • Sequence presence does not prove viability or activity.
  • Reference databases are incomplete and biased toward well-studied environments.
  • Amplicon primer choice systematically favours some taxa over others.
  • MAGs are drafts — chimeric bins and missing genes are common.
  • Low-biomass samples can be dominated by reagent and kit contaminants.

Common confounders

  • Batch effects from extraction kit, sequencing run or operator.
  • Host DNA swamping microbial reads in tissue and swab samples.
  • Storage time and freeze–thaw cycles altering community profiles.
  • Uneven sequencing depth mistaken for biological difference.
  • Compositional artefacts: one taxon rising forces all others down.

Live data

Search the public record

Read-only searches against public databases. Madomic presents and explains the records; the databases named below remain their source and owner.

Madomic Research Explorer

Search public MGnify studies

Live, read-only search of public metagenomic studies. Results come directly from MGnify and are capped at eight records per search.

Searches are capped and rate limited out of respect for the public API. Nothing you type is stored.

Reference

Glossary

Essential terms, in plain English.

Amplicon
A targeted marker region, such as 16S rRNA, amplified by PCR before sequencing.
OTU / ASV
Operational taxonomic unit or amplicon sequence variant — the clustered or exact sequence used as a taxonomic unit.
MAG
Metagenome-assembled genome: a draft genome binned from assembled metagenomic contigs.
Biome
The standardised environment classification of a sample, e.g. root:Environmental:Aquatic:Marine.
Alpha diversity
Diversity within one sample: richness and evenness.
Beta diversity
Dissimilarity between samples, usually shown as an ordination.
Compositional data
Data constrained to a total, so abundances are relative and not independent.
Negative control
A blank processed alongside samples to reveal reagent and workflow contamination.
Functional profiling
Assigning reads to gene families or pathways to estimate functional potential.

For students

Learning path and a practical activity

Everything below uses public data only. No samples, credentials or paid services are needed.

Learning path

  1. 01Learn how DNA extraction and PCR bias shape everything downstream.
  2. 02Compare amplicon and shotgun designs, and what each can and cannot answer.
  3. 03Work through one public study end to end: metadata, depth, controls.
  4. 04Learn alpha and beta diversity, and read ordination plots critically.
  5. 05Understand compositional data and why raw proportions mislead.
  6. 06Read a MAG paper and check reported completeness and contamination.
  7. 07Practise writing a methods section detailed enough to reproduce.

Activity — compare public studies and form a testable hypothesis

  1. 01Use the explorer below to search MGnify for a biome that interests you, for example “soil” or “human gut”.
  2. 02Open two studies in MGnify and record accession, biome lineage, sample count and sequencing strategy.
  3. 03Follow each study to its ENA project and note the run count, platform and read length.
  4. 04List every metadata field present in one study but missing in the other, and say what that prevents you comparing.
  5. 05Write one hypothesis you could test with these public data, naming the comparison, the replication level and the confounders you would need to control.
  6. 06State how you would verify your answer in the originating databases and the associated publications.

Attribution

Research resources

Authoritative public resources. Omicser is independent of each of them and links to the official source.