Omics field
Metagenomics
Metagenomics is the study of genetic material recovered directly from environmental or clinical samples, without culturing the organisms first. It describes which organisms are present in a community and, with deeper sequencing, what genes and pathways that community carries.
Questions this field can answer
- Which taxa are present in a sample, and at what relative abundance?
- How does community composition differ between sites, treatments or time points?
- What functional potential — enzymes, pathways, resistance genes — does the community carry?
- Can genomes of uncultured organisms be reconstructed from the sequence data?
- Is an observed difference plausibly biological, or an artefact of sampling and processing?
Method
Research workflow
Established practice, from framing a question to depositing data others can reuse.
- 01
Study design
Define the comparison, the unit of replication and the number of biological replicates before sampling. Randomise collection and processing order across groups.
- 02
Sampling
Standardise collection, transport and storage. Record location, time, temperature and matrix. Always include negative field and extraction controls plus a mock community if possible.
- 03
Assay and data generation
Choose amplicon (16S/18S/ITS) for cheap composition, or shotgun sequencing for functional potential and metagenome-assembled genomes. Fix primers, region, kit and depth for the whole study.
- 04
Quality control
Trim and filter reads, remove host and adapter sequence, and screen blanks for contaminant taxa. Report read depth per sample and drop samples below a stated threshold.
- 05
Analysis
Taxonomic profiling against a named reference release; functional profiling or assembly plus binning for MAGs. Treat counts as compositional — use rarefaction or ratio-based transforms, not raw proportions.
- 06
Interpretation
Composition is not activity. DNA persists after cell death, and relative abundance shifts when any one taxon changes. State the effect size and the uncertainty.
- 07
Reproducibility
Deposit raw reads in ENA or SRA with structured sample metadata, and record tool versions, database releases and parameters.
- 08
Deposition
Submit to MGnify for standardised public analysis, cite the study accession in the manuscript, and share analysis code.
Practice
Samples, technologies and what can go wrong
Method choice sets the ceiling on what an analysis can show. Limitations are part of the method, not an afterthought.
Sample types
- Soil, sediment and rhizosphere
- Freshwater, marine and wastewater
- Gut, stool and other host-associated sites
- Skin, oral and respiratory swabs
- Built-environment and air surfaces
- Low-biomass clinical samples, where contamination dominates most easily
Major technologies
- Amplicon sequencing (16S/18S/ITS)
- Cheap and robust for community composition; limited taxonomic resolution and no direct functional readout.
- Shotgun metagenomics
- Sequences all DNA present; supports functional profiling, strain resolution and assembly, but costs more and is sensitive to host DNA.
- Metagenome-assembled genomes (MAGs)
- Assembly plus binning reconstructs draft genomes of uncultured organisms; report completeness and contamination estimates.
- Long-read sequencing
- Improves assembly contiguity and repeat resolution at higher per-base error rates than short reads.
Limitations
- Sequence presence does not prove viability or activity.
- Reference databases are incomplete and biased toward well-studied environments.
- Amplicon primer choice systematically favours some taxa over others.
- MAGs are drafts — chimeric bins and missing genes are common.
- Low-biomass samples can be dominated by reagent and kit contaminants.
Common confounders
- Batch effects from extraction kit, sequencing run or operator.
- Host DNA swamping microbial reads in tissue and swab samples.
- Storage time and freeze–thaw cycles altering community profiles.
- Uneven sequencing depth mistaken for biological difference.
- Compositional artefacts: one taxon rising forces all others down.
Live data
Search the public record
Read-only searches against public databases. Madomic presents and explains the records; the databases named below remain their source and owner.
Madomic Research Explorer
Search public MGnify studies
Live, read-only search of public metagenomic studies. Results come directly from MGnify and are capped at eight records per search.
Searches are capped and rate limited out of respect for the public API. Nothing you type is stored.
Reference
Glossary
Essential terms, in plain English.
- Amplicon
- A targeted marker region, such as 16S rRNA, amplified by PCR before sequencing.
- OTU / ASV
- Operational taxonomic unit or amplicon sequence variant — the clustered or exact sequence used as a taxonomic unit.
- MAG
- Metagenome-assembled genome: a draft genome binned from assembled metagenomic contigs.
- Biome
- The standardised environment classification of a sample, e.g. root:Environmental:Aquatic:Marine.
- Alpha diversity
- Diversity within one sample: richness and evenness.
- Beta diversity
- Dissimilarity between samples, usually shown as an ordination.
- Compositional data
- Data constrained to a total, so abundances are relative and not independent.
- Negative control
- A blank processed alongside samples to reveal reagent and workflow contamination.
- Functional profiling
- Assigning reads to gene families or pathways to estimate functional potential.
For students
Learning path and a practical activity
Everything below uses public data only. No samples, credentials or paid services are needed.
Learning path
- 01Learn how DNA extraction and PCR bias shape everything downstream.
- 02Compare amplicon and shotgun designs, and what each can and cannot answer.
- 03Work through one public study end to end: metadata, depth, controls.
- 04Learn alpha and beta diversity, and read ordination plots critically.
- 05Understand compositional data and why raw proportions mislead.
- 06Read a MAG paper and check reported completeness and contamination.
- 07Practise writing a methods section detailed enough to reproduce.
Activity — compare public studies and form a testable hypothesis
- 01Use the explorer below to search MGnify for a biome that interests you, for example “soil” or “human gut”.
- 02Open two studies in MGnify and record accession, biome lineage, sample count and sequencing strategy.
- 03Follow each study to its ENA project and note the run count, platform and read length.
- 04List every metadata field present in one study but missing in the other, and say what that prevents you comparing.
- 05Write one hypothesis you could test with these public data, naming the comparison, the replication level and the confounders you would need to control.
- 06State how you would verify your answer in the originating databases and the associated publications.
Attribution
Research resources
Authoritative public resources. Omicser is independent of each of them and links to the official source.
MGnify
Maintained by EMBL-EBI
Standardised public analysis of metagenomic and metatranscriptomic studies, with biome classification, taxonomic and functional results, and an open API.
MGnify API v2
Maintained by EMBL-EBI
Programmatic, key-free access to MGnify studies, samples, analyses and downloads — the source used by the explorer on this page.
European Nucleotide Archive (ENA)
Maintained by EMBL-EBI
Archive of raw sequencing reads, assemblies and project metadata; the deposition target for most metagenomic studies.
Research and education use only. This page is not a clinical diagnosis, medical advice or a laboratory result. Verify every finding in the originating database and in the peer-reviewed literature before drawing conclusions.
