Human genome

The reference, the history and the evidence behind human genome research

A human genome is both a biological sequence and a coordinate system for research. This guide separates established facts, changing reference assemblies and interpretation limits.

Genome at a glance

Scale, organization and function

The genome is not a list of genes. It is a layered system of coding sequence, regulatory elements, repeats, structural regions and mitochondrial DNA.

~3.1 billion

base pairs

A haploid nuclear reference is about three billion DNA letters long; exact totals depend on the assembly.

23 pairs

chromosomes

Humans normally carry 22 pairs of autosomes and one pair of sex chromosomes, plus a small mitochondrial genome.

<2%

protein-coding sequence

Most sequence does not encode protein. Regulatory, structural, repetitive and still-uncertain regions are biologically important.

99.9%+

shared sequence

Human genomes are overwhelmingly similar, but millions of inherited and new variants make each genome distinct.

Sequence layers

Genes and transcripts sit within regulatory DNA, repetitive elements, centromeres, telomeres and large structural segments.

A reference is a model

A reference assembly is not a universal or ideal human genome. It is a versioned scaffold used to name coordinates and compare observations.

Variation needs context

Single-nucleotide changes, indels, copy-number and structural variants require population, family, tissue and phenotype evidence.

History

From coordinated mapping to an open reference

The Human Genome Project was an international public consortium, not the work of one person or one laboratory.

  1. 1988–1990

    The U.S. National Research Council endorsed a coordinated program; NIH and DOE launched the Human Genome Project in 1990.

  2. 1996

    The Bermuda Principles established rapid public release of large-scale sequence data, shaping open genomic science.

  3. 2000–2001

    The international public consortium announced a working draft; its analysis appeared in Nature in 2001.

  4. 2003

    The project announced a finished essential sequence, covering about 99% of gene-containing regions at 99.99% accuracy.

  5. 2013–2022

    GRCh38 became the major coordinate reference; the Telomere-to-Telomere consortium later completed previously unresolved regions with T2T-CHM13.

  6. 2023 onward

    The Human Pangenome Reference Consortium began representing global diversity with many high-quality, phased assemblies rather than one linear genome.

People and collaboration

Scientists, centres and public infrastructure

Thousands of researchers at 20 sequencing centres across six countries contributed. The names below help orient the history; they do not replace the consortium credit.

James D. Watson

First director of the U.S. National Center for Human Genome Research, 1989–1992.

Francis S. Collins

Led the U.S. public genome program from 1993 through completion of the Human Genome Project.

John Sulston

Led the Wellcome Trust Sanger Centre contribution and advocated rapid, public sequence release.

Robert Waterston & Eric Lander

Major leaders of public sequencing centres and consortium analysis.

Ari Patrinos

Directed the U.S. Department of Energy genome program during key project years.

Craig Venter & Celera

Led a separate private-sector draft effort; important historically, but distinct from the international public consortium.

Detailed analysis

A reproducible human-variant research workflow

A responsible workflow keeps identity, evidence, uncertainty and provenance visible from beginning to end.

1. Frame

Define phenotype, population, tissue, inheritance model and the question before opening a database.

2. Identify

Record genome assembly, chromosome, position, reference and alternate alleles. Normalize the variant.

3. Annotate

Map genes, transcripts, consequence, frequency, conservation and regulatory context using versioned sources.

4. Compare

Check population frequencies, phenotype databases, functional evidence and primary literature. Avoid treating prediction as proof.

5. Validate

Confirm technical quality and, where required, use an independent validated laboratory method.

6. Report

Preserve provenance, access dates, software/database versions, uncertainty and limitations. Protect participant privacy.

Interpretation guardrails

A computational score is a hypothesis-support tool, not a diagnosis. Absence from a database is not evidence of harmlessness; correlation is not causation; ancestry imbalance can distort frequency estimates; and reference bias can hide structural diversity. Clinical classification requires accredited workflows, current professional standards and qualified specialists.

Ethics and ELSI

Genomic data describes people, families and populations

Consent, privacy, governance and equitable representation are scientific quality requirements.

Responsible practice

  • Use purpose-specific informed consent and approved data access.
  • Minimize identifiable data; genomic sequence cannot be made truly anonymous.
  • Plan for familial findings, recontact, withdrawal and controlled sharing.
  • Report population descriptors precisely and avoid biological claims based on social categories.

What this page does not do

It does not interpret a person’s genome, establish pathogenicity, recommend care or replace genetic counselling. It is a research and education map to primary sources.