Human mitochondrial DNA (mtDNA) is a compact genome with unusual analytical properties. It is abundant, predominantly maternally inherited, and supported by a detailed population phylogeny. Those features make it useful in population genetics, ancient-DNA research, forensics, genealogy, and studies of mitochondrial biology. They also make it easy to overinterpret: an mtDNA sequence is not a miniature nuclear genome, a haplogroup is not a complete ancestry profile, and a consensus sequence does not describe every mitochondrial molecule in a specimen.
The practical question is therefore not only “what sequence did we obtain?” but also “which molecules, tissue, assay, reference, and decision rules does that sequence represent?”
A 16,569-base genome with dense organization
The human mitochondrial reference genome is 16,569 base pairs long and conventionally represented as a circle. It contains 37 genes: 13 protein-coding genes that contribute subunits of oxidative-phosphorylation complexes, 22 transfer RNA genes, and two ribosomal RNA genes. Most bases are transcribed, genes are tightly packed, and some genes overlap. A noncoding control region contains important elements for replication and transcription.
The two strands are called heavy (H) and light (L), reflecting their nucleotide composition. Most genes are encoded on the heavy strand. Mitochondrial translation also differs in several codon assignments from the standard nuclear genetic code, so nuclear-genome annotation assumptions cannot simply be reused.
Although the molecule is circular, sequence files and coordinate systems must choose a breakpoint. Human mtDNA convention places that breakpoint between positions 16,569 and 1. Regions crossing it, including the control region, may therefore be represented as two coordinate intervals. Circularity also matters for alignment: reads spanning the artificial breakpoint can be lost or soft-clipped if a mapper treats the reference as an ordinary linear chromosome.
Copy number changes what a “genotype” means
A diploid autosomal locus usually has two chromosomal copies per nucleated cell. mtDNA exists in many copies distributed among multiple mitochondria, and copy number varies by cell type, physiological state, extraction, and specimen quality. High copy number often helps recover mtDNA when nuclear DNA is scarce, but it creates a population of molecules rather than a single pair of alleles.
If essentially all sampled molecules carry the same state at a position, that state is described as homoplasmic. If more than one state is present, it is heteroplasmic. The observed fraction is an assay-specific estimate from a particular specimen. It can differ among tissues, cells, and time points and can be distorted by amplification, sequencing error, contamination, or nuclear insertions of mitochondrial origin.
A consensus FASTA normally records one symbol per position. It can suppress minority states, encode them as IUPAC ambiguity symbols, or insert N, depending on the upstream workflow. Without the calling policy and underlying evidence, a consensus alone cannot establish homoplasmy or quantify heteroplasmy.
Inheritance supports lineage analysis, with boundaries
Human mtDNA is transmitted predominantly through the maternal line. Because recombination is generally treated as absent for routine human mtDNA phylogenetics, mutations accumulated along maternal lineages form a nested tree. A haplogroup caller compares an observed variant pattern with branch-defining mutations in a specified phylogeny.
This inheritance model has several consequences:
- People sharing a maternal-line ancestor can share an mtDNA haplotype even when most of their genomes differ.
- A match does not uniquely identify an individual; many maternally related or unrelated people may share the same common haplotype.
- mtDNA says nothing directly about the many other ancestors represented by autosomal and sex-chromosome DNA.
- A haplogroup is a phylogenetic classification, not a precise geographic origin, ethnicity, or date.
Rare reports of paternal transmission do not justify assuming it in routine analysis. If a result appears to require an exceptional inheritance mechanism, sample identity, contamination, alignment, nuclear mitochondrial sequences, and technical replication should be examined first.
The nuclear genome can imitate mtDNA
Segments derived from mtDNA have accumulated in the nuclear genome; these are nuclear mitochondrial DNA segments, or NUMTs. An assay can co-amplify or map reads from a NUMT, particularly when mitochondrial template is limited, primers are not specific, or short reads cannot distinguish the loci. NUMTs may create false differences from the mitochondrial reference or apparent low-level mixtures.
Mitigation depends on the assay. Long-range mtDNA enrichment, mtDNA-specific primer design, mapping-quality filters, paired-read context, coverage patterns, and comparison with a nuclear-aware reference can all contribute. No single filter is universally sufficient. The report should state how NUMTs were considered rather than merely asserting that reads “mapped to chrM.”
From specimen to interpretable result
A defensible workflow preserves transformations and decision points:
- Define the specimen. Record tissue or material, collection, extraction, sample identifier, controls, and chain of custody where relevant.
- Preserve primary data. Retain chromatograms or reads, instrument metadata, and checksums or immutable identifiers for uploaded files.
- Document assay scope. State primers or enrichment, sequenced coordinate range, platform, strand coverage, and known blind spots.
- Review quality locally. A high mean quality cannot rescue a weak base at the position driving an interpretation. Inspect trace shape or read-level support, alignment context, and coverage.
- Name the reference. Report the sequence accession or reconstruction, coordinate convention, and whether variants are expressed relative to rCRS or RSRS.
- Separate calling from classification. Preserve the consensus or variant profile before haplogroup assignment, then record caller, tree build, parameters, quality score, missing expected variants, and private variants.
- Interpret at the supported resolution. Partial control-region data should not be reported as if it were a complete mitogenome. Mixed signals need an assay-specific heteroplasmy method.
Common interpretation mistakes
Treating every reference difference as unusual. Most differences are ordinary lineage-associated variation. Their meaning depends on phylogenetic context and population frequency, not merely on being different from rCRS.
Equating read count with independent molecules. PCR duplicates and amplification bias can make thousands of reads derive from far fewer templates. Molecular depth and sequencing depth are not interchangeable.
Ignoring the reference breakpoint. Poor handling around positions 16,569/1 can produce an artificial coverage gap or missed variants.
Calling an ambiguous Sanger peak heteroplasmy. Dye artifacts, baseline noise, compression, carryover, and poor extension can also produce secondary peaks. Confirmation and a validated mixture threshold are needed.
Using a haplogroup quality score as a biological probability. It measures fit under a tool’s scoring model; it is not the probability that all sequence calls are correct or that a geographic narrative is true.
Review and reporting checklist
| Domain | Minimum record | Question for review |
|---|---|---|
| Specimen | Material or tissue, extraction, sample ID, controls | Is the analyzed material the one the interpretation refers to? |
| Assay | Target range, primers or enrichment, platform | Could dropout, NUMTs, or the circular breakpoint affect the result? |
| Primary evidence | Raw file, quality values, trace or read support | Can a reviewer return to the evidence behind a key base? |
| Sequence | Consensus policy, ambiguous-base policy, coverage | What information was discarded when the consensus was made? |
| Reference | Name, accession or build, coordinate convention | Are all reported variants in the same notation system? |
| Classification | Tool, version, tree build, range, quality | Does the stated haplogroup resolution match the observed range? |
| Interpretation | Purpose, limits, population context | Are identity, ancestry, or biological claims stronger than the data? |
Where GeneFlow fits
GeneFlow keeps uploaded AB1, FASTA, FASTQ, and VCF files associated with project samples. For AB1 and quality-bearing consensus FASTQ records, it stores base calls and Phred values; AB1 records also retain chromatogram channels. It can summarize Phred quality, flag weak intervals, suggest trim coordinates, classify sequence or rCRS-relative VCF inputs with HaploGrep3, build project trees with MAFFT and FastTree, and generate reports linked to the project.
These functions support provenance and review, but they do not validate an assay or make clinical interpretations. GeneFlow’s current workflow also does not detect or quantify heteroplasmy. A consensus, ambiguous symbol, secondary chromatogram peak, or HaploGrep3 mutation list must not be presented as a heteroplasmy measurement without an external, validated calling procedure.