Mitochondrial Haplogroups: Assignment, Evidence, and Interpretation
Mitochondrial DNA Fundamentals

Mitochondrial Haplogroups: Assignment, Evidence, and Interpretation

How mtDNA haplogroups are inferred from phylogenetic variant patterns, why resolution and quality vary, and how to review and report a call.

Photo by Terry Vlisidis on Unsplash · View photo

A mitochondrial haplogroup is a named branch in a phylogeny of mtDNA lineages. It is inferred from a pattern of variants, not read directly from one diagnostic position. A useful result therefore includes the evidence that places a sample on the tree: observed range, expected and missing branch markers, private differences, reference system, tree build, and the classifier’s scoring output.

The label is a compact summary of maternal-line phylogenetic placement. It is not a unique identifier, a complete ancestry analysis, or a biological interpretation of every variant in the sample.

Haplotypes, haplogroups, and trees

A haplotype is the observed sequence or variant combination over a defined range. A haplogroup is a clade containing related haplotypes that descend from a phylogenetic node. The distinction matters because two samples can receive the same haplogroup while carrying different private variants, and the same short haplotype can be compatible with several deeper lineages when decisive coding-region positions were not sequenced.

Human mtDNA phylogenies are reconstructed from many published sequences. Mutations are assigned to branches under explicit analytical conventions. Names are hierarchical but historically accumulated, so their syntax is not always a simple taxonomy that can be parsed reliably from characters alone. Tree releases change as additional genomes reveal new subclades, recurrent mutations, or earlier errors.

A haplogroup should therefore be understood as: “best-supported placement under this tree and method, given these observed positions.” Removing any part of that statement changes its meaning.

How a classifier reaches a call

For a candidate branch, a caller compares the sample profile with mutations expected along the path from the tree root to that branch. A typical scoring procedure rewards expected mutations that are found and penalizes expected mutations that are absent, while accounting for sequence range, recurrent sites, reversions, and extra variants. Different tools use different distance functions and weighting schemes.

Three evidence classes are especially informative:

  • Found expected variants support placement on a branch.
  • Missing expected variants can reflect limited coverage, a no-call, sequencing error, reversion, incorrect notation, or a poor candidate branch.
  • Private or remaining variants are observed differences not assigned to the selected path. They may be genuine recent changes, known recurrent sites, alignment artifacts, contamination, or base-calling errors.

No category should be interpreted mechanically. A missing marker outside the sequenced range is not contradictory evidence. A missing marker within a high-quality callable region is more consequential. A private variant at a mutational hotspot deserves different weight from a novel, well-supported coding-region difference.

Resolution follows information, not label length

Whole-mitogenome data usually support finer placement than HVS-I or control-region data because many clades share control-region motifs but differ at coding-region positions. However, nominally complete sequencing does not guarantee complete information. Coverage gaps, ambiguous calls, primer dropout, or poor alignment can remove the decisive sites.

A classifier may return a specific-looking label from partial data because that branch scores best among available candidates. Specificity is not the same as resolution. Review the reported range and missing markers, and if possible compare the call with the most recent common ancestor of similarly scoring candidates.

Conversely, a broad call from a full mitogenome is not automatically a failure. The sample may sit near a currently named node, contain a lineage not yet represented in the tree, or carry unresolved data problems. Forcing every sequence into a deeply named terminal branch overstates current phylogenetic knowledge.

Why callers or runs disagree

Two defensible classifications can differ because they use:

  • different phylogeny releases or reference systems;
  • different scoring functions and mutation weights;
  • different handling of heteroplasmic or ambiguous symbols;
  • different indel normalization and control-region alignment;
  • different input ranges or no-call representations;
  • different rules for genotyping arrays, VCFs, and complete sequences.

Agreement at a parent clade with disagreement only in a downstream subclade is not the same as conflict between distant branches. Compare placements on the tree, not just whether label strings are identical.

A phylogenetic plausibility check is also valuable QC. A profile that combines expected markers from distant branches may indicate sample mixture, contamination, artificial recombination from assembly, reference conversion errors, or transcription mistakes. The tree is not merely an ancestry output; it is a structured expectation against which the sequence can be checked.

Quality scores are not posterior probabilities

Haplogroup tools commonly return a quality or distance score. Its definition belongs to that tool and version. A score of 0.9 should not be reported as “90% probability that the sample belongs to this haplogroup” unless the method explicitly defines and validates it as such. It does not measure sample identity, population membership, or correctness of every observed base.

Inspect the components behind the score. A high score supported by a short hypervariable segment may still permit several more specific placements. A lower score from a full genome may expose one contradictory high-quality position that warrants laboratory review.

Interpreting a haplogroup responsibly

Because mtDNA follows one predominantly maternal lineage, a haplogroup represents one path among a person’s many genealogical ancestors. Population frequencies can provide historical context at an appropriate geographic and temporal scale, but present-day labels do not establish nationality, ethnicity, migration date, or the location of a particular recent ancestor.

Haplogroup frequencies are properties of sampled populations and are sensitive to sampling design. The absence of a lineage from a database is not proof that it is absent from a population. A haplogroup match between two specimens does not uniquely identify a source; common haplotypes may be shared widely, and maternally related people are expected to match closely.

Haplogroup assignment should also be separated from variant interpretation. A branch-defining variant can be common within a lineage. Being listed in a disease-oriented database does not by itself establish an effect in a particular person, and GeneFlow does not provide clinical interpretation.

Common interpretation mistakes

Calling one mutation diagnostic in isolation. Recurrent mutation and back mutation occur. Evaluate the complete observed motif and range.

Ignoring missing expected variants. A top label without its contradictions conceals the most useful QC information.

Treating partial and complete sequences as equivalent. Their labels can differ in depth even when they are compatible.

Comparing labels without tree versions. Nomenclature and branch definitions evolve.

Calling a de novo sequence tree a haplogroup tree. A MAFFT/FastTree tree estimates relationships among submitted sequences under its own model. It does not automatically map samples onto the curated global mtDNA nomenclature.

Review and reporting checklist

ItemReportReview question
InputFASTA, VCF, HSD, array profile; referenceWas the representation compatible with the caller?
ScopeSequenced and callable coordinatesWere decisive branch markers observed?
MethodTool/version, tree/build, distance functionCan the call be reproduced?
PlacementHaplogroup plus relevant parent cladeIs the stated specificity supported?
Supporting evidenceFound expected variantsAre key markers high quality?
ContradictionsMissing expected variants in callable rangeCould they indicate error or an alternate placement?
Additional evidencePrivate/remaining variants and hotspot contextDo any require trace or alignment review?
QualityTool score with its definitionIs it being mistaken for a posterior probability?
InterpretationMaternal-line scope and population sourceAre geographic or identity claims appropriately limited?

Where GeneFlow fits

GeneFlow currently classifies sequence-bearing AB1, FASTA, and consensus FASTQ samples with HaploGrep3; rCRS-relative VCFs can be passed to HaploGrep3 directly. It stores the top haplogroup, HaploGrep quality, found polymorphisms, and the extended raw output, which includes range and found, not-found, and remaining polymorphism fields. Failures are retained rather than silently discarded.

The shipped automated classifier is HaploGrep3 alone. Although the data model can represent other tools, a current “match” consensus status is a persisted single-tool result, not independent agreement between HaploGrep3 and MitoMaster. Researchers who require cross-tool concordance must run and document that comparison explicitly.

GeneFlow can also align project sequences with MAFFT and infer a FastTree tree when at least three sequence-bearing samples are available. That tree is useful for within-project exploration, but GeneFlow does not automatically harmonize partial and complete ranges or convert it into a curated haplogroup phylogeny. Input comparability remains the researcher’s responsibility.

Further reading