A useful mtDNA report should let a reader understand both the result and how it was produced. The exact fields depend on whether the work is exploratory research, population analysis, forensic testing, or a validated clinical workflow.
The report is a controlled view of evidence, not a replacement for source data, workflow records, or expert review. Define its audience and purpose on the first page. A research report should not inherit clinical or forensic language unless it was produced under the corresponding validated system.
1. Identify scope and status
State whether the document is draft, reviewed, or superseded; whether it covers one sample or a cohort; and the precise question it addresses. Include a stable report identifier, version, generation date/time zone, project, author/generator, reviewers, and approval date where applicable.
Separate the analysis date from the report generation date. Regenerating a PDF later should not imply that tools or databases were rerun. Cite the exact pipeline run and result versions included.
2. Describe samples and inputs
| Item | Minimum reporting detail |
|---|---|
| Sample identity | stable non-identifying code; biological/technical replicate relationships |
| Context | specimen/material type and collection context needed to interpret the research question |
| Assay | extraction/amplification/sequencing method, primers or capture design, instrument/platform |
| Target | complete mitogenome or stated region; expected and reported coordinates |
| Source data | filenames or artifact IDs, formats, checksums, acquisition/run identifiers |
| Reference | name, accession and version/checksum, topology and coordinate convention |
| Exclusions | excluded samples/files/intervals with reasons and decision-maker |
Avoid placing direct participant identifiers in analytical filenames or report tables. Follow the project’s approved data-governance process for any link to identity.
3. Report quality evidence and decisions
A quality label without its rule is not interpretable. For Sanger data, report the base caller when known, quality threshold and window, proposed and accepted trim intervals, low-quality regions, strand coverage, control/replicate results, and manual chromatogram review. Distinguish automated flags from reviewer dispositions.
For high-throughput data, report the metrics material to the conclusion: read counts before/after filtering, mapping/reference strategy, on-target and duplicate handling, depth distribution and callable range, base/mapping-quality filters, strand balance, contamination checks, and any NUMT-aware procedure. Do not transplant Sanger Phred summaries into an NGS acceptance claim.
For every QC exception, include the observation, decision, rationale, reviewer, and effect on downstream analysis. If a failed or warning sample remains included, make that visible in result tables.
4. Explain consensus or variant construction
Report whether downstream analysis used AB1 base calls, a reviewed consensus FASTA, a quality-bearing consensus, or VCF variants. Then provide:
- software/version and complete material parameters;
- primer/adapter trimming and orientation rules;
- read overlap, base selection, depth, ambiguity, no-call, and indel rules;
- coordinate transformations around the circular origin;
- manual edits with original state and rationale;
- final sequence/VCF identifier and checksum;
- callable/covered interval and missing-data proportion.
If reporting a VCF, define FILTER, INFO, and FORMAT fields used in interpretation. Absence from a sparse VCF must not be described as a reference call unless callability supports it. If reporting a consensus, explain what N, IUPAC symbols, and gaps mean in that workflow.
5. Present findings with supporting evidence
Lead with a concise result table, then provide evidence and exceptions. Keep observations distinct from interpretations.
For each haplogroup result, report the classifier and version, phylogenetic tree/database build, input sequence or VCF version, covered range, returned haplogroup, tool quality/confidence field, found/supporting variants, expected-but-missing variants when available, private/remaining variants, warnings, and failed classifications. A software quality score is tool-specific; do not translate it into an assay accuracy claim.
When multiple classifiers are used, list each result before presenting a consensus. Explain whether differences are nomenclature depth, database version, input handling, or a substantive conflict. Name the person and evidence used for any manual resolution.
For a phylogenetic result, include the input cohort and consensus versions, common analyzed interval, alignment and site exclusions, inference software/versions/commands, substitution model, support method, rooting/outgroup, label map, Newick tree, and alignment. Captions should define branch-length units and node labels. State clearly if the tree is unrooted.
6. Make limitations result-specific
Useful limitations explain how uncertainty affects the conclusion. Consider:
- partial sequence or nonuniform coverage;
- unresolved positions, strand conflicts, low-quality regions, and manual edits;
- detection limits for low-level variation under the assay and validation status;
- contamination, sample mixture, index misassignment, or NUMT risk where relevant;
- reference choice, circular-coordinate and indel-representation effects;
- classifier database/tree coverage and nomenclature version;
- cohort size, sampling bias, missing comparators, and model assumptions;
- steps performed manually or by an external service without complete provenance.
Use direct language: “positions 1-72 were not assessed” is more useful than “minor limitations may apply.” Do not claim that a matching mtDNA sequence uniquely identifies an individual, or that a tree proves a pedigree or migration history.
7. Attach a reproducibility package
The report should point to, or be distributed with, a manifest containing:
- source and derivative artifact IDs/checksums;
- sample/replicate map and inclusion table;
- workflow definition, code revision, tool/database versions, parameters, and environment;
- per-step run records, logs, warnings, and errors;
- reviewed consensus/VCF, alignment, Newick, and machine-readable result tables;
- manual decision log and report change history.
If data cannot be shared, state the reason and describe the controlled access route or which metadata remain available. A PDF alone is rarely sufficient for reanalysis.
8. Perform a release review
Before release, have a reviewer verify:
- Every sample and result maps to the intended source and run.
- References, tools, tree/database builds, units, and coordinate systems are explicit.
- Summary values agree with machine-readable outputs.
- Warning/failed cases and missing data are not presented as successful results.
- Figures have label maps, legends, scales, and reproducible source artifacts.
- Interpretations do not exceed assay, sampling, or method capability.
- Superseded versions remain traceable and sensitive metadata are appropriate for the audience.
Where GeneFlow fits
GeneFlow defines QC, haplogroup, phylogenetic, and full report records with project, optional sample, structured JSON payload, stored downloadable artifact, creator, and an optional pipeline-run relationship. Current generated PDFs provide project-level QC/haplogroup summary tables or a single-sample full view containing source filename/kind, sequence length, Phred summary and trim suggestion, HaploGrep3 consensus/quality, and found mutations.
These are operational summaries, not complete researcher-grade reports. Generation does not add assay details, source checksums, accepted chromatogram decisions, consensus rules, exact executable/database versions, alignment, rooting/support methods, or a limitations narrative. A defined phylogenetic report type should not be assumed to contain the underlying alignment and tree methodology. Also verify and populate the relevant pipeline-run relationship rather than assuming report generation inferred it. Researchers must review the structured payload against source evidence, add external/manual processing records and interpretation, and follow their own validated release process where one applies.