Skip to main page content
U.S. flag

An official website of the United States government

Dot gov

The .gov means it’s official.
Federal government websites often end in .gov or .mil. Before sharing sensitive information, make sure you’re on a federal government site.

Https

The site is secure.
The https:// ensures that you are connecting to the official website and that any information you provide is encrypted and transmitted securely.

Access keys NCBI Homepage MyNCBI Homepage Main Content Main Navigation
Review
. 2016 Oct;29(4):881-913.
doi: 10.1128/CMR.00001-16. Epub 2016 Sep 7.

A Primer on Infectious Disease Bacterial Genomics

Affiliations
Review

A Primer on Infectious Disease Bacterial Genomics

Tarah Lynch et al. Clin Microbiol Rev. 2016 Oct.

Abstract

The number of large-scale genomics projects is increasing due to the availability of affordable high-throughput sequencing (HTS) technologies. The use of HTS for bacterial infectious disease research is attractive because one whole-genome sequencing (WGS) run can replace multiple assays for bacterial typing, molecular epidemiology investigations, and more in-depth pathogenomic studies. The computational resources and bioinformatics expertise required to accommodate and analyze the large amounts of data pose new challenges for researchers embarking on genomics projects for the first time. Here, we present a comprehensive overview of a bacterial genomics projects from beginning to end, with a particular focus on the planning and computational requirements for HTS data, and provide a general understanding of the analytical concepts to develop a workflow that will meet the objectives and goals of HTS projects.

PubMed Disclaimer

Figures

FIG 1
FIG 1
Example of three HTS applications for infectious disease bacterial genomics. These applications use HTS data to answer common questions regarding bacterial pathogenesis in a public health/clinical microbiological research setting from bacterial typing to molecular epidemiology and in-depth pathogenomic investigations. These applications are shown in a feedback loop to demonstrate that HTS provides data that can be analyzed to various degrees (both depth and breadth) based on the hypotheses under test and the number of isolates included for comparative genomics.
FIG 2
FIG 2
General HTS project stages and timeline. The importance and time requirements for the project planning stage and ongoing project management are sometimes underestimated but are invaluable for large-scale HTS projects for which hundreds of samples and terabytes of data are produced. The stages are divided by natural timeline progression and also increasing depth of investigation and specialized analysis requirements.
FIG 3
FIG 3
Overview of bacterial genome annotation. Structural annotation identifies the location of genes on the contigs of an assembled bacterial genome. Protein-encoding locations are identified, followed by automated assignment of gene function by comparison to existing databases. Non-protein-encoding genes are annotated by identifying key signatures for each type of gene. The resulting annotations are combined with an optional manual curation that can be performed before the final annotated genome is produced.
FIG 4
FIG 4
General phylogenetic concepts. (A) The general structure of a phylogenetic tree consisting of nodes and branches. A terminal node represents a particular organism (or sequence) under study. Terminal nodes are connected to internal nodes representing hypothetical ancestors. A common ancestor and all of its descendants are referred to as a clade. The ancestor to all descendants in the tree is called the root. (B) A cladogram showing the relative recency of common ancestry. Branch lengths are not informative in a cladogram. (C) A phylogram showing both the relative recency of common ancestry and evolutionary distance. Branch lengths are scaled to reflect the evolutionary distance between an organism and its ancestors.
None
None
None
None
None

References

    1. Portny SE, Austin J. 12 July 2002. Project management for scientists. American Association for the Advancement of Science, Washington, DC: http://www.sciencemag.org/careers/2002/07/project-management-scientists.
    1. Sandve GK, Nekrutenko A, Taylor J, Hovig E. 2013. Ten simple rules for reproducible computational research. PLoS Comput Biol 9:e1003285. doi: 10.1371/journal.pcbi.1003285. - DOI - PMC - PubMed
    1. Vos RA. 2016. Ten simple rules for managing high-throughput nucleotide sequencing data. bioRxiv doi: 10.1101/049338. - DOI
    1. Liolios K, Schriml L, Hirschman L, Pagani I, Nosrat B, Sterk P, White O, Rocca-Serra P, Sansone SA, Taylor C, Kyrpides NC, Field D. 2012. The Metadata Coverage Index (MCI): a standardized metric for quantifying database metadata richness. Stand Genomic Sci 6:438–447. doi: 10.4056/sigs.2675953. - DOI - PMC - PubMed
    1. Gargis AS, Kalman L, Berry MW, Bick DP, Dimmock DP, Hambuch T, Lu F, Lyon E, Voelkerding KV, Zehnbauer BA, Agarwala R, Bennett SF, Chen B, Chin EL, Compton JG, Das S, Farkas DH, Ferber MJ, Funke BH, Furtado MR, Ganova-Raeva LM, Geigenmuller U, Gunselman SJ, Hegde MR, Johnson PL, Kasarskis A, Kulkarni S, Lenk T, Liu CS, Manion M, Manolio TA, Mardis ER, Merker JD, Rajeevan MS, Reese MG, Rehm HL, Simen BB, Yeakley JM, Zook JM, Lubin IM. 2012. Assuring the quality of next-generation sequencing in clinical laboratory practice. Nat Biotechnol 30:1033–1036. doi: 10.1038/nbt.2403. - DOI - PMC - PubMed

LinkOut - more resources