A Computational Protocol for Enhanced Characterization of Sequence Differences Among Salmonella enterica Isolates within SNP Clusters Identified in the NCBI Pathogen Detection System.
Hiruni R Wijesena, Dayna M Harhay, Tatum S Katz, Tommy L Wheeler, John W Schmidt, Terrance M Arthur +1 more
Journal of food protection
Abstract
Salmonella enterica remains a major public health concern in the United States due to its high serotype diversity and its association with foodborne outbreaks. Genomic surveillance through the Pathogen Detection Isolates Browser (PDIB) has facilitated tracking of Salmonella isolates, but limitations in draft genome quality and reliance on single nucleotide polymorphism (SNP) analysis alone may obscure important genetic differences between isolates. To address this gap, we developed a two-part comparative genomic protocol that integrates both core chromosome and variable genome analyses by including nearest-neighbor complete closed reference genomes in the evaluation. This protocol enhanced the resolution of genetic relatedness assessment among Salmonella isolates within SNP clusters identified in the PDIB attributed to four meat commodities (beef, chicken, pork, and turkey). Genomic relatedness analysis was conducted on 354 isolates spanning 59 SNP clusters and 21 serotypes. While our core chromosome SNP analysis aligned closely with PDIB data, the variable genome analysis revealed additional genetic differences such as large deletions, prophages, and plasmid variability, which were not detected by the PDIB analysis. In outbreak case studies involving serotypes Newport and Hadar, our method uncovered genomic features that could be beneficial for distinguishing related isolates from those unrelated to the outbreaks. These findings demonstrate that reliance on SNP analysis alone likely masks relevant genomic variation, potentially undermining outbreak traceability and risk assessment. The protocol presented here offers a more complete view of Salmonella genomic diversity, surpassing the capabilities offered by public databases, while also providing a practical, resource efficient tool for conducting private genomic data analyses for Salmonella without the requirement of uploading sequences to public databases.