An Integrated Computational Framework for Replication Site and Regulatory Motif Discovery in Coronavirus Genomes
Contributors
Dr. Pushpa Susant Mahapatro
Keywords
Proceeding
Track
General Track
License
Copyright (c) 2026 Sustainable Global Societies Initiative

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
Abstract
Coronavirus genomes contain regulatory regions that play an important role in genome replication, transcription, and other stages of the viral life cycle. Identifying these regions computationally is difficult because coronavirus sequences show considerable variation, while regulatory motifs are often short, weakly conserved, and susceptible to nucleotide changes. Although several computational approaches have been used to study regulatory sequences in coronaviruses, there is still limited comparative evaluation of motif discovery methods when the target motifs contain mutations or non-contiguous patterns. This study proposes a computational framework for locating candidate replication-associated regions and identifying regulatory motifs in coronavirus genomes. Complete coronavirus genome sequences will be collected from publicly available databases and subjected to quality filtering, alignment, and comparative analysis. Genome-wide sequence characteristics and conservation patterns will be examined to identify regions that may be associated with replication. Four motif discovery techniques—Greedy Motif Search, Greedy Motif Search with Pseudocounts, Randomized Motif Search, and Gibbs Sampling—will be evaluated using a common experimental framework. The analysis will specifically consider motifs containing one or two nucleotide mutations, as well as non-contiguous patterns that may be overlooked by conventional exact motif searches. The performance of the four approaches will be assessed in terms of motif identification accuracy, consistency, tolerance to mutations, and computational efficiency. Candidate motifs will further be assessed against conserved genomic regions, existing annotations, and available biological and structural evidence. The framework is intended to provide a reproducible approach for prioritizing potentially important coronavirus genomic regions for further experimental investigation. Such computationally prioritized regions may subsequently support studies of antiviral targets and contribute to genomic surveillance of emerging coronavirus variants.