Porject title: Independent Bioinformatics Verification of the Hepatitis Delta Virus Sequence Analysis Reported by [Butler et al. (2018). DOI:10.1038/s41598-018-30078-5]
- Introduction
Hepatitis Delta Virus (HDV) is a defective, single-stranded circular RNA virus belonging to the family Kolmioviridae and is recognized as the smallest known human pathogen capable of causing viral hepatitis. Unlike conventional viruses, HDV is replication-defective and requires the hepatitis B virus (HBV) for its life cycle, utilizing the HBV surface antigen (HBsAg) to assemble infectious virions and facilitate transmission. Consequently, HDV infection occurs only in individuals infected with HBV, either through simultaneous coinfection or superinfection of chronic HBV carriers. While coinfection often results in acute hepatitis that resolves spontaneously, superinfection frequently progresses to chronic HDV infection and is associated with rapid liver disease progression, including severe hepatitis, liver cirrhosis, hepatic decompensation, and hepatocellular carcinoma (HCC).
- Rational
In recent years, reproducibility has become a fundamental principle of computational biology and bioinformatics. As bioinformatics analyses often involve multiple computational steps—including sequence retrieval, quality control, sequence alignment, phylogenetic reconstruction, genetic diversity estimation, and statistical analyses—the availability of reproducible computational workflows is essential for validating published findings, ensuring transparency, and facilitating future research. Reproducible research practices, supported by open-source programming languages such as Python, interactive computational environments such as JupyterLab, and publicly accessible sequence repositories, enable independent verification of published analyses while promoting confidence in scientific conclusions. The increasing availability of public nucleotide sequence databases and open-source bioinformatics software provides an opportunity to independently verify published molecular epidemiological studies. Such computational reproduction studies are particularly valuable because they evaluate the robustness of previously reported findings, identify potential methodological discrepancies, and provide reusable analytical workflows that can be adapted for future investigations. Furthermore, reproducibility studies contribute to the growing movement toward open science by making computational methods transparent, accessible, and verifiable.
- Aim
The present study aimed to independently reproduce and verify the principal bioinformatics analyses reported by Butler et al. (2018) using publicly available HDV nucleotide sequences obtained from the National Center for Biotechnology Information (NCBI). Using Python implemented within JupyterLab on a Linux environment, an open and fully reproducible computational workflow was developed to perform sequence quality control, genotype classification, genetic diversity analysis, prevalence statistics, and variable site entropy analysis. By comparing the reproduced results with those originally reported by Butler et al. (2018), this study assesses the reproducibility of the published analyses while providing a transparent workflow that can be reused for future HDV molecular epidemiology studies.
- Materials and Methods
Publicly available HDV nucleotide sequences corresponding to those analyzed by Butler et al. (2018) were retrieved from the NCBI GenBank database. Analyses were performed in Python using JupyterLab in a Linux environment. Quality control included inspection of sequence integrity, removal of problematic records where applicable, and preparation of FASTA datasets for downstream analyses. Genotype classification was performed by comparison with HDV reference sequences using phylogenetic/bioinformatic approaches implemented in Python. Genetic diversity was assessed through pairwise genetic distance analyses and visualized with heatmaps. Genotype prevalence was summarized using descriptive statistics. Sequence conservation and variability were quantified using Shannon entropy calculated from multiple sequence alignments. All scripts were designed to provide a transparent and reproducible computational workflow.
- Results
This Project folder contains the following;
A. Files
- Emily K. Butler et al 2018.pdf
- HDV sequences.fasta
- README.md
B. Folders with files
-
Sequence QC
- My_sequence_QC.ipynb
- hdv_sequence_qc_report.csv
- hdv_qc_plots.png
- README.md
-
Genotype classification
- My_GT_classif.ipynb
- hdv_references.fasta
- HDV_filtered_sequences.fasta
- unified_dataset_for_alignment.fasta
- aligned_hdv_dataset.fasta
- Phylogenetic tree files
- README.md
-
Genetic diversity
- My_Gen_divers,ipynb
- hdv_aligned.fasta
- Pairwise_Genetic_Distance.png
- HDV Genetic Diversity Across Genome.png
- README.md
-
Prevalence statistics
- My_Prev_Stats.ipynb
- Recomputed Prevalence Statistics.PNG
- README.md
-
Variable sit-entropy analysis
- My_VarSite_Ent_Anal.ipynb
- hdv_cameroon_aligned.fasta
- hdv_entropy_results.csv
- hdv_entropy_profile.png
- README.md
- General conclusion
A reproducible Python-based workflow successfully replicated the principal HDV bioinformatics analyses reported by Butler et al. (2018). The workflow provides a reusable framework for future molecular epidemiology studies of HDV and other viral pathogens. Overall agreement with the published findings supports the robustness of the original analyses while demonstrating that the workflow can be reproduced using publicly available sequence data.