Variant calling is the process of identifying genetic differences from high-throughput sequencing data. By analyzing the genome of an individual or a population, variant calling pipelines can identify several types of genetic variation, including Single Nucleotide Polymorphisms (SNPs), insertions and deletions (Indels), structural variants (SVs), and copy number variations (CNVs). These variants can then be used in downstream analyses to investigate genetic diversity and correlations with inherited conditions and diseases. Therefore, efficient and accurate variant calling is an important step in genomic analysis.
Variant Calling pipeline
The variant calling pipeline combines several bioinformatics applications implemented in different programming languages, including C++, Java, and Perl. Its core components are based on the GATK4 toolkit, one of the most widely used frameworks for variant calling. Other tools used in the pipeline include Trim Galore for read preprocessing, Bowtie2 for sequence alignment, and Annovar for variant annotation.
The pipeline contains several computationally intensive stages that can be executed independently. In particular, three levels of parallelism are introduced using the scatter feature of the Common Workflow Language (CWL). The first scatter distributes patient samples across independent instances of the generate sample recalibrated sub-workflow. A second scatter distributes chromosome intervals across the togvcf step, while a third distributes chromosome intervals across the generate chromosome vcf sub-workflow.
The original pipeline was developed using Snakemake, a widely adopted workflow system in bioinformatics.By tuning the application parameters and introducing additional scatter operations, we obtained a 3× speedup compared to the original pipeline.
StreamFlow application
We used this pipeline to evaluate StreamFlow against Snakemake and, in particular, to investigate the benefits of executing the workflow across hybrid cloud-HPC infrastructures. The original Snakemake workflow was first ported to CWL, allowing its computational structure to be described independently of the underlying execution environment. The resulting workflow was executed on two resources provided by CINECA in Bologna, Italy: the Galileo100 HPC cluster and a virtual machine on the Ada Cloud infrastructure.
The experiments were designed to compare workflow execution in Cloud and HPC environments and to evaluate the hybrid workflow paradigm supported by StreamFlow. This paradigm allows individual workflow steps to be executed on the infrastructure that is best suited to their computational requirements, combining Cloud and HPC resources within the same application. At the time of the experiments, this capability was not available in Snakemake.
The results show that StreamFlow and Snakemake achieve comparable execution times when running the pipeline on either Cloud or HPC resources. More importantly, distributing the workflow across the two environments using StreamFlow’s hybrid execution capabilities resulted in an additional 20% improvement in execution time.
These results demonstrate that StreamFlow can provide performance comparable to established bioinformatics workflow systems while enabling a hybrid Cloud-HPC execution model. By combining the strengths of different infrastructures and exploiting parallelism at multiple levels of the pipeline, StreamFlow can improve the overall efficiency of complex genomic workflows.
A. Mulone, S. Awad, D. Chiarugi, and M. Aldinucci, “Porting the Variant Calling Pipeline for NGS data in cloud-HPC environment,” in 47th IEEE Annual Computers, Software, and Applications Conference, COMPSAC 2023, p. 1858-1863, 2023. doi:10.1109/COMPSAC57700.2023.00288