Internship Presentations

Development and Implementation of ISAWGS- LOC: A Nanopore WGS Pipeline for Integration Site Identification and Characterization​

Keertana Chagari

Mentor: Dr. Sandrine Moreira, Head of Bioinformatics, PathoQuest.

Date/Time: August 25th, 2026 at 2:15 PM.

Abstract: Gene delivery is used in the development of gene and cell therapies and other genetically modified biological products, creating a need to characterize where introduced genetic material integrates into the host genome. Identifying these integration sites can support the genomic characterization and safety assessment of client-submitted samples. Long-read whole-genome sequencing (WGS) can capture host and transgene sequences within individual reads, enabling analysis of both integration location and structure. We developed ISAWGS-LOC, a bioinformatics pipeline for identifying and characterizing transgene integration sites from Oxford Nanopore WGS data.

ISAWGS-LOC uses a modular Snakemake workflow to identify reads containing transgene sequence and evaluate their alignments to the host genome and vector. Confirmed integration reads are mapped to a combined host–vector reference, allowing host–vector junctions to be detected and genomic breakpoints to be estimated. Breakpoints supported by nearby reads are clustered to define integration sites. Read-level alignment segments are retained to preserve complex integration patterns for further site-specific sequence characterization.

The pipeline was successfully implemented and executed on Oxford Nanopore WGS data, producing integration-site calls with supporting-read, breakpoint, junction, and alignment information. The workflow also preserves reads containing multiple host and vector alignment segments for further characterization. Additional validation using known and simulated integration sites will evaluate breakpoint accuracy, sensitivity, false-positive detection, and supporting-read assignment.

ISAWGS-LOC provides a framework for determining where transgene integrations occur in the genome while retaining sequence-level information needed to characterize their structure. This approach supports more comprehensive analysis of transgene integration from long-read WGS data.

Tagged
Summer 2026
Summer 2026 #1