AI Tool Improves Long-Read Detection of Cancer Mutations

By LabMedica International staff writers
Posted on 29 Jul 2026

Accurate identification of somatic mutations remains difficult, particularly in structurally complex genomic regions. Short-read sequencing pipelines can miss clinically relevant variants, limiting research and precision oncology. Long-read sequencing expands genomic coverage but requires models that can be trained despite limited tumor datasets. A new study describes a deep-learning approach to improve small-variant detection across multiple tumor types.

The University of Hong Kong (HKU) School of Computing and Data Science has developed ClairS, a deep-learning algorithm for detecting cancer mutations using long-read sequencing. The system targets somatic small variants present in tumors but absent from matched normal tissue. It was created to overcome limitations of short-read methods that struggle in complex genomic regions.


Image: Cover artwork featuring the ClairS study (Photo courtesy of Professor Ruibang Luo/HKU)

ClairS introduces a training-data synthesis strategy to address the shortage of high-quality labeled cancer datasets. The team mixes sequencing reads from normal human samples to create realistic synthetic tumor–normal pairs, generating virtually unlimited training examples that span different tumor purities, sequencing depths, and mutation burdens. This yields a flexible, robust artificial intelligence (AI) model suited to real-world cancer genomic analysis.

The algorithm was tested on cell-line datasets from breast cancer, lung cancer, and melanoma, demonstrating high accuracy across cancer types and sequencing conditions. Additional evaluations across multiple cancer cell lines, including pancreatic models, reinforced its performance for detecting small mutations in long-read data. The evaluations highlight how long-read sequencing can reveal mutations that might otherwise be missed.

ClairS has been integrated into Oxford Nanopore Technologies’ official somatic variant-calling workflow, placing it within a practical commercial analysis pipeline. The research appears in Nature Methods under the title “ClairS: a deep-learning method for long-read tumor–normal pair somatic small variant calling.” The software is open source and available on GitHub. The work highlights a scalable way to train medical AI when real clinical data are scarce.

“Long-read sequencing is transforming how we study cancer genomes, especially in regions that were previously difficult to analyse. ClairS makes it possible to train powerful AI models even when real cancer training data is limited, supporting more reliable cancer mutation discovery from long-read sequencing data,” said Ruibang Luo, Assistant Director (Learning Experience & Student Enrichment) and Associate Head of the Department of AI & Data Science, School of Computing and Data Science, The University of Hong Kong.

Related Links
The University of Hong Kong


Latest Molecular Diagnostics News