Skip to content
Pawsey Supercomputing Research Centre
  • Home
  • About Us
    • About Us
      • History
      • The Pawsey Centre
      • Discover Setonix
      • Our green credentials
      • Diversity, equity and inclusion
    • Our People
      • Board & Management
      • Staff List
      • Meet the Pawsey Champions
      • Pawsey’s business response to COVID-19
    • Documentation
      • Documents Library & Annual Reports
    • Contact
      • Pawsey Friends newsletter
      • Jobs
  • Services
    • Supercomputing and Data Processing
      • Setonix
      • Nectar Cloud
    • Data Workflows
      • Acacia
      • Banksia
      • Data Portal
    • Analysis and Visualisation
      • Nebula
      • Setonix RemoteVis
      • Visualisation Lab
      • Remote Virtual Reality
    • Support & Consultancy
      • PULSE Collaborations (formerly Uptake Projects)
      • Support
      • Pawsey Centre for Extreme Scale Readiness (PaCER)
    • Training, Education and Engagement
      • Training
      • Intern Alumni Profiles
      • Pawsey Internships: Call for Projects
      • Pawsey Internships: Call for Students
      • Intern Project Guidelines for Assessment
      • Upcoming events
      • Past events
  • New Technologies
    • New Technologies
      • Setonix-Q
      • Technology Refresh
      • Quantum Technologies
  • Science Showcase
    • Science Showcase
      • Case Studies
      • Project Allocations
      • 2026 Allocations
      • 2025 Allocations
      • 2024 Allocations
      • 2023 Allocations
      • Publications
      • HPC Hearts & Minds podcast
      • Voices of Science
      • COVID-19 Accelerated Access Initiatives
  • News & Events
    • News & Events
      • News
      • Events
      • Pawsey Friends newsletter
      • Friends Newsletter
      • Technical Newsletter
      • Pawsey Conference/Event Support Guidelines
      • Pawsey In The Media
  • Support
  • User Portal
  • Status
Apply now
Pawsey Supercomputing Research Centre
Apply now
  • Home
  • About Us
    • About Us
      • History
      • The Pawsey Centre
      • Discover Setonix
      • Our green credentials
      • Diversity, equity and inclusion
    • Our People
      • Board & Management
      • Staff List
      • Meet the Pawsey Champions
      • Pawsey’s business response to COVID-19
    • Documentation
      • Documents Library & Annual Reports
    • Contact
      • Pawsey Friends newsletter
      • Jobs
  • Services
    • Supercomputing and Data Processing
      • Setonix
      • Nectar Cloud
    • Data Workflows
      • Acacia
      • Banksia
      • Data Portal
    • Analysis and Visualisation
      • Nebula
      • Setonix RemoteVis
      • Visualisation Lab
      • Remote Virtual Reality
    • Support & Consultancy
      • PULSE Collaborations (formerly Uptake Projects)
      • Support
      • Pawsey Centre for Extreme Scale Readiness (PaCER)
    • Training, Education and Engagement
      • Training
      • Intern Alumni Profiles
      • Pawsey Internships: Call for Projects
      • Pawsey Internships: Call for Students
      • Intern Project Guidelines for Assessment
      • Upcoming events
      • Past events
  • New Technologies
    • New Technologies
      • Setonix-Q
      • Technology Refresh
      • Quantum Technologies
  • Science Showcase
    • Science Showcase
      • Case Studies
      • Project Allocations
      • 2026 Allocations
      • 2025 Allocations
      • 2024 Allocations
      • 2023 Allocations
      • Publications
      • HPC Hearts & Minds podcast
      • Voices of Science
      • COVID-19 Accelerated Access Initiatives
  • News & Events
    • News & Events
      • News
      • Events
      • Pawsey Friends newsletter
      • Friends Newsletter
      • Technical Newsletter
      • Pawsey Conference/Event Support Guidelines
      • Pawsey In The Media
  • Support
  • User Portal
  • Status

Nanopore Basecalling on Pawsey Supercomputer using AMD GPUs

By Bonson Wong  and Dr Hasindu Gamaarachchi, UNSW

 

Nanopore sequencing has become a popular technology for genomics research due to its cost-effectiveness and ability to sequence long reads. A nanopore sequencer generates a time-series “raw signal” or squiggles, which are converted into DNA bases (A, C, G, T) through a process called “basecalling”. Basecalling is essentially the first computational step in a nanopore sequencing experiment. Modern nanopore basecallers use deep learning-based models to significantly (by at least 10%) improve the accuracy of predicting a DNA base from the raw input signal compared to traditional non-deep learning-based basecallers. Oxford Nanopore Technologies (ONT) is the leading company in producing nanopore sequencers. “Dorado” is ONT’s most recent open-source basecaller that uses deep learning models for basecalling.

Dorado supports both CPUs and GPUs, but using GPUs is essential for practical runtime. Running Dorado on CPU alone would take a few weeks for a single sample, while with GPUs this is in order of hours or days.  Currently, Dorado supports NVIDIA GPUs, but not AMD GPUs. Pawsey’s Setonix cluster has +750 AMD GPUs that cannot be utilised for basecalling. We recently enabled basecalling on AMD GPUs, thus enabling basecalling on Pawsey, by implementing AMD GPU support onto “slorado”, a simplified version of Dorado that we developed. Slorado, which is completely open-source, supports both NVIDIA and AMD GPUs.

We benchmarked slorado on a Setonix GPU node with the four AMD Instinct MI250X GPUs (one AMD GPU has 2 dies, therefore 8 GPU dies in total), using a full HG002 human dataset (~20X coverage, 16 million reads) sequenced on an ONT R10.4.1 PromethION flowcell. We tested with all three v4.2.0 basecalling models and the execution time was ~21 hours for super accuracy basecalling (sup), ~7.5 hours for high accuracy basecalling (hac), and ~5 hours for fast basecalling (fast):

 

Basecalling model Execution time (hours)
super accuracy (sup) 21.1
high accuracy (hac) 7.5
fast basecalling (fast) 4.8

 

We indeed thoroughly verified the accuracy of basecalling on Pawsey’s AMD GPUs using slorado. For this, we aligned the basecalled reads to the hg38 genome using minimap2 and calculated the statistics (e.g., mean, median) for the identity scores. We compared these values to what is obtained from the same model versions in the ONT’s original Dorado software when run on a system with four NVIDIA Tesla V100 GPU. The identity scores from Slorado were very similar (even slightly better in some cases) to those from the original Dorado, as shown in Fig. 1 below.

Fig 1: Accuracy comparison of slorado vs Dorado using mean and median identity scores for a 20 million-read HG002 dataset

If you are interested in basecalling on Pawsey, simply follow the instructions here to get started. To install Slorado, just download the compiled binaries, or even get it compiled by yourself – please refer to the readme here.

Slorado currently supports v4 basecalling models (LSTM-based models), in line with what is available on ONT’s Dorado version used for live-basecalling on the PromethION sequencer (Dorado-server that comes with MinKNOW). The latest stand-alone version of ONT Dorado has v5 basecalling models and the work on supporting the new v5 models (transformer-based models) on Slorado is underway, and you can expect to see this in a future release.

For those who want details, let us go into some technical details and how we implemented AMD GPU support in Slorado.

Technical Details

The background story of slorado

Slorado was originally motivated as a research tool to benchmark and test basecalling on different computing architectures. Because of Dorado’s complicated setup, we needed a simplified stripped-down version of the basecaller. This way, we could learn about the details of every step in basecalling, measure execution time for different steps separately, and make the development process for ourselves manageable.

The Anatomy of Dorado and Koi

For the sake of brevity, basecalling on Dorado can be broken down into 2 major components; inference, and decoding. Inference happens when squiggle data is sent to the neural network. The neural network takes this data squiggle and outputs a table of scores for the probability of a base present at each timestep. In Dorado, this is mostly handled by the open-source torch library, which contains accelerated GPU and CPU implementations. The second step, decoding, interprets these scores and outputs the most probable nucleobase sequence. In Dorado, this step is only open-source for CPUs and closed-source on NVIDIA GPUs. The koi library, found precompiled in Dorado, contains accelerated NVIDIA GPU code for the entire decoding step and parts of the inference step.

Can we avoid closed-source koi?

To make Slorado a truly open-source and portable implementation, we would have to opt out of using the precompiled closed-source koi binaries. To assess the impact of koi on performance, we implemented slorado with and without the closed-source koi library.  The main difference in this open-source implementation is that instead of performing an optimised version of inference and decoding on the GPU, we use the standard torch version of inference on the GPU and then decode on the CPU.

We then measured the end-end basecalling times on a small subset of 20,000 reads with NVIDIA Tesla V100 GPUs, which are plotted in Fig. 2.

Fig 2: End-end basecalling times on NVIDIA Tesla V100 with v4.2.0 models for fast, hac, and sup. CPU decode (blue) indicates that the original torch code (GPU torch) was used for inference and the decoding was done on the CPU. Orange bars represent the time when koi is used for both inference and decoding.

Here we observe that basecalling without the GPU-accelerated koi library is extremely slow. For example, without koi, it is nearly 64 times slower than basecalling with the 4.2.0 model. This slowdown makes it unrealistic to perform basecalling on a practical amount of nanopore data (which can be millions of reads) without NVIDIA GPUs that can run koi.

Execution Breakdown

To identify the optimisations lost without the koi library, we profiled the end-end basecalling time for the version of Slorado without koi (the one represented by blue bars in Fig. 2),  to see where all our time is being spent. The breakdown of the execution time is plotted in Fig. 3.

Fig 3: Breakdown of the execution time for basecalling in slorado when inference is performed using torchlib (GPU) and the decoding is performed on CPU. Performed on a system with NVIDIA V100 GPUs.

As you can see, most of the time is spent copying the output tensor of the inference step from the device (GPU) to the host (CPU) before decoding (orange bars in Fig. 3). Therefore, although we can perform inference without koi’s inference optimisations, and decoding is quite fast on the CPU, transferring this data becomes impractical for anything other than a small dataset.

Since hosting inference on the CPU side (torch CPU version) would take ages even for a tiny dataset, our only option was to develop an open-source decoder for the GPU side. Decoding on the GPU would mean copying only the resulting DNA sequence and quality scores determined by the decoder. For orders of magnitude smaller than the output tensor, copying the decoding result would mean we are no longer limited by the throughput of the PCI express bus that connects the GPU with CPU.

Openfish

Openfish is our open-source implementation of the decoder tailored toward nanopore signal data.

The decoder’s job is to determine the most likely nucleobase sequence for our inference output. The decoding phase consists of several distinct steps. The first part of decoding involves calculating the posterior probabilities and back guide from our output tensor. This is an embarrassingly parallelisable problem, very suitable for the GPU. Here, we could map each batch produced by our neural network model to a single threadblock in the GPU. Each thread in our block calculates the probability for a single state at each timestep. We saw a massive speed up for this step when executed on GPU compared to CPU. After this, most of the time is spent on the beam search. Beam search is the part that searches the probabilities for the best nucleobase sequence. Since this search is mostly linear, it is not ideal for the GPU. However, a batch still requires hundreds of searches per chunk basis, so not being limited to the number of threads on a CPU still gives a speed boost.

On top of these speed-ups, avoiding that significant overhead to copy tensors from GPU to CPU, makes Slorado practically usable now with Openfish integrated. The results comparing the version of slorado with openfish to slorado without openfish or koi, are in Fig. 4.

Fig 4: End-end performance for slorado basecalling. CPU decode (blue) indicates torch inference and an optimised version of the CPU decoder. GPU (openfish) Decode (orange) refers to the same torch inference but with the GPU-optimised openfish decoder.

Read news story about the development of Slorado: https://pawsey.org.au/pawsey-enables-more-flexible-and-scalable-dna-analysis/

Back to top
  • Home
  • About Us
  • Science Showcase
  • News & Events
  • Support
  • User Portal
  • Contact
The Pawsey Supercomputing Centre is supported by $90 million funding as part of the Australian Government's measures to support national research infrastructure under the National Collaborative Research Infrastructure Strategy and related programs through the Department of Education.
Pawsey Computing Centre logo
Copyright Pawsey Computing Centre 2026
Website by harmonic