English

A Targeted Reference Database for Improved Analysis of Environmental 16S rRNA Oxford Nanopore Sequencing Data

Molecular Ecology Resources ()

https://doi.org/10.1111/1755-0998.70036

Tilgang krever betaling

1 Akvaplan-niva (nåværende ansatt)

1 Akvaplan-niva (tidligere ansatt)

Forfattere (9)
  1. Melcy Philip
  2. Tonje Nilsen
  3. Sanna Kristiina Majaneva
  4. Ragnhild Pettersen
  5. Morten Stokkan
  6. Jessica Louise Ray
  7. Nigel Keeley
  8. Knut Rudi
  9. Lars‐Gustav Snipen

Abstract

ABSTRACT The Oxford Nanopore Technologies (ONT) sequencing platform is compact and efficient, making it suitable for rapid biodiversity assessments in remote areas. Despite its long reads, ONT has a higher error rate compared to other platforms; necessitating high‐quality reference databases for accurate taxonomic assignments. However, the absence of targeted databases for underexplored habitats, such as the seafloor, limits ONT's broader applicability for exploratory analysis. To address this, we propose an approach for building environmentally targeted databases to improve 16S rRNA gene (16S) analysis using Oxford Nanopore Technologies (ONT), using seafloor sediment samples from the Norwegian coast as an example. We started by using Illumina short‐read data to create a database of full‐length or near full‐length 16S sequences from seafloor samples. Initially, amplicons are mapped to the SILVA database, with matches added to our database. Unmatched amplicons are reconstructed using METASEED and Barrnap methodologies with amplicon and metagenome data. Finally, if the previous strategies did not succeed, we included the short‐read sequences in the database. This resulted in AQUAeD‐DB, which contains 14,545 16S sequences clustered at 95% identity. Comparative database analysis reveals that AQUAeD‐DB provides consistent results for both Illumina and Nanopore read assignments (median correlation coefficient: 0.50), whereas a standard database showed a substantially weaker correlation. These findings also emphasise its potential to recognise both high and low abundance taxa, which could be key indicators in environmental studies. This work highlights the necessity of targeted databases for environmental analysis, especially for ONT‐based studies, and lays the foundations for future extension of the database.

Fra , siste endring

Registrert i Nasjonalt vitenarkiv