For the complete documentation index, see llms.txt. This page is also available as Markdown.

πŸ—“οΈThe Real Time Assembly Pipeline

Always on pipeline generating public data sets for key pathogens

About

The Centre for Genomic Pathogen Surveillance runs a genome assembly pipeline providing daily updates for the WHO priority pathogens, along with a few other key species. The aim is provide real time support for local outbreak detection and characterisation within the global context, encouraging and maximising the value of rapid release of genomics data in the public archives. The genomes produced by this pipeline is automatically imported in Pathogenwatch and made freely available for community use.

Viewing the data

Species supported by the "Always on" pipeline are marked in the Genomes Browser drop down menus with a check mark on the right hand side. Within a specific species browser, the public genomes can be viewed by selecting the "Always on" project within the Project filter on the left hand side.

Species that are updated daily are indicated with checkmarks in the drop down menus
Within a species the Project filter can be used to just select the public data

The assembly pipeline

Implementation

See "Short Read Assembly" for details on the short read pipeline.

AMR.watch uses the the assemblies and results produced by Pathogenwatch. To see a summaries of the number of genomes imported, and the impact of the different filters, visit https://amr.watch/summary and https://amr.watch/summary/all.

Selection Constraints

The pipeline is currently restricted to short read assembly, and is focused on genomic epidemiology rather than complete coverage of species diversity. A genome will be selected for assembly provided it:

  1. Has been assigned to one of priority species

  2. Meets assembly requirements:

    1. Illumina paired-end whole genome sequence

    2. Two FASTQs present

    3. >20x mean coverage according to the base count

    4. Available for download from the SRA

  3. Meets minimum metadata requirements

    1. Has at least the year as sample date

    2. The location can be resolved to a country

    3. Associated with a single sample accession

  4. Is the only representative for that sample

    1. The sample doesn't have a representative already

    2. If there is more than one possibility, it is the FASTQ pair with the most coverage

Citations

[1] - Prjibelski, A., Antipov, D., Meleshko, D., Lapidus, A., & Korobeynikov, A. (2020). Using SPAdes de novo assembler. Current Protocols in Bioinformatics, 70, e102. doi: 10.1002/cpbi.102

Last updated