ποΈThe Real Time Assembly Pipeline
Always on pipeline generating public data sets for key pathogens
Last updated
Always on pipeline generating public data sets for key pathogens
The Centre for Genomic Pathogen Surveillance runs a genome assembly pipeline providing daily updates for the WHO priority pathogens, along with a few other key species. The aim is provide real time support for local outbreak detection and characterisation within the global context, encouraging and maximising the value of rapid release of genomics data in the public archives. The genomes produced by this pipeline is automatically imported in Pathogenwatch and made freely available for community use.
Species supported by the "Always on" pipeline are marked in the Genomes Browser drop down menus with a check mark on the right hand side. Within a specific species browser, the public genomes can be viewed by selecting the "Always on" project within the Project filter on the left hand side.


See "Short Read Assembly" for details on the short read pipeline.
AMR.watch uses the the assemblies and results produced by Pathogenwatch. To see a summaries of the number of genomes imported, and the impact of the different filters, visit https://amr.watch/summary and https://amr.watch/summary/all.
The pipeline is currently restricted to short read assembly, and is focused on genomic epidemiology rather than complete coverage of species diversity. A genome will be selected for assembly provided it:
Has been assigned to one of priority species
Meets assembly requirements:
Illumina paired-end whole genome sequence
Two FASTQs present
>20x mean coverage according to the base count
Available for download from the SRA
Meets minimum metadata requirements
Has at least the year as sample date
The location can be resolved to a country
Associated with a single sample accession
Is the only representative for that sample
The sample doesn't have a representative already
If there is more than one possibility, it is the FASTQ pair with the most coverage
Genomes with updated metadata are not automatically detected and imported, but will be added on a more manual basis. If there are genomes you know have been updated with time or location data and now meets requirements, please let us know and we will add them.
Last updated