Skip to content

Pipeline Overview

PoolSeqFlow is a set of Nextflow DSL2 modules, each an independent file under scripts/ and each responsible for its own resume logic. Steps 0 to 8 are the analysis; a tenth module handles promotion, moving each finished artifact to permanent storage once nothing needs it any more. This page is the map; the detail is in Steps.

Raw FASTQ reads
      │
      ▼
[Step 0] Verify environment, parameters and folder structure
      │
      ▼
[Step 1] Build reference dictionaries (BWA, SAMtools, SnpEff)
      │
      ▼
[Step 2] QC & trimming (FastQC → Trim Galore → composition-aware clipping)
      │
      ▼
[Step 3] Alignment (BWA-MEM)
      │
      ▼
[Step 4] BAM cleanup (name-sort → fixmate → coord-sort → markdup → addRG → filter → index)
      │
      ▼
[Step 5] Reports, and the depth ceiling for each sample
      │
      ▼
[Step 6] Cap each BAM, then variant calling (BCFtools mpileup + call)
      │
      ├────────────────────────────────┐
      ▼                                ▼
[Step 7] VCF → frequency tables   [Step 8] Annotation (optional)

What runs per sample and what runs once

This distinction explains most of the pipeline's runtime behavior.

Step Granularity Notes
0 Verify Once Gate for everything else
1 Dictionaries Once Three parallel index builds
2 Trim & clip Per sample Two sub-steps, the second depends on the first's FastQC output
3 Align Per sample
4 Cleanup Per sample
5 Reports Per sample Three sub-steps, independent; one of them sets that sample's depth ceiling
6 Variant call Capping per sample, then once, jointly Capping runs only for samples that have a ceiling; then one bcftools mpileup over every BAM
7 Frequencies Once, five sub-steps Serial chain; the last runs twice, on SNPs and on INDELs
8 Annotation Once Optional; runs on step 6's output, not step 7's

Step 6 is the pipeline's barrier: it needs every BAM before it can start, so a single slow sample delays the whole run from that point. It is also why samples must share a reference, and why the sample column order needs deciding (in metadata.csv).

Under a run table, "once" means once per variant rather than once overall: runs that agree on everything a step reads share that step's single task, and only diverging work is repeated. See Multi-run.

The two output branches

After variant calling the workflow forks, and the branches never rejoin:

Step 7 takes the raw VCF through major-allele normalization, the cross-sample false-positive filter, depth and quality filtering, a SNP/INDEL split, and conversion to frequency tables. This is the analytical path.

Step 8 takes the same raw VCF — not step 7's output — splits multiallelic sites, and annotates with SnpEff. So the annotated VCF contains sites step 7 removed, in the original reference encoding rather than the major-allele one. Joining the two requires matching on CHROM/POS and expecting unmatched rows (details).

Where the work happens

Every step follows the same pattern:

  1. Check whether its output already exists — in storageDir, or on the working volume waiting to be promoted. If so, symlink it and exit.
  2. Otherwise do the work in mainDir/work/.
  3. Move the result out atomically: to mainDir/Utilized/ if a later step will read it, or straight to storageDir if nothing will.
  4. Symlink it back into the task's working directory.
  5. Copy .command.log and .command.err into Logs/.

Then, separately: when the last step that needed an artifact has succeeded, it is moved from Utilized/ to storageDir. Each byte crosses between the two volumes exactly once, and something enters Utilized/ precisely when it will be read again — the unpaired reads, the trimming reports and the FastQC output have no consumer at all, so they go straight to permanent storage and never appear there.

That pattern is what makes the pipeline resumable without Nextflow's cache, keeps files from being duplicated, and survives an interrupted move. The reasoning is in Design Decisions; the mechanics are in Resume Logic.

Error handling

nextflow.config sets errorStrategy = 'finish': on a failure, running tasks are allowed to complete and no new ones start. Nothing is left half-written, and a re-run picks up from whatever genuinely finished.

ClipReads is the exception, with errorStrategy 'retry' and maxRetries 3 — it is the one step whose failures are commonly transient.

cleanup = true removes task working directories after a successful run, leaving only empty hash-prefix folders under work/. PoolSeqFlow clean clears those.

These five are yours to change, from parameters.config — errorStrategy, maxRetries, maxErrors, cleanup and conda.enabled. They are defaults the installation sets before it reads your config, so anything you write wins. Nextflow scopes go outside the params { } block:

params {
    // ... your settings ...
}

cleanup = false          // keep work/ for debugging
process {
    errorStrategy = 'retry'
    maxRetries = 5
}

Two settings in nextflow.config are not yours to change, and a value you write for either is ignored. workDir is mainDir/work, which is where clean and reset look for it — move mainDir if you need the work directory elsewhere. env.PATH is what puts dir.bin in front of every task, and it is how the process scripts find the helpers they call by bare name; both of those follow the installation, and a run that lost them would fail at its first atomic_mv.sh.