Skip to content

Resume Logic

PoolSeqFlow implements its own resume strategy rather than using Nextflow's. This page covers how it behaves; the reasoning behind replacing -resume is in Design Decisions.

Two directories

mainDir Where the pipeline runs and where everything it works on lives — your reads, your reference, the dictionaries built from it, work/, and outputs still being read. Not scratch: it holds your inputs, so it has to survive between runs
storageDir Where finished results are kept. Network storage, a group volume, a different mount

They cannot be the same path, and the run stops if they are. The two are storage tiers with different jobs, and an output moving from one to the other is the event that marks it finished — which cannot mean anything if they are one place.

Each output is moved out of work/ and a symlink left behind. Consequences:

  • No duplication. Nothing the pipeline produces exists twice on disk.
  • One crossing. An artifact moves between the two volumes exactly once, when the last step that needed it has finished.
  • Automatic step-skipping. Every step looks for its own outputs — in storageDir first, then on the working volume — and skips itself if they are there, regardless of the state of work/.
  • Resilience. The check depends on nothing but the storage the results are already in, so it survives cluster timeouts, reboots and work/ cleanups.

The order of that search matters. Permanent storage is consulted first so that a stray copy left on the working volume can never outrank a promoted one.

There is no -resume

This strategy replaces Nextflow's -resume, and the wrapper never passes that flag. Two reasons it could not work here even if it were passed:

  • cleanup = true deletes task working directories once a run completes. -resume replays task outputs from those directories; after a successful run there is nothing to replay.
  • Several steps delete their own inputs once consumed. That leaves the upstream task's recorded outputs dangling, which invalidates the cache entry regardless.

So PoolSeqFlow run is both "start" and "resume". PoolSeqFlow resume survives as a deprecated alias and prints a notice.

To start genuinely from scratch:

PoolSeqFlow reset

It lists exactly what it is about to remove, across both directories, and requires typing DELETE_MY_ANALYSIS to confirm. That covers Output/ and Logs/ in storageDir; on mainDir, the dictionaries built from your reference, anything not yet promoted, and work/; Nextflow's own history; and the .parameters.config, .multirun.csv and .poolseqflow_* records describing all of it. Those records go too, because leaving them would have the next run comparing your configuration against outputs that no longer exist.

Your reads, your reference and your two configuration files are not touched.

What a resumed run looks like

Every process is still submitted. Step-skipping happens inside each task, not before it, so a fully resumed run submits roughly one job per process per sample. Those jobs test for a file, create a symlink, copy two log files and exit — but on a scheduler they are real submissions with real queue time.

The log lines to look for are the COMPLETED messages that follow a "Found existing" line:

ALIGNING Sample1: Found existing BAM file
ALIGNING Sample1: Found: /storage/project/Output/Aligned/Sample1.bam
ALIGNING Sample1: Creating symbolic link...
ALIGNING Sample1: COMPLETED

Partial-stage resume

Step 7 is a chain of five sub-steps, and each checks for the outputs of every later stage as well as its own. If the frequency tables already exist, the earlier sub-steps create an empty placeholder and exit rather than redoing work whose result was superseded. This is why a partially completed step 7 resumes correctly even though its intermediates have been deleted.

Interrupted moves

A plain mv across a filesystem boundary is a copy followed by an unlink. A job killed mid-move would leave a truncated file under its final name, which existence-based resume would then accept as a completed step.

All cross-filesystem moves go through bin/atomic_mv.sh, which stages inside a temporary directory of its own and renames into place. An interrupted move leaves nothing any check looks for, and the step simply runs again.

What resume does not protect you from

"The output exists" is not "the output is correct for your current settings". A file produced under poolSize = 50 is indistinguishable from one produced under poolSize = 100.

That gap is closed by the consistency guards at the start of a run, which record the release, the analysis parameters, the run table and the analysis-affecting parts of metadata.csv behind a set of outputs, and stop the run when any of them has changed. See Step 0.

Cleaning up

Command Removes
PoolSeqFlow clean Nextflow work directories — the empty hash-prefix folders cleanup = true leaves behind — and any staging directory left by a killed transfer
PoolSeqFlow dryclean The empty directory tree dryrun created as a preview
PoolSeqFlow reset All progress, across both directories, after listing it and asking you to type a confirmation

clean is safe at any time and does not affect resume — nothing in work/ is consulted by the skip logic. reset deletes results.

Every file the pipeline and the analysis layer publish is written into a staging directory beside its destination, verified there, and only then renamed into place, so that a partly-written file is never visible under a final name. That staging directory is removed on every path the transfer can reach, including failure — but a process that is killed reaches none of them, and leaves one behind. clean collects them from both mainDir and storageDir, naming each one before it goes. They are named .atomic_mv.*, .restore.* and .analysis_results.*, and finding one is a sign a run was killed rather than allowed to stop.

The preview itself is built in dryRunDir, which is dryrun/ in the directory you launch from, beside parameters.config, unless you point it somewhere else. It is outside the Output/ tree on both roots, so a preview is never mixed in among results. dryclean asks the same parameter where to look, and both commands check what they are about to delete first: a directory holding anything other than empty folders, a README.txt and a members.txt is not a preview, so they list what is in it and remove nothing. dryrun replaces a preview that passes that check without asking, because there is nothing in one to lose.

dryRun is the other half, and it is the pipeline's own flag rather than a setting for you to make: the dryrun command sets it, and it is what stops step 0 writing. Every check still runs and every comparison is still made — the report says would record where a real run says recording. That is what keeps a preview from leaving a baseline behind for results that are never produced, and it is why dryRun and dryRunDir are among the handful of parameters kept out of the recorded manifest: a preview must not be able to look like a run.

Do not delete either directory's contents while a run is in flight

Task working directories contain symlinks into the volume the output was moved to. Removing the target breaks links that are actively in use, and the failure will not be obvious. That applies to the working volume as well as to permanent storage — an artifact waiting to be promoted is being read from where it is.