Design Decisions¶
PoolSeqFlow departs from stock Nextflow practice in several places. Each departure was made for a reason, and each one costs something. This page states both, so you can tell whether a behavior you are seeing is a bug or the design working as intended.
Configuration is a file, never a flag¶
The decision. Every setting lives in parameters.config. The wrapper takes a subcommand, not settings. There is no --poolSize 50.
Why. Two reasons, one about reproducibility and one about correctness.
A run described entirely by a file is a run you can version, diff, publish alongside a paper, and hand to someone else with a guarantee they will get the same numbers. Once any setting can come from the command line, the file stops being a complete record and the shell history becomes part of the method.
The correctness reason is sharper. Nextflow passes --param values as strings. So this:
sets annotate to the string "false", and Groovy evaluates any non-empty string as true. Annotation stays switched on, no warning is printed, and the only symptom is that step 8 runs when you asked it not to. Written in the config file:
it is a real boolean and behaves as expected. This class of failure is silent and type-dependent, and the only reliable fix is to remove the path that creates it.
What it costs. Sweeping a parameter across values means editing a file between runs rather than scripting a loop over flags. For a parameter sweep, copy the project directory or keep several config files and swap them into place.
Resume is filesystem-based¶
The decision. Every step checks whether its own outputs already exist, and skips itself if they do. It looks in permanent storage first and then on the working volume, so an output that has been produced but not yet moved to storageDir counts as done. This replaces Nextflow's -resume entirely; the wrapper never passes that flag. PoolSeqFlow run is both "start" and "resume".
Why. Nextflow's cache lives in work/. It is invalidated by anything that removes or changes those directories, which for this pipeline is routine:
cleanup = trueinnextflow.configdeletes task working directories once a run completes.-resumereplays task outputs from those directories, so after a successful run there is nothing left to replay.- Several steps delete their own inputs once the next stage has consumed them — the trimmed reads are removed after clipping, and each VCF is removed after the next filter produces its successor. That leaves the upstream task's recorded outputs dangling, which invalidates the cache entry regardless.
- On HPC, jobs hit walltime, nodes reboot, and scratch is purged on a schedule. A cache that lives in scratch does not survive the failure modes that actually interrupt long runs.
A check for "does this output file already exist" survives all of that, because it depends on nothing but the storage the results are already in.
What it costs. Two things worth knowing.
Step-skipping happens inside each task rather than before it, so a re-run still submits every process to the scheduler. Those jobs exit almost immediately — they test for a file, create a symlink and copy two log files — but they are real submissions. Expect roughly one short job per process per sample on a fully resumed run.
More importantly, "the output exists" is not the same as "the output is correct for your current settings". A file produced under poolSize = 50 looks identical to one produced under poolSize = 100. That gap is closed separately, by the guardrails below.
Symbolic links instead of copies¶
The decision. Each output is moved out of the task's working directory to where it belongs, and a symbolic link is left behind pointing at it. Nothing is copied, and no file the pipeline produces exists twice.
Why. Pool-seq intermediates are large. A run with a dozen pools moves through hundreds of gigabytes of BAMs and VCFs. Nextflow's default is to publish outputs by copying them out of work/, which means every large file exists twice for as long as work/ survives.
An output is not moved to its final home immediately, though, and that is the second half of the design. It first lands on mainDir, the working volume, and moves to storageDir only once the last step that needed it has finished. A BAM is read several times on its way to a frequency table, and reading it repeatedly across a network mount is the slowest thing a run does.
The two directories therefore have distinct jobs, and cannot be the same path:
mainDiris the fast one. Your reads, your reference and everything in progress live here, and it is where the work happens.storageDiris the durable one, and holds what you keep.
On a cluster this maps onto a node's local disk and the network volume behind it; on a laptop it is one directory that churns and one you back up. Either way an output crosses between them exactly once, at the point where it stops being working material and becomes a result.
What it costs. Symbolic links, which is why Windows is not supported — including WSL under some filesystem configurations. It also means a task's working directory is not self-contained: deleting either directory while a run is in flight breaks links that are already in use. And mainDir is not scratch — it holds your inputs and your configuration, so it has to survive between runs.
Threads are budgeted, not divided¶
The decision. One threads value sizes the whole run. Every tool's core count is derived from it through a fixed ladder, and each process reserves what it actually uses.
Why. The obvious alternative — divide the available cores evenly among concurrent tasks — assumes tools scale linearly with threads. They do not. Each tool here is quantized to the point where its published scaling flattens out, so extra cores go to another task instead of into diminishing returns.
The sharper reason is that a tool's advertised thread count is not always what it spawns. Trim Galore's --cores N runs N+4 threads: N workers, two decompressors, a batcher and a writer. A process that declares cpus 4 and then passes --cores 4 is really using eight. Nextflow decides how many tasks to run concurrently by comparing cpus against available resources, so an under-declared task causes oversubscription — the machine ends up running twice the work it thinks it is.
PoolSeqFlow reserves Trim Galore's full footprint and maps back to the worker count in the script, so the declaration and the reality agree. The full ladder is in Resources.
What it costs. Honest accounting is slower than optimistic accounting on a small machine. At threads = 8, a single trimming task reserves all eight cores, so samples are trimmed one at a time. Earlier behavior ran three concurrently at twelve threads each on an eight-core box — faster in wall-clock, and a 4.5× oversubscription. A request larger than the machine now fails immediately rather than quietly degrading:
The run refuses to mix settings¶
The decision. The run stops when the pipeline version, the analysis parameters, the run table or metadata.csv have changed since the existing outputs were produced.
Why. This is the direct consequence of filesystem-based resume. Because a step skips itself when its output file exists, and the file carries no record of what produced it, changing poolSize and re-running would leave one Frequencies/ folder holding tables computed under two different thresholds. Nothing downstream could detect that, and the mixture would be invisible in the output.
So a record of what produced them is kept beside the results:
| Record | Covers |
|---|---|
.poolseqflow_version |
The release that produced these results. A mismatch is a hard stop, not a warning — see A project belongs to one release |
.poolseqflow_params |
The analysis-affecting parameters, mirrored to a readable run_parameters.txt beside it |
.parameters.config |
Your configuration file, copied verbatim as you wrote it |
.multirun.csv |
Your run table, copied verbatim, if you used one |
.poolseqflow_metadata |
The parts of metadata.csv that can change a result — the read-group tags, the pool sizes, and the row order that sets column order. Kept beside the results it describes |
That last one is also why most of your own columns cost nothing to change. A column recording experimental design, or a result from somewhere else — how the samples were grouped, when they were collected, what was measured about them — is not in the record, so adding or editing one never invalidates results you already have. That holds for now. Once the analysis layer arrives, the columns it reads will start to matter, and that freedom will narrow.
Path and resource parameters are excluded, since they change where and how fast the work happens rather than what the answer is; so are the software entries that name the commands to run. Anything added in a later release counts as analysis-affecting until decided otherwise, which is the conservative direction to err in.
The two verbatim copies serve a different purpose from the rest. They are not what the comparison is made against — they are the citable record, the exact files you ran kept next to the numbers they produced.
What it costs. You cannot change a threshold and re-run to see the difference in place. The report names the folders to delete, and deleting them is what clears it — that is deliberate, because the alternative is a folder of results you can no longer attribute to a setting.
Moves across filesystems are atomic¶
The decision. All cross-filesystem moves stage inside a temporary directory of the mover's own and rename into place, via bin/atomic_mv.sh.
Why. A plain mv across a filesystem boundary is a copy followed by an unlink, not an atomic rename. A job killed mid-move — walltime, preemption, a node failure — leaves a truncated file under its final name. Combined with existence-based resume, that is the worst possible failure: the next run sees the file, concludes the step is done, and builds everything downstream on a partial BAM.
Staging under a temporary name and renaming means an interrupted move leaves nothing any existence check looks for, and the step simply runs again. The staging directory belongs to the one move: two jobs racing for the same destination stage separately, so neither can be seen writing into the other's copy.
Steps delete their own inputs¶
The decision. Once a stage's output is safely in permanent storage, several steps delete the input they consumed — trimmed reads after clipping, each VCF after the next filter produces its successor.
Why. Peak disk usage on a Pool-seq run is dominated by intermediates that nobody needs once the next stage has run. Keeping every one of them would roughly multiply the storage requirement by the number of filter stages, for files that exist only to be consumed.
What it costs. You cannot inspect an intermediate after the fact without re-running from an earlier point, and it is part of why Nextflow's own cache cannot be used.
Note that this applies to the permanent copies too, not just the scratch ones: the step deletes the file the symlink resolves to. After a complete run, Output/VCF/ holds the raw call set, the fully filtered VCF, and the annotated VCF if you enabled it — the per-stage intermediates between them are gone. Exactly which files survive is listed in Directory Layout.