Configuration¶
Everything the pipeline does is set in parameters.config. There are no command-line overrides (why). The optional analysis layer has settings of its own, in a file of its own — see Analysis Layer.
This page sorts the parameters by what they actually affect, which is the distinction that matters most: some change your numbers, some change only where files land or how fast the run goes, and some are computed for you and should not be edited at all.
Parameters, by what changing one does¶
Parameters that change your results¶
Change one of these and your output changes. Step 0 records them and refuses to run if they differ from what produced your existing outputs, so that one folder never holds results from two settings.
| Parameter | Effect | Page |
|---|---|---|
poolSize |
Individuals per pool; sets the minimum credible allele frequency. Can be set per pool in metadata.csv |
Filtering & Frequency |
ploidy |
Ploidy; same threshold. Can be set per run | Filtering & Frequency |
filterFalsePositives.sampleThreshold |
Fraction of samples that must support an allele | Filtering & Frequency |
capBAM.maxDepth |
The depth ceiling put on each BAM. Can be set per sample in metadata.csv |
Variant Calling |
variantCall.* |
Pileup and calling behavior, including a flat depth cap on top of the measured one | Variant Calling |
vcffilter.minDP, vcffilter.minQUAL |
Post-call depth and quality filtering | Filtering & Frequency |
bwa.minScoreOutput, bwa.batchSize, bwa.options |
How reads are aligned in the first place, and whether that is reproducible across machines | Step 3 |
cleanBAM.filter, cleanBAM.required, cleanBAM.mapq |
Which alignments reach the pileup | Alignment & Cleaning |
cutadapt.at_gc_error |
Composition tolerance driving the clip points | Trimming & Clipping |
trim_galore.quality, .autodetect, .adapter1/2 |
What is trimmed off the reads | Trimming & Clipping |
annotate, gffFile, snpEff.* |
Whether step 8 runs, against what, and with which SnpEff options | Annotations |
metadata.csv |
Which FASTQ pairs are one pool, each pool's size and depth ceiling, and column order — the RG_* and param_* columns only |
Metadata |
| The run table | Whatever it varies, per run. Every column in it is a parameter | Multi-run |
Parameters that change speed, not answers¶
Safe to tune between runs. Step 0 does not track them, precisely because they cannot change a result.
| Parameter | Effect | Page |
|---|---|---|
threads |
Cores a single task may use; drives every tool's thread count | Resources |
memory |
Memory ceiling for a single task | Resources |
java.heapSize |
JVM heap for FastQC and SnpEff | Resources |
fastqc.memory |
FastQC's own memory setting, in megabytes | Resources |
software.* |
Paths to executables, if not using the conda environment | below |
capBAM.histogramMax |
How deep step 5's depth histogram looks | Variant Calling |
capBAM.histogramMax is the odd one here: it changes nothing about speed. It is in this group because it bounds what step 5 looks at rather than what it decides — a run whose histogram would be truncated stops instead of choosing, so every value a run completes at gives the same ceiling. That is what makes it safe to raise on a project that already holds results.
Parameters that change where files go¶
| Parameter | Effect |
|---|---|
mainDir |
Working directory — your inputs, work/, and everything in progress. Where you run from |
storageDir |
Permanent storage — the finished results. Must be a different path from mainDir |
dataSource |
Subdirectory of mainDir holding the FASTQs |
readPattern |
Glob matching paired FASTQs; needs a {1,2} group |
referenceFile, gffFile |
Input filenames within mainDir/Reference |
metadataFile |
Name of the sample table, in mainDir |
multiRun, multiRunFile |
Whether to read a run table, and what it is called |
vcf.fileName |
Base name for the VCFs and frequency tables. See Filtering & Frequency |
dryRunDir |
Where dryrun builds its preview and where dryclean looks for one. See Cleaning up |
Parameters that change a report, not a result¶
One parameter changes only what a published report contains, and nothing that flows onward:
| Parameter | Effect | Page |
|---|---|---|
fastqc.options |
What FastQC reports on the clipped reads | Trimming & Clipping |
Step 0 does not refuse a change to these, because your results do not depend on them. It does refuse two runs of a table that share a step and disagree about one — only one of the values can have produced the single report sitting in the shared directory, so the run stops rather than publishing an ambiguous file.
Do not edit: derived values¶
A large part of parameters.config is computed. The cores block derives every tool's thread count from threads; the dir block builds every path from mainDir and storageDir; filterFalsePositives.sensitivity is computed from poolSize and ploidy; snpEff.db is derived from gffFile.
Beyond those, eight named values are assembled from the filenames and roots you set. They appear in the config as ordinary assignments, so they can be overridden — but each has an input that is the thing you actually mean to change:
| Parameter | What it is | Set instead |
|---|---|---|
referenceFa |
Your reference filename with .gz removed — the decompressed name step 1 works with |
referenceFile |
referencePath |
Full path to the reference, under mainDir/Reference |
referenceFile, dir.references |
gffPath |
Full path to the annotation file, alongside it | gffFile |
metadataPath |
Full path to the sample table, in mainDir |
metadataFile |
multiRunPath |
Full path to the run table, in mainDir |
multiRunFile |
dir.allOutputs |
Results root for things belonging to the whole invocation rather than one run. storageDir/Output, or storageDir/Output/All_Runs when multiRun is on |
storageDir, multiRun |
dir.allLogs |
The same for logs: storageDir/Logs, or storageDir/Logs/All_Runs |
storageDir, multiRun |
dir.sessionReports |
Where the four Nextflow session reports go — PoolSeqFlow_pipeline_report.html, _timeline.html, _trace.txt and _dag.html, named in nextflow.config |
storageDir, dir.subpath.reports |
The three All_Runs values are why a multi-run project does not scatter session-level output through the individual run directories: a report describing the whole invocation has one home, and it is chosen by multiRun rather than by each step guessing.
Editing these by hand breaks the invariant that makes the pipeline predictable — that one number sizes the run, and one pair of paths places everything. Change the input, not the derivation.
Setting one is supported rather than forbidden, and how you do it depends on which: the cores block and the options strings ship commented out, so uncommenting a line is what pins it, while filterFalsePositives.sensitivity, referenceFa and snpEff.db are written out as formulas, so pinning one means replacing the formula with a value. Either way a derived value you set by hand is used exactly as written, and nothing is derived from its inputs any more. Pin variantCall.mpileupOptions and variantCall.maxDepth stops meaning anything for that run. That is a reasonable thing to want when you need full control of a command line — it is only a trap when it happens by accident.
The paths are the exception, and a value written over one of them is replaced rather than used. referencePath, gffPath, metadataPath, multiRunPath and the dir block are rebuilt for every run out of the roots and filenames that run holds — under a run table dir.outputs and dir.logs carry the run's own name, and a pin there would send two runs to one directory. Change the root or the filename instead, which is what the table above names for each.
Where the pipeline can tell, it says so. The verification step at the beginning of a run reports when your trimming options have been pinned rather than derived, and if you are using a run table it names any column that sets a computed value directly.
Where a parameter can be set¶
The tables above sort parameters by what they affect. The other axis is how widely a value applies, and it has three levels:
| Level | Where you write it | Applies to |
|---|---|---|
| Global | parameters.config |
Every sample of every run |
| Per run | a column in the run table | Every sample of one run — Multi-run |
| Per sample | a param_* column in metadata.csv |
The rows that carry a value |
Any parameter can be set globally or per run. The run table takes any parameter name as a column, spelled as parameters.config spells it.
Per-sample is a closed list, because each one needs code that knows to look for it:
| Column | Overrides | Read at |
|---|---|---|
param_poolSize |
poolSize |
step 7 |
param_capMaxDepth |
capBAM.maxDepth |
step 5 |
param_adapter1, param_adapter2 |
trim_galore.adapter1/2 |
step 2 |
A param_ column that is not on this list is refused rather than ignored — a name the pipeline cannot act on is a typo, not a preference, and silently recording it would leave you with a setting you can see in your own file and that never took effect.
What a per-sample value costs you in a run table¶
The step a per-sample parameter is read at decides how much two runs can still share. Runs sharing a value share the work up to the step that first reads it, and diverge from there on. So the same kind of override is cheap in one place and expensive in another:
| Two runs differ in… | They still share | They repeat |
|---|---|---|
param_poolSize |
trimming, alignment, BAM cleanup, reports, calling | the frequency tables only |
param_capMaxDepth |
trimming, alignment, BAM cleanup | reports, capping, calling, and everything after |
param_adapter1/2 |
nothing | the whole analysis |
Two runs can only differ in a param_* value by using different metadata files — metadataFile is itself a parameter, so it can be a column in the run table. Step 0 prints the resulting split before any compute is spent, so check there rather than inferring it: it names each results directory, which runs share it, and which steps it holds.
What to decide before your first run¶
In rough order of how expensive it is to get wrong:
metadata.csv— which FASTQ pairs share anRG_Sample. Wrong here means valid results that answer a different question, and fixing it invalidates every BAM. →poolSizeandploidy— these set the frequency floor.poolSizecan be given per pool inmetadata.csv, andploidyapplies to a whole run. →filterFalsePositives.sampleThreshold— decides whether alleles seen in few pools survive. The default removes them. →capBAM.maxDepth— leave it at-1unless you know your libraries need otherwise. It is the one item here you can safely decide after the first run, because the depth reports tell you what it did. →threads— must fit the machine, or the run fails at submission. →
Changing any of items 1–4 after outputs exist means deleting those outputs. That is enforced, not advisory.
Using system tools¶
The software block maps each tool to a command:
Replacing a command with an absolute path makes the pipeline use a system installation instead of the conda environment. This is supported but not recommended: the environment pins exact builds because Pool-seq results depend on the precise behavior of the pileup and filtering tools, and a version mismatch will not announce itself. Use it to work around a genuine packaging problem, not as a default.
Three of them cannot be repointed. rsync, diff and find are checked at the start of a run along with everything else, so a missing one is reported before any work begins — but a path you give for them is not used. They are called by name from the small scripts that move a finished file into permanent storage, and those scripts read none of your settings. If you need a different one, change what is on your PATH.