Skip to content

Configuration

Everything the pipeline does is set in parameters.config. There are no command-line overrides (why). The optional analysis layer has settings of its own, in a file of its own — see Analysis Layer.

This page sorts the parameters by what they actually affect, which is the distinction that matters most: some change your numbers, some change only where files land or how fast the run goes, and some are computed for you and should not be edited at all.

Parameters, by what changing one does

Parameters that change your results

Change one of these and your output changes. Step 0 records them and refuses to run if they differ from what produced your existing outputs, so that one folder never holds results from two settings.

Parameter Effect Page
poolSize Individuals per pool; sets the minimum credible allele frequency. Can be set per pool in metadata.csv Filtering & Frequency
ploidy Ploidy; same threshold. Can be set per run Filtering & Frequency
filterFalsePositives.sampleThreshold Fraction of samples that must support an allele Filtering & Frequency
capBAM.maxDepth The depth ceiling put on each BAM. Can be set per sample in metadata.csv Variant Calling
variantCall.* Pileup and calling behavior, including a flat depth cap on top of the measured one Variant Calling
vcffilter.minDP, vcffilter.minQUAL Post-call depth and quality filtering Filtering & Frequency
bwa.minScoreOutput, bwa.batchSize, bwa.options How reads are aligned in the first place, and whether that is reproducible across machines Step 3
cleanBAM.filter, cleanBAM.required, cleanBAM.mapq Which alignments reach the pileup Alignment & Cleaning
cutadapt.at_gc_error Composition tolerance driving the clip points Trimming & Clipping
trim_galore.quality, .autodetect, .adapter1/2 What is trimmed off the reads Trimming & Clipping
annotate, gffFile, snpEff.* Whether step 8 runs, against what, and with which SnpEff options Annotations
metadata.csv Which FASTQ pairs are one pool, each pool's size and depth ceiling, and column order — the RG_* and param_* columns only Metadata
The run table Whatever it varies, per run. Every column in it is a parameter Multi-run

Parameters that change speed, not answers

Safe to tune between runs. Step 0 does not track them, precisely because they cannot change a result.

Parameter Effect Page
threads Cores a single task may use; drives every tool's thread count Resources
memory Memory ceiling for a single task Resources
java.heapSize JVM heap for FastQC and SnpEff Resources
fastqc.memory FastQC's own memory setting, in megabytes Resources
software.* Paths to executables, if not using the conda environment below
capBAM.histogramMax How deep step 5's depth histogram looks Variant Calling

capBAM.histogramMax is the odd one here: it changes nothing about speed. It is in this group because it bounds what step 5 looks at rather than what it decides — a run whose histogram would be truncated stops instead of choosing, so every value a run completes at gives the same ceiling. That is what makes it safe to raise on a project that already holds results.

Parameters that change where files go

Parameter Effect
mainDir Working directory — your inputs, work/, and everything in progress. Where you run from
storageDir Permanent storage — the finished results. Must be a different path from mainDir
dataSource Subdirectory of mainDir holding the FASTQs
readPattern Glob matching paired FASTQs; needs a {1,2} group
referenceFile, gffFile Input filenames within mainDir/Reference
metadataFile Name of the sample table, in mainDir
multiRun, multiRunFile Whether to read a run table, and what it is called
vcf.fileName Base name for the VCFs and frequency tables. See Filtering & Frequency
dryRunDir Where dryrun builds its preview and where dryclean looks for one. See Cleaning up

Parameters that change a report, not a result

One parameter changes only what a published report contains, and nothing that flows onward:

Parameter Effect Page
fastqc.options What FastQC reports on the clipped reads Trimming & Clipping

Step 0 does not refuse a change to these, because your results do not depend on them. It does refuse two runs of a table that share a step and disagree about one — only one of the values can have produced the single report sitting in the shared directory, so the run stops rather than publishing an ambiguous file.

Do not edit: derived values

A large part of parameters.config is computed. The cores block derives every tool's thread count from threads; the dir block builds every path from mainDir and storageDir; filterFalsePositives.sensitivity is computed from poolSize and ploidy; snpEff.db is derived from gffFile.

Beyond those, eight named values are assembled from the filenames and roots you set. They appear in the config as ordinary assignments, so they can be overridden — but each has an input that is the thing you actually mean to change:

Parameter What it is Set instead
referenceFa Your reference filename with .gz removed — the decompressed name step 1 works with referenceFile
referencePath Full path to the reference, under mainDir/Reference referenceFile, dir.references
gffPath Full path to the annotation file, alongside it gffFile
metadataPath Full path to the sample table, in mainDir metadataFile
multiRunPath Full path to the run table, in mainDir multiRunFile
dir.allOutputs Results root for things belonging to the whole invocation rather than one run. storageDir/Output, or storageDir/Output/All_Runs when multiRun is on storageDir, multiRun
dir.allLogs The same for logs: storageDir/Logs, or storageDir/Logs/All_Runs storageDir, multiRun
dir.sessionReports Where the four Nextflow session reports go — PoolSeqFlow_pipeline_report.html, _timeline.html, _trace.txt and _dag.html, named in nextflow.config storageDir, dir.subpath.reports

The three All_Runs values are why a multi-run project does not scatter session-level output through the individual run directories: a report describing the whole invocation has one home, and it is chosen by multiRun rather than by each step guessing.

Editing these by hand breaks the invariant that makes the pipeline predictable — that one number sizes the run, and one pair of paths places everything. Change the input, not the derivation.

Setting one is supported rather than forbidden, and how you do it depends on which: the cores block and the options strings ship commented out, so uncommenting a line is what pins it, while filterFalsePositives.sensitivity, referenceFa and snpEff.db are written out as formulas, so pinning one means replacing the formula with a value. Either way a derived value you set by hand is used exactly as written, and nothing is derived from its inputs any more. Pin variantCall.mpileupOptions and variantCall.maxDepth stops meaning anything for that run. That is a reasonable thing to want when you need full control of a command line — it is only a trap when it happens by accident.

The paths are the exception, and a value written over one of them is replaced rather than used. referencePath, gffPath, metadataPath, multiRunPath and the dir block are rebuilt for every run out of the roots and filenames that run holds — under a run table dir.outputs and dir.logs carry the run's own name, and a pin there would send two runs to one directory. Change the root or the filename instead, which is what the table above names for each.

Where the pipeline can tell, it says so. The verification step at the beginning of a run reports when your trimming options have been pinned rather than derived, and if you are using a run table it names any column that sets a computed value directly.

Where a parameter can be set

The tables above sort parameters by what they affect. The other axis is how widely a value applies, and it has three levels:

Level Where you write it Applies to
Global parameters.config Every sample of every run
Per run a column in the run table Every sample of one run — Multi-run
Per sample a param_* column in metadata.csv The rows that carry a value

Any parameter can be set globally or per run. The run table takes any parameter name as a column, spelled as parameters.config spells it.

Per-sample is a closed list, because each one needs code that knows to look for it:

Column Overrides Read at
param_poolSize poolSize step 7
param_capMaxDepth capBAM.maxDepth step 5
param_adapter1, param_adapter2 trim_galore.adapter1/2 step 2

A param_ column that is not on this list is refused rather than ignored — a name the pipeline cannot act on is a typo, not a preference, and silently recording it would leave you with a setting you can see in your own file and that never took effect.

What a per-sample value costs you in a run table

The step a per-sample parameter is read at decides how much two runs can still share. Runs sharing a value share the work up to the step that first reads it, and diverge from there on. So the same kind of override is cheap in one place and expensive in another:

Two runs differ in… They still share They repeat
param_poolSize trimming, alignment, BAM cleanup, reports, calling the frequency tables only
param_capMaxDepth trimming, alignment, BAM cleanup reports, capping, calling, and everything after
param_adapter1/2 nothing the whole analysis

Two runs can only differ in a param_* value by using different metadata files — metadataFile is itself a parameter, so it can be a column in the run table. Step 0 prints the resulting split before any compute is spent, so check there rather than inferring it: it names each results directory, which runs share it, and which steps it holds.

What to decide before your first run

In rough order of how expensive it is to get wrong:

  1. metadata.csv — which FASTQ pairs share an RG_Sample. Wrong here means valid results that answer a different question, and fixing it invalidates every BAM. →
  2. poolSize and ploidy — these set the frequency floor. poolSize can be given per pool in metadata.csv, and ploidy applies to a whole run. →
  3. filterFalsePositives.sampleThreshold — decides whether alleles seen in few pools survive. The default removes them. →
  4. capBAM.maxDepth — leave it at -1 unless you know your libraries need otherwise. It is the one item here you can safely decide after the first run, because the depth reports tell you what it did. →
  5. threads — must fit the machine, or the run fails at submission. →

Changing any of items 1–4 after outputs exist means deleting those outputs. That is enforced, not advisory.

Using system tools

The software block maps each tool to a command:

software {
    samtools = 'samtools'
    bcftools = 'bcftools'
    // …
}

Replacing a command with an absolute path makes the pipeline use a system installation instead of the conda environment. This is supported but not recommended: the environment pins exact builds because Pool-seq results depend on the precise behavior of the pileup and filtering tools, and a version mismatch will not announce itself. Use it to work around a genuine packaging problem, not as a default.

Three of them cannot be repointed. rsync, diff and find are checked at the start of a run along with everything else, so a missing one is reported before any work begins — but a path you give for them is not used. They are called by name from the small scripts that move a finished file into permanent storage, and those scripts read none of your settings. If you need a different one, change what is on your PATH.