Skip to content

Quick Start

This page takes you from an installed PoolSeqFlow to a running pipeline. It assumes only that Install is done; everything to do with your own data starts here.

1. Make your project

A project is a directory of your own. It is not the installation — that is a tool, shared by any number of projects and replaced wholesale when you upgrade, so a project kept inside it would not survive one. The pipeline refuses to start if you point it at the installation.

mkdir -p /path/to/project
cd /path/to/project
PoolSeqFlow init

The paths are yours to choose; the structure inside them is not. A project ready to run looks like this — init makes everything except the reads, the reference and metadata.csv:

/path/to/project/                ← mainDir, and the directory you run from
├── parameters.config            ← your settings, read from wherever you launch
├── metadata.csv                 ← what your samples are; you write this one
├── metadata.csv.example         ← every column explained, to write it from
├── Data/                        ← dataSource names this folder
│   ├── Sample1_R1.fq.gz
│   ├── Sample1_R2.fq.gz
│   ├── Sample2_R1.fq.gz
│   ├── Sample2_R2.fq.gz
│   └── …
└── Reference/
    ├── reference.fasta.gz       ← referenceFile
    └── reference.gff.gz         ← gffFile, only if annotate = true

The file names are examples. readPattern is what finds the reads — *_R{1,2}.fq.gz by default, which is what matches the pairs above — and referenceFile and gffFile name the two in Reference/.

The reads may sit in subfolders of Data/ instead, or as well. However your sequencing facility handed them over is how you can leave them — one folder per sample, one per run or lane, nested as deep as you like, or all together as above. All three of these are read the same way:

Data/                    Data/                      Data/
├── Sample1_R1.fq.gz     ├── Sample1/               ├── run_A/
├── Sample1_R2.fq.gz     │   ├── Sample1_R1.fq.gz   │   ├── Sample1_R1.fq.gz
├── Sample2_R1.fq.gz     │   └── Sample1_R2.fq.gz   │   └── Sample1_R2.fq.gz
└── Sample2_R2.fq.gz     └── Sample2/               └── run_B/
                             ├── Sample2_R1.fq.gz       ├── Sample2_R1.fq.gz
                             └── Sample2_R2.fq.gz       └── Sample2_R2.fq.gz

A sample is named by its file, never by its folder. Sample1_R1.fq.gz is sample Sample1 wherever it lies, so metadata.csv does not change and neither does readPattern — the folders are yours to arrange and the pipeline does not read meaning into them. A pair whose mates are in two different folders is still that pair, and is taken as one.

What step 0 checks is the pair itself: every sample has exactly one of each mate. Three ways to break it, all of which the reader would otherwise accept in silence:

in Data/ what it is
Sample1_R1.fq.gz with no Sample1_R2.fq.gz anywhere a sample that would be left out of the run without a word
run_A/Sample1_R1.fq.gz and run_B/Sample1_R1.fq.gz the same mate twice, which would be aligned against itself
a full pair under both run_A/ and run_B/ one sample over again, each copy overwriting the other's results

If two folders really do hold two different samples, give them two names. If they are one sample sequenced twice, that is two rows in metadata.csv sharing an RG_Sample — see What a unit is.

Hidden folders are skipped, and the run says so. Anything beginning with a dot is not searched — .snapshot on NetApp storage holds a copy of every file per snapshot, and searching it would find every sample many times over. If one of them does hold reads, step 0 names it so you know what was left out.

init never overwrites. Running it again in a project you have already filled in reports what is there and changes nothing.

It copies parameters.config for you because that is a settings file you edit in place. It does not write metadata.csv, because that is a table describing your experiment — which FASTQ pairs are one pool, how many individuals each holds, and the order your result columns come out in — and a copied one would describe someone else's. Write it yourself, starting from metadata.csv.example, and read Metadata before your first run rather than after it.

Finished results do not go here. They go to storageDir, which has to be a different directory. See Requirements.

Then edit parameters.config: mainDir, storageDir, readPattern, referenceFile, poolSize and ploidy at minimum. The reference and the annotation may be gzipped or plain — the pipeline takes either and unpacks what it needs into Reference/Dictionaries/.

Analyzing one set of reads under several parameter sets — two reference genomes, say? Run PoolSeqFlow init_multi instead. It does everything init does, switches multiRun on, and copies multi-run.csv.example into the project. It does not write the run table itself, for the same reason it does not write metadata.csv.

2. Verify it any time

There are two checks and they answer different questions, so check takes a word and refuses without one. An installation is a tool; a project is your configuration and your data. Neither answers the other, and a bare check would have to guess which you meant — leaving the other unchecked without saying so.

PoolSeqFlow check install    # the tools and helpers this release is built to run
PoolSeqFlow check project    # the configuration and the commands this project names

Both end by reporting how many checks passed, and fail loudly if any did not rather than summarizing.

check install

Tools
  from ~/.local/opt/miniconda3/envs/PoolSeqFlow-<version>

  nextflow       nextflow     OK       26.04.6 build 12646
  cutadapt       cutadapt     OK       5.2
  …
  samtools       samtools     OK       samtools 1.24
  bcftools       bcftools     OK       bcftools 1.24
  …

Pipeline helpers

  atomic_mv.sh                 OK
  depth2freq.awk               OK
  …

Every command the release is built to run, with the version each reports. The list is the canonical one — what this release expects its environment to provide.

And each has to come from that environment. Every tool on the list is pinned in install/environment.yml, so one that resolves from anywhere else means the environment is missing a package and your system's copy is standing in — at whatever version it happens to be, and only on this machine. That is reported, not passed over:

  samtools       samtools     OUTSIDE THE ENVIRONMENT  /usr/bin/samtools

It is the failure worth catching, because it is the quiet one: the pipeline runs, the results look fine, and nothing reproduces anywhere else. Reinstalling is the fix — PoolSeqFlow install.

Every helper in bin/, present and executable. nextflow.config puts that directory on PATH and the process scripts call the helpers by bare name, so a lost executable bit fails mid-run rather than at startup. bin/ is enumerated rather than listed, so a helper added to a release is checked without anyone remembering to say so.

The three check_*.sh scripts live there too — bin/ is where everything that is run rather than sourced lives — and are the one thing skipped: they are run by the command line, never by a process script, so a run does not depend on them. lib/ is not enumerated at all, because what is in there is sourced rather than run, which is the whole reason the two directories are separate.

It reads no parameters.config and needs no project. Run it from anywhere, including straight after installing and before you have a project at all.

check project

Run it from your project directory.

Configuration

  parameters.config            WRITTEN FOR THIS RELEASE
  parameters.config            PARSES
  metadata.csv                 PARSES
  runs.csv                     not in use

Tools, as this project configures them

  nextflow       nextflow               OK       26.04.6 build 12646
  samtools       /usr/bin/samtools      OK       samtools 1.19
  bcftools       bcftools               OK       bcftools 1.24
  …

  tool list from: params.software in parameters.config

Your files parse — parameters.config through Nextflow itself, and metadata.csv and the run table through the same parsers step 0 uses, so what you are told here is what a run would tell you. It also says whether the config was written for this release, which is the one thing that stops a run before anything else is read.

Every command, as you configured it. The list comes from params.software, so a command repointed at a system binary is checked the way the run will call it — and the second column shows what will actually be invoked. That override is the setting most likely to be wrong and least likely to announce itself.

A path outside the environment is not a finding here. check install treats one as a fault, because the installation is supposed to provide its own tools; check project does the opposite, because repointing one is a thing you are allowed to do and this is where you see the result. It reports what will be invoked and whether it runs, and leaves the judgment to you.

A project that repoints nothing gets the same tool answers from both checks. That is the point: the difference between them is exactly your own configuration.

If params.software cannot be read, check project says so and checks no tool. It never falls back to the canonical list — the whole reason the section exists is that it reads yours.

3. Fill in the essentials

Open parameters.config. Nine settings need your attention before a first run; everything else has a working default.

params {
    mainDir       = "/path/to/working/directory"  // where you run, and where your inputs live
    storageDir    = "/path/to/permanent/storage"  // where finished results are kept
    dataSource    = 'Data'                        // subdirectory of mainDir holding the FASTQs
    readPattern   = "*_R{1,2}.fq.gz"              // glob matching your paired FASTQ files
    referenceFile = 'reference.fasta.gz'          // reference genome, in mainDir/Reference
    gffFile       = 'reference.gff.gz'            // annotation (only if annotate = true)
    poolSize      = 50                            // individuals per pool
    ploidy        = 2                             // copies of the genome per individual
    annotate      = true                          // run SnpEff annotation (step 8)
}
Setting How to choose it
mainDir The working directory you ran init in. Holds Data/, Reference/, your configuration and everything actively processed, so it has to persist between runs — not node scratch that is wiped between jobs. Put it on the fastest storage that satisfies that.
storageDir Where finished results are kept. A different directory from mainDir, and the pipeline refuses to start otherwise: the two are storage tiers, and outputs move from the first to the second as each step that needs them completes.
readPattern Must match both mates with a {1,2} group. If your files end _1.fastq.gz/_2.fastq.gz, write "*_{1,2}.fastq.gz".
poolSize Individuals in one pool, not the total across pools. Sets the smallest allele frequency worth believing — see The Filter Chain. If your pools differ in size, give each its own value in metadata.csv; this setting is the default for any that do not.
ploidy Ploidy of the organism: 2 for diploid, 1 for haploid, 4 for tetraploid.
annotate false skips step 8 and makes gffFile unnecessary.

Set these through the file, never the command line

PoolSeqFlow rejects command-line parameter overrides. Nextflow delivers --param values as strings, so --annotate false sets the string "false" — which Groovy evaluates as true, leaving annotation switched on with no warning. Why →

The two directories cannot be the same path. They are storage tiers, not a preference: the pipeline works on mainDir and moves each output to storageDir once the last step that needed it has finished, which cannot mean anything if they are one place. On a cluster this is the difference between a node's fast disk and the archive it is backed by; on a laptop, make them two directories and the same reasoning still holds — one is churn, the other is what you keep.

4. Size the run to your machine

threads = 8          // cores a single task may use
memory  = '24 GB'    // memory ceiling for a single task

Set threads to the cores you actually have — on HPC, the size of one node. Every tool's thread count follows from this one number; do not set the per-tool counts by hand. A request larger than the machine fails immediately:

Process requirement exceeds available CPUs -- req: 12; avail: 8

Details and the full ladder: Resources.

5. Write metadata.csv

One row per FASTQ pair. SampleID must match the sample name readPattern takes from your filenames — with *_R{1,2}.fq.gz, the file Sample1T1Rep1_R1.fq.gz gives Sample1T1Rep1.

SampleID,RG_Sample,RG_Library,RG_Platform,param_poolSize,exp_population,exp_time,pt_resistance,cov_temperature
Sample1T1Rep1,Sample1T1,Lib1,ILLUMINA,50,Pop1,T1,susceptible,21.5
Sample1T1Rep2,Sample1T1,Lib1,ILLUMINA,50,Pop1,T1,susceptible,21.5
Sample1T2Rep1,Sample1T2,Lib1,ILLUMINA,50,Pop1,T2,resistant,22.1
Sample1T2Rep2,Sample1T2,Lib1,ILLUMINA,50,Pop1,T2,resistant,22.1
Sample2T1Rep1,Sample2T1,Lib1,ILLUMINA,40,Pop2,T1,susceptible,18.0
Sample2T1Rep2,Sample2T1,Lib1,ILLUMINA,40,Pop2,T1,susceptible,18.0
Sample2T2Rep1,Sample2T2,Lib1,ILLUMINA,40,Pop2,T2,resistant,19.4
Sample2T2Rep2,Sample2T2,Lib1,ILLUMINA,40,Pop2,T2,resistant,19.4

Three things this file decides:

  • RG_Sample decides what counts as a sample. Rows sharing one are merged into a single VCF column and their read depths add together. The eight rows above produce four columns, not eight — two populations at two timepoints, each sequenced twice.
  • param_poolSize sets that pool's detection limit. It describes the pool rather than the row, so rows sharing an RG_Sample have to agree on it. Leave the column out and every pool uses the global poolSize.
  • Row order decides column order in the VCF and the frequency tables.

The prefix is what a column means. No pipeline step reads any of the three below — analysis modules do, and they build their whole design from them:

Prefix What it holds Example above
exp_ something you set — the experiment's own structure exp_population, exp_time
pt_ a phenotype you measured on the pool, and are testing against pt_resistance
cov_ measured too, but neither set nor the thing being tested cov_temperature

Add as many of each as your experiment needs; the names after the prefix are yours. A column with no prefix is recorded just as faithfully and read by nothing, so a variable you might analyze later is worth prefixing now. Editing any of them never invalidates results you already have. All three describe the pool — what differs between two rows of one pool takes no prefix, because once their reads are merged nothing can tell them apart again. Metadata has the full account.

exp_time is the one name treated specially. A project that has it must set analysis.timeVar.kind before any analysis will run — the pipeline itself does not care. Leave the column out if this is not a time course. See The time axis.

All of it is covered in Metadata. Getting RG_Sample wrong is the most common way to end up with results that are valid but not what you meant.

6. Run

PoolSeqFlow run

That is also the resume command. Every step checks whether its outputs already exist and skips itself if they do, so an interrupted run picks up where it left off with no extra flag. There is no -resume.

Run it from your project directory — the one holding parameters.config. The pipeline reads its settings from wherever you launch, so running from anywhere else stops with a message saying so.

To see where a run's results would go before spending any compute on it, PoolSeqFlow dryrun builds that directory tree empty and changes nothing else; PoolSeqFlow dryclean removes the preview.

To start genuinely from scratch, use PoolSeqFlow reset first — it requires typing DELETE_MY_ANALYSIS to confirm.

7. Check the output

storageDir/
├── Logs/
└── Output/
    ├── Frequencies/        ← the result: <name>_snp_freq.tsv, <name>_indel_freq.tsv
    ├── VCF/
    ├── Ready/              ← cleaned, indexed BAMs
    ├── Reports/
    ├── run_parameters.txt  ← the settings these results were produced under
    └── …

<name> is vcf.fileName, which is Test until you change it.

Start with Output/Reports/Depth/ and Output/run_parameters.txt. The first says what depth ceiling each sample was given and why; the second is a read-only record of exactly which settings produced these files.

Your inputs stay where they were, under mainDir — Data/, Reference/ and the dictionaries built from your reference are working material, not results, and none of them are copied here.

If you are running a table of several runs, Output/ and Logs/ gain a level: work every run shared sits under All_Runs/, work some of them shared under Shared_<N>/, and whatever one run did alone under its own RunID. A single run keeps the plain tree above.

How to read the tables: Interpreting Results.


Commands

Command Description
PoolSeqFlow install Create this release's conda environment, install the pipeline, then verify both
PoolSeqFlow init Populate the current directory as a project (what it writes)
PoolSeqFlow init_multi The same, for a project running several parameter sets over one set of reads
PoolSeqFlow check install Verify an installation — the tools and helpers it is built to run (what it covers)
PoolSeqFlow check project Verify a project — its configuration, and the commands it names (what it covers)
PoolSeqFlow run Start — or resume — the pipeline
PoolSeqFlow dryrun Create the directory tree the run would write, empty, so the layout can be approved before any compute is spent. Records nothing and changes none of your files
PoolSeqFlow dryclean Remove the preview dryrun made
PoolSeqFlow migrate_config Carry an older parameters.config onto the current template (details)
PoolSeqFlow clean Remove Nextflow work directories
PoolSeqFlow reset Remove all progress and start fresh (requires typed confirmation)
PoolSeqFlow analysis <command> The analysis layer, which carries a word of its own — listed below
PoolSeqFlow version Print the version of this copy
PoolSeqFlow cite Print how to cite this copy, and which DOI to use (why it matters)
PoolSeqFlow list List the pipelines and conda environments installed on this machine
PoolSeqFlow uninstall Remove one installed version, environment and pipeline together, after confirmation
PoolSeqFlow uninstall_all Remove every PoolSeqFlow environment and installation, after confirmation

Before anything is installed there is no PoolSeqFlow on your PATH, so the first command is ./PoolSeqFlow install, run from the folder you downloaded. Everything after that uses the installed command.

Each subcommand takes no arguments of its own, with one exception: analysis carries a word — the analysis command, or the module to run — and analysis modules carries a second. Anything else is rejected. Two environment variables adjust the wrapper instead: POOLSEQFLOW_PREFIX, where install puts things and where list and uninstall look, and POOLSEQFLOW_HOME, to run a checkout without installing it.

Naming a version. Every installed release is also on your PATH under its own name, so PoolSeqFlow-2.1.0 run uses that release and plain PoolSeqFlow uses the newest. This matters most for uninstall: with several installations present it lists them and asks which to remove, and if nothing is attached to ask — a script, a CI job — it refuses and tells you to name one, rather than guessing at which installation to delete. PoolSeqFlow-2.1.0 uninstall names it and is never asked.

What that list contains, and what removing one takes with it. Every installed version, plus the unversioned PoolSeqFlow environment if a release from before per-version environments left one behind. A version whose analysis layer is installed is marked, because an installation is removed whole — its environment, its pipeline, and its analysis environment together:

Installed under /home/you/.local:
  1) PoolSeqFlow                   (unversioned - predates per-version environments)
  2) PoolSeqFlow-2.1.0
  3) PoolSeqFlow-<version>             (this wrapper, analysis installed)

Choosing 3 removes PoolSeqFlow-<version> and PoolSeqFlow-<version>-analysis. To remove only an analysis layer and keep the pipeline that produced your results, use PoolSeqFlow analysis uninstall, which never touches anything else.

Both then list exactly what will go and ask before removing any of it, every time — including when there is only one installation and nothing to choose between. Choosing which is not the same as agreeing to the removal. Answering anything but y removes nothing, and with no terminal attached to ask — a script, a CI job — the command refuses rather than proceeding unasked. That is the same rule uninstall_all has always followed.

PoolSeqFlow resume still works as a deprecated alias for run and prints a notice.

The analysis layer hangs off analysis, the one subcommand that carries a word of its own. It ships with the same PoolSeqFlow install and is versioned with it:

Command Description
PoolSeqFlow analysis install Create this release's analysis conda environment, which carries R, then verify it
PoolSeqFlow analysis check Verify an existing analysis installation — tools, R packages, entry point
PoolSeqFlow analysis modules available List the modules published for this release's table contract
PoolSeqFlow analysis modules install <module> [version] Install one, newest readable version unless you name it. Also installs what it runs on
PoolSeqFlow analysis modules list List the modules installed for this release, and name any that are broken
PoolSeqFlow analysis modules uninstall <module> Remove one module, after confirmation. Results it produced are untouched
PoolSeqFlow analysis complete Move finished analyses and shared intermediates from mainDir/Analysis/ to storageDir/Analysis/, after confirmation. A name already taken in permanent storage stops it with nothing moved
PoolSeqFlow analysis version Print the version of this copy
PoolSeqFlow analysis cite Print how to cite this copy, plus R and every package a module runs on
PoolSeqFlow analysis uninstall Remove the analysis environment, after confirmation. The pipeline, its environment, and your results are untouched

The environment is separate from the pipeline's and is built only when you ask for it. PoolSeqFlow uninstall removes both environments of the version it is removing.

cite reads R's own citation records, so what it prints is the version of each package actually installed rather than a list kept in the documentation.

PoolSeqFlow analysis also accepts the name of an installed module. Modules are installed separately from the pipeline and each is documented with itself; the commands in the table above are the ones a release ships with.