Skip to content

Getting Started

Requirements

Operating system Linux. macOS and Windows are not supported — see below
glibc 2.28 or newer — RHEL and Rocky 8, Debian 10, Ubuntu 18.10, or anything later. ldd --version prints what you have. CentOS 7 is the one common machine below it, and is end of life. Why →
Conda Conda or Miniconda
Git Optional, for cloning
Working directory mainDir — where you launch the pipeline. Holds your reads, the reference, this project's configuration and everything actively processed. It has to persist between runs; it is not scratch space
Permanent storage storageDir — where finished results are kept. Must be a different path from mainDir, and neither may be the installation itself. The pipeline will refuse to start otherwise
Disk mainDir needs room for your inputs and the files produced along the way; storageDir for the BAMs, VCFs and tables you keep

Everything else is installed for you. ./PoolSeqFlow install builds an isolated conda environment from install/environment.yml, which pins every tool to an exact build:

Tool Version Role
Nextflow 26.04.6 Workflow engine
FastQC 0.12.1 Read quality metrics
Trim Galore 2.3.0 Adapter and quality trimming
Cutadapt 5.2 Composition-aware clipping
BWA 0.7.19 Alignment
SAMtools 1.24 BAM processing and filtering
BAMtools 2.5.3 Alignment statistics
BCFtools 1.24 Variant calling and VCF manipulation
VCFtools 0.1.17 VCF filtering and splitting
SnpEff 5.4.0c Variant annotation (optional)
OpenJDK 25 Runtime for FastQC, SnpEff and Nextflow

Pinning is deliberate. Pool-seq results depend on the exact behavior of the pileup and filtering tools, and an unpinned environment would make two runs of the same config non-comparable.

Why Linux only

Windows. The pipeline moves each file it produces to where it belongs and leaves a symbolic link behind, and it relies on Unix path semantics throughout. Neither behaves correctly on native Windows filesystems, and WSL only works under some filesystem configurations — which is not a guarantee worth documenting. See Symbolic links instead of copies.

macOS. Every tool a release runs is pinned to an exact build in install/environment.yml and install/environment-analysis.yml, and those files are exports from a Linux machine: they name builds that exist for linux-64 and for no other platform. Conda cannot solve either of them on a Mac, so the install fails before anything else is reached. Supporting macOS means a second set of pinned files exported and verified on a Mac — one for Apple Silicon and one for Intel, since conda treats them as different platforms — and that is a piece of work rather than a flag. Earlier releases listed macOS as supported; that was never true of the pinned environments, and saying so was the error.

The glibc floor

The floor is glibc 2.28, and it applies to both environments. In distribution terms that is RHEL and Rocky 8, Debian 10, Ubuntu 18.10, and everything released after them. The one common machine it excludes is CentOS 7, which reached end of life in June 2024.

__glibc is conda's name for the glibc your machine has, and ldd --version prints it. Below the floor, conda refuses the solve rather than installing something that would fail later, and the message names a virtual package instead of a distribution:

LibMambaUnsatisfiableError: Encountered problems while solving:
  - nothing provides __glibc >=2.28 needed by rsync-3.4.4-hffd6c76_1

No flag works around it. Conda offers one that forces the solve, and using it would install binaries the system cannot run — turning a refusal you get in seconds into a crash you get hours into a run. If your machine is below the floor, use a newer one.

What sets it is rsync, which both environments carry. Every artifact the pipeline produces is moved to permanent storage by a copy that is verified before the original is removed, and rsync is what stages that copy. So the floor is a property of the whole release rather than of the analysis layer: the analysis environment additionally carries a C and C++ compiler, because every module builds its hot path on your machine rather than shipping a binary, but that toolchain reaches back further than rsync does and is not what decides this.


Ready

Two pages. Install gets the tool onto your machine and verifies it, which is two steps and a wait. Quick Start is everything to do with your own data — making the project, filling in the three files that describe it, and starting the run.

Install PoolSeqFlow Quick Start

Already have it installed and upgrading from an earlier version? Read Upgrading first — your parameters.config is never touched by an update and can be missing parameters the new code expects.