Getting Started¶
Requirements¶
| Operating system | Linux. macOS and Windows are not supported — see below |
| glibc | 2.28 or newer — RHEL and Rocky 8, Debian 10, Ubuntu 18.10, or anything later. ldd --version prints what you have. CentOS 7 is the one common machine below it, and is end of life. Why → |
| Conda | Conda or Miniconda |
| Git | Optional, for cloning |
| Working directory | mainDir — where you launch the pipeline. Holds your reads, the reference, this project's configuration and everything actively processed. It has to persist between runs; it is not scratch space |
| Permanent storage | storageDir — where finished results are kept. Must be a different path from mainDir, and neither may be the installation itself. The pipeline will refuse to start otherwise |
| Disk | mainDir needs room for your inputs and the files produced along the way; storageDir for the BAMs, VCFs and tables you keep |
Everything else is installed for you. ./PoolSeqFlow install builds an isolated conda environment from install/environment.yml, which pins every tool to an exact build:
| Tool | Version | Role |
|---|---|---|
| Nextflow | 26.04.6 | Workflow engine |
| FastQC | 0.12.1 | Read quality metrics |
| Trim Galore | 2.3.0 | Adapter and quality trimming |
| Cutadapt | 5.2 | Composition-aware clipping |
| BWA | 0.7.19 | Alignment |
| SAMtools | 1.24 | BAM processing and filtering |
| BAMtools | 2.5.3 | Alignment statistics |
| BCFtools | 1.24 | Variant calling and VCF manipulation |
| VCFtools | 0.1.17 | VCF filtering and splitting |
| SnpEff | 5.4.0c | Variant annotation (optional) |
| OpenJDK | 25 | Runtime for FastQC, SnpEff and Nextflow |
Pinning is deliberate. Pool-seq results depend on the exact behavior of the pileup and filtering tools, and an unpinned environment would make two runs of the same config non-comparable.
Why Linux only¶
Windows. The pipeline moves each file it produces to where it belongs and leaves a symbolic link behind, and it relies on Unix path semantics throughout. Neither behaves correctly on native Windows filesystems, and WSL only works under some filesystem configurations — which is not a guarantee worth documenting. See Symbolic links instead of copies.
macOS. Every tool a release runs is pinned to an exact build in install/environment.yml and install/environment-analysis.yml, and those files are exports from a Linux machine: they name builds that exist for linux-64 and for no other platform. Conda cannot solve either of them on a Mac, so the install fails before anything else is reached. Supporting macOS means a second set of pinned files exported and verified on a Mac — one for Apple Silicon and one for Intel, since conda treats them as different platforms — and that is a piece of work rather than a flag. Earlier releases listed macOS as supported; that was never true of the pinned environments, and saying so was the error.
The glibc floor¶
The floor is glibc 2.28, and it applies to both environments. In distribution terms that is RHEL and Rocky 8, Debian 10, Ubuntu 18.10, and everything released after them. The one common machine it excludes is CentOS 7, which reached end of life in June 2024.
__glibc is conda's name for the glibc your machine has, and ldd --version prints it. Below the floor, conda refuses the solve rather than installing something that would fail later, and the message names a virtual package instead of a distribution:
LibMambaUnsatisfiableError: Encountered problems while solving:
- nothing provides __glibc >=2.28 needed by rsync-3.4.4-hffd6c76_1
No flag works around it. Conda offers one that forces the solve, and using it would install binaries the system cannot run — turning a refusal you get in seconds into a crash you get hours into a run. If your machine is below the floor, use a newer one.
What sets it is rsync, which both environments carry. Every artifact the pipeline produces is moved to permanent storage by a copy that is verified before the original is removed, and rsync is what stages that copy. So the floor is a property of the whole release rather than of the analysis layer: the analysis environment additionally carries a C and C++ compiler, because every module builds its hot path on your machine rather than shipping a binary, but that toolchain reaches back further than rsync does and is not what decides this.
Ready¶
Two pages. Install gets the tool onto your machine and verifies it, which is two steps and a wait. Quick Start is everything to do with your own data — making the project, filling in the three files that describe it, and starting the run.
Install PoolSeqFlow Quick Start
Already have it installed and upgrading from an earlier version? Read Upgrading first — your parameters.config is never touched by an update and can be missing parameters the new code expects.