Changelog¶
All notable changes to PoolSeqFlow will be documented in this file.
The format follows Keep a Changelog, and this project adheres to Semantic Versioning.
The public API is what a result depends on: the published table formats, parameters.config, metadata.csv, and what an analysis folder contains. The command line is documented, not frozen — a verb may gain a word or a listing may get shorter without that being a breaking change, because nothing already computed becomes unreproducible. A version moves when something changes, never on a schedule.
In practice a version number says what upgrading will cost you: a third number is an inconvenience fixed, a second means the tool works mostly fine but there is a caveat worth reading and an incompatibility with its immediate fix, and a first is a different experience.
3.2.0 - 2026-09-24¶
Your raw reads can stay where they were handed to you. Data/ no longer has to be one flat folder (one folder per sample, one per sequencing run, nested as deep as you like, or all together as before). Nothing you already have needs moving, not one pinned tool changed version, and no result is affected. As with every release, the new installation's module store starts empty and your modules are installed into it again.
One thing to read before upgrading, under Changed: a sample with only one of its two mates is now refused, where it used to be dropped from the run without a word. If a project of yours has ever been missing a mate, this release will tell you.
Added¶
- Reads may sit in subfolders of
Data/. However your facility handed them over is how you can leave them. A sample is named by its file, never by its folder —Sample1_R1.fq.gzis sampleSample1wherever it lies — sometadata.csvdoes not change,readPatterndoes not change, and the pipeline reads no meaning into the folders you use. A pair whose two mates are in different folders is still that pair. What does not work is the same file name in two places: that is one sample twice over, each copy would overwrite the other's results, and the run stops and names both directories. - Hidden folders are skipped. Anything beginning with a dot is not searched.
.snapshotis the one that matters: NetApp exposes it read-only inside every directory on a great deal of shared storage, and it holds a copy of every file per snapshot, so searching it would find every sample many times over. If one of them does hold reads, step 0 names it, because a folder full of reads ignored in silence is how you lose samples without being told. - Tab completion, installed with the pipeline. It completes every command, the two targets
checktakes, the analysis commands, and the modules you actually have installed — soPoolSeqFlow analysis <TAB>lists what is in your store rather than a fixed list. bash finds it in your next shell with nothing to do; zsh needs two lines, which the installer prints.PoolSeqFlow analysis modules install <TAB>offers nothing on purpose: those names come from the catalogue over the network, and a keystroke should not make a network request.
Changed¶
- A sample missing one of its mates is refused.
readPatterntakes the two together, and a lone file was simply not a pair — so it was left out of the run, silently, with nothing to say a sample had gone. The only check against it asked whether the total number of FASTQ files was even, which two samples each missing a mate satisfy between them. Every sample is now required to have exactly one of each mate, and the run names the ones that do not. This is the change most likely to stop a project that used to start: it is telling you about data that was already incomplete. - macOS is no longer listed as a supported platform. No released version has installed on one. It ran in early development, when the environment was a plain list of names and versions. Conda solved that per platform, picking builds for whatever machine it is on. Once the environment was frozen by export it stopped being portable, because an export names the exact builds one machine received:
install/environment.ymlhas pinnedlinux-64since the repository's first commit, so conda cannot solve it on a Mac and the install fails before anything else is reached. The Requirements table had said "Linux or macOS" since before 3.1.0 and that was never true of a release anyone could download. The manual now says Linux, and explains what supporting macOS would take — three sets of pinned files instead of one, since Apple Silicon and Intel are separate platforms again. It is planned, not prioritized.
Removed¶
resume. It printed a deprecation notice and then did exactly whatrundoes. Resuming is filesystem-based and always has been: every step checks whether its outputs already exist and skips itself if they do, soPoolSeqFlow runis both "start" and "resume". Nothing about that behavior changes.
Commits¶
- (263f266) rehash addition reverted, incorrect fix
- (d7ac846) Test suite fixed and reordered
- (65f66f7) Tab completion added, deprecated resume command removed completely
- (796811c) Test suite redundancies resolved
- (7034b19) new release check added
- (b1b856a) Data directory is now more permissive for different data subfolder structures
- (8ff7334) prep-release stale message removed
- (d411635) americanize language
- (b5c63ba) Removed wordy comment
- (1f48890) Test suite fixes
3.1.2 - 2026-09-23¶
A compatibility release. Nothing you set changes and no result moves. Not one tool that computes anything moved a version; what changed is which machines the analysis layer installs on, plus three annoyances on machines that are not the one it was built on.
Fixed¶
analysis installfailed on any machine with glibc older than 2.39, which is most clusters.install/environment-analysis.ymlpinnedsysroot_linux-64=2.39, and that is a demand on your machine rather than something conda installs —__glibcis conda's name for the glibc you already have. Nothing asked for it: the solver took the newest the exporting machine allowed. Both environments now install on anything from glibc 2.28 up. The pipeline layer was never affected, so ifPoolSeqFlow installworked whileanalysis installdid not, this was why.rehash: command not foundon every conda call, where something in your environment exportsZSH_VERSION. conda runs\rehash— zsh's name forhash -r— whenever that is set, including inside the bash wrapper, where the command does not exist. Nothing was failing underneath it. The wrapper now definesrehash.~/.local/bin is NOT on your PATHwhen it is. The check compared$PATHas text, so a trailing slash, a doubled slash or a symlinked home each made it answer no. It resolves directories now. Worth fixing for what follows it: the message tells you to edit a shell profile that was correct.uninstallreportedDirectory not emptyand stopped there, on network storage, where removing a file that is still open leaves a.nfs*placeholder — and the process holding them open isPoolSeqFlowitself. It now finishes the uninstall and tells you the placeholders go when it exits. Previously it left~/.local/binpointing at the release you were removing.
Changed¶
- The manual states the host requirement. No document had named a glibc minimum, so there was no way to know before installing. It now carries the floor of 2.28, why no flag works around it, and the error text under Troubleshooting.
rsyncis what sets the floor, in both environments. The only common machine below it is CentOS 7, end of life since June 2024.
Commits¶
- (d4c4028) development notes
- (45737e3) The analysis install error due to glibc dependency is resolved by decreasing the floor
- (f4876a9) Manual updates
- (6ec83c4) glibc floor raised to 2.28 rsync requires it
- (8936bbc) Version prep fix
- (aa72721) Potential solution to prep-version.sh error that blocks version preparation
- (3b745cb) fixes and debug on test suite
- (de1459e) prep_version is now verbose
- (aad713c) environment files are exported for v3.1.2
- (9e4c36a) wrapper fixes
- (af7939b) wrapper fix
- (d9737de) Development notes added
- (b143021) rm error fix
- (2e3ead6) test suite fix, version bump has revert
- (430f987) test suite fixes
- (fd01d41) Test suite fixed, now uses right conda env
3.1.1 - 2026-09-10¶
Inconveniences fixed. No known incompatibilities. Nothing you set changes, nothing you have installed needs touching, and no result computed under 3.1.0 is affected. This tidies two things that were merely noisy and one page that was actively misleading.
Fixed¶
analysis modules availablelisted every published version of every module. The catalogue holds one row per version — that is whatinstallresolves against — so a listing of rows was each module's history rather than something to act on, and it grew with every publish. It now shows one line per module: the newest version your release can run, which is exactly whatinstall <name>would take. Both commands ask the same question through the same code, so they cannot tell you different things. Older versions stay published and stay installable by naming one.- Installing a module asked conda to install packages that were already there. A module declares everything it needs whether or not the release's own environment already carries it, and most of what a manifest names normally is already there — so installing one meant a network round trip and a solve to be told nothing had to happen. Only what is actually missing is installed now, and a module needing nothing new says so instead of listing seven packages it is not going to touch. What a manifest declares is unchanged, and uninstalling still reasons over the whole list.
- The published modules page offered libraries as though you could install them. Published modules rendered every catalogue row, so the five libraries appeared beside the modules — and a library is never installed by name. It now lists modules only, newest version each, matching what
availableshows, and names the libraries as reference. The page corrects itself on the next site deploy rather than on upgrade.
Changed¶
- The CHANGELOG says what the public API is. Semantic Versioning requires a project to declare one and this had not, which left every release number a judgment call. It is what a result depends on — the published table formats,
parameters.config,metadata.csv, and what an analysis folder contains. The command line is documented, not frozen.
Commits¶
- (a7d68a5) Published basicstats, association, and mds modules
- (8270a4e) module list was showing every version, fixed to latest version
- (028b391) for analysis install redundant conda install is removed
- (61c34f8) Documentation update
- (b31e291) Clarification on versioning on CHANGELOG
3.1.0 - 2026-09-10¶
This version makes the module system do what it was built for. 3.0.0 introduced modules that are published, versioned and installed on their own timetable — and then shipped three of them inside the release, which is the one thing that design was meant to avoid. A module in the payload is a module that moves when the pipeline moves. Now nothing ships: analysis/modules/ is a store that arrives empty and holds what you put in it.
Upgrading leaves you with an empty module store, and that is the whole of the upgrade. Ask the release you are leaving what it has, then install each into the new one. Nothing else moves: your projects, your configuration and every analysis already published are untouched, and the old release keeps its own modules and still runs.
PoolSeqFlow-3.0.0 analysis modules list # what the old release has
PoolSeqFlow analysis modules install mds # and again, into the new one
PoolSeqFlow check now takes a word. There are two questions — is this installation sound, and is this project sound — and one command answering both meant answering neither well. A bare check is refused rather than guessing which you meant, because whichever it picked would leave the other unchecked while reporting success.
Changed¶
- No module ships inside a release, and no library either.
analysis/modules/is the install store: gitignored, absent from the tarball, and empty until you install something.PoolSeqFlow analysis modules install <name>puts one there. This is what lets a module be fixed, improved or published without waiting for a pipeline release — and the cost is that a new installation starts with nothing in it and you choose what goes back. - A module arrives with the libraries it declares. The shared arithmetic more than one module wants — effective pool size, gene diversity, per-site allele frequencies, Nei's distance, the chunking — is now five libraries, each published and versioned like a module and installed into
analysis/modules/lib/. You never ask for one by name: it arrives with whatever needs it, and leaves when nothing installed still declares it. A published result still carries the library code folded into the script that produced it, so a result explains itself whatever the store holds later. checkis two commands.check installverifies the installation — every tool the release is built to run, and every helper inbin/.check projectverifies a project — thatparameters.configis current and parses, thatmetadata.csvand the run table parse, and that every command resolves as that project configures them, so a tool repointed at a system binary is checked the way the run will call it. Runcheck installfrom anywhere, including before you have a project; runcheck projectfrom your project directory.check installasks the release's own environment rather than yourPATH. Every tool it looks for is pinned ininstall/environment.yml, so one that resolves from anywhere else means the environment is missing a package and your system's copy is standing in — at some other version, on your machine only. That is now reported asOUTSIDE THE ENVIRONMENTand fails the check. It is worth catching because it is quiet: the pipeline runs, the results look fine, and nothing reproduces anywhere else.- The installation directories say what they are for.
bin/holds everything that is run rather than sourced, the check scripts included;lib/holds what is sourced;install/holds the two pinned environment files and nothing else; andcitations/is new, holding the pipeline's ownreferences.biband thecitations.jsongenerated from it. Nothing you set moves, and no project is affected.
Added¶
PoolSeqFlow check project— the configuration and the commands a project names, checked without spending a run. It readsparameters.configthrough Nextflow itself and the two tables through the same parsers step 0 uses, so what it tells you is what a run would tell you.
Fixed¶
- Uninstalling a module could take a package another module still needs. The keep-list is built by reading one manifest per installed module, and the reader did not terminate its last line — so with several modules in the store the last package of one and the first of the next arrived joined, and a name at that boundary dropped out of the list of things to keep. The same defect applied to libraries. The shipped modules declared no packages, so this could only be reached by a module published against 3.0.0 that declared its own.
- Uninstalling a module could take a package the release itself is built on. A module declares what it needs whether or not the baseline already carries it, so
r-ggplot2appears in a manifest and ininstall/environment-analysis.ymlboth. The removal now subtracts the baseline, and nothing the release provides can leave with a module.
Commits¶
- (91e027f) Publish the three shipped modules
- (45770fd) Release notes now reads from changelog
- (7428880) install no longer checks for paramters.config, which is not in the install folder anymore
- (6d38c88) Full rework of modules
- (598d955) wrapper check is aligned with the current file structure
- (4f8a313) citations move to their own folder
- (dbc522d) modules rework continued
- (c9b26e2) check project fix
- (bdd3c93) Major bug fixes related to the migration of files to different folders
- (ae1af19) Release notes updated
3.0.0 - 2026-09-10¶
This version is about accessibility. I tried to do as much engineering as possible using the most common tools and knowledge to make sure that the pipeline can create reproducible results for the users. The outputs now contain, not only the parameter set used in each analysis, there is a list of citations for all the tools used for each portion of the analysis. The pipeline refuses to run when parameter combination is changed mid-run, this is because, one cannot say which one is used for certain analysis if they change it mid-run. This was a reproducibility choice. However, if the user wants to compare multiple parameter combinations, multi-run feature is added. The pipeline handles it in the most efficient way, by finding where the divergent parameter applies and creates separate workflows for each parameter combination. Analysis layer is built to accommodate different ploidies and multiallelic sites. I also improved the manual/website which now has all explanation and history about the tool.
Upgrading is not automatic and it is not optional reading. Your project now has two directories instead of one, RGTags.csv is replaced by a file that does not convert from it, and the depth ceiling is measured per sample rather than fixed. Run ./PoolSeqFlow migrate_config before anything else — it carries your settings across, prints the mv commands for the files that have to move, and explains each change in place. A configuration from 2.2.0 is now refused rather than half-read, so there is no way to discover this partway through a run.
The pipeline also grew a second half. ./PoolSeqFlow analysis runs statistical modules over a finished run's frequency tables: three ship in this release, more install from a repository without waiting for a PoolSeqFlow release, and every published result carries the script that produced it, the assumptions it was computed under, and a PDF report. It runs in its own conda environment and does not touch the pipeline's.
Changed¶
- A project now has two roots and they must be different paths.
mainDirholds your reads, reference,parameters.configandmetadata.csv, and is where you launch;storageDirholds finished results. Before 3.0 there was one directory doing both, calledprojectDir.migrate_configrenames the parameter, reports it underRenamed this release, and prints the moves for the files that were on the old root — it never moves anything itself. RGTags.csvis replaced bymetadata.csv, and it is not a rename. The old file carried SAM read-group tags and nothing else. The new one describes the experiment: it names each sample, decides which rows merge into one pool throughRG_Sample, and carries per-sample pool sizes and adapters. It also has somewhere to put the experiment itself —exp_for what you set,pt_for what you measured as a response,cov_for what you measured alongside — which is what the analysis layer reads. Nothing converts the old file, and the run stops at step 0 until the new one exists. Start frommetadata.csv.template, which documents every column.- The depth ceiling is measured per sample instead of fixed at 2000. Step 5 reads each sample's own depth histogram and step 6 caps that sample's BAM before calling, so a shallow library is no longer judged at a deep one's ceiling.
capBAM.maxDepth = -1is that measurement;variantCall.maxDepthbecomes a second ceiling on top of it and ships as0, which mpileup reads as no limit. Your old2000is not carried across, andmigrate_configreports it underFormat changed this releasewith the reason. To reproduce 2.2.0 results exactly:variantCall.maxDepth = 2000andcapBAM.maxDepth = 0. - Pool size, ploidy and detection sensitivity can vary per sample, set in
metadata.csvthroughparam_poolSize. One number for a whole run judged a pool of 10 at a pool of 500's resolution. Rows sharing anRG_Samplemust agree, and a blank cell counts as a different answer rather than as agreement. - Parameters renamed for what they do rather than which tool runs them.
samtools.*iscleanBAM.*,bcftools.*isvariantCall.*,diploidyisploidy— the pipeline was never limited to diploids and the name said otherwise.migrate_configcarries every value across. - A configuration from an older release is refused rather than partly read. Nextflow reads
parameters.configas given, so an absent parameter used to interpolate into a path as the literal stringnulland the run started anyway.run,resume,dryrun,reset,analysis completeand running a module now stop and namemigrate_config.clean,drycleanandmigrate_configitself are unaffected. - The
coresblock and the tooloptionsstrings are computed for you and ship commented out. They are not gone:migrate_configreports them underStill yours to set, and uncommenting a line takes one back. Coming from 2.2.0 that is thirteen parameters — the eightcoresvalues and the fiveoptionsstrings your file already had.
Added¶
- The analysis layer.
./PoolSeqFlow analysis <module>runs a module over a finished run's frequency tables, in a conda environment of its own that./PoolSeqFlow analysis installcreates. Three modules ship:basicstats(site counts, depth, effective pool size and gene diversity per pool),association(each allele's frequency regressed on a phenotype measured per pool, with a permutation p), andmds(the pools placed by Nei's minimum distance, corrected for sampling, on a classical MDS). Each publishes the script that produced its numbers, the shared library folded in, areferences.bibfor the methods it used, and a PDF report of the whole folder. - Every module states what it cannot answer. A module's manifest carries its assumptions and its limits as text, and the run prints them beside the result — the permutation floor a small design cannot go below, why an MDS distance can be negative and that this is correct, that a capped BAM is biased toward reads mapping earliest. A result that is model-based says so where it is read.
- A module store and a repository.
./PoolSeqFlow analysis modules {list|available|install|uninstall}installs a module published separately from the pipeline, with its conda packages, checked against a checksum and against what this release can run. Installing one that needs a GPL package tells you so: the pipeline stays Apache-2.0 and each module carries its own license. - Every run writes its own citations. A pipeline run leaves
citations.txtinOutput/, naming each tool it actually invoked with the version that tool reported — probed at run time, so a tool repointed at a system binary is recorded as what ran rather than as what shipped. A published analysis carriesCITATIONS.mdandreferences.bibbeside its results, covering the methods each module used as well as the software. What to cite stops being something you reconstruct months later. - Multi-run projects. One data source, several parameter sets, described in
runs.csvwithmultiRun = true. Runs sharing an input share the work rather than repeating it, and step 0 prints what is shared before any compute is spent. ./PoolSeqFlow dryrunanddryclean— check the configuration and preview what a run would do, writing nothing into either root, then remove the preview../PoolSeqFlow init— populate an empty directory with the template configuration and metadata, ready to edit.- Several versions install side by side. Each release has its own environment and payload, so an in-flight project can finish on the release it started on.
./PoolSeqFlow listshows what is installed anduninstallasks which. - A test suite, nineteen suites split by what they cost, so a change runs only the cases it can reach:
test/run_tests.sh --changedpicks them from the file you edited. A module ships its own cases inside its own directory. - Both conda environments are pinned to exact builds.
install/environment-analysis.ymlwas a hand-written specification and is now an export of an environment the full suite passed against, asinstall/environment.ymlalready was.
Fixed¶
- FastQC was given up to eight threads and needed two. Its
-tcounts files processed simultaneously, not threads per file, and step 2 hands it one pair. Measured on a pair of 2M-read files,-t 2is 1.93× faster than-t 1while-t 4,-t 6and-t 8are no faster at all — and every thread past the second costs roughly 250 MB of resident memory, twice per sample. On a memory-constrained machine that is the difference between a run finishing and being killed. The documented ladder always said two; the code had drifted. DepthProfiledeclared one of the three files it publishes, so the depth histogram and the depth report were outside Nextflow's tracking and outside the skip that avoids rebuilding them.- Two data-loss defects in
atomic_mv.sh, both found by reading rather than by a failure: moving a directory onto an existing one could lose a file that only the destination had. snpEff's configuration file is settable throughparameters.configrather than fixed, and multiple annotation databases are supported.- The step 7 intermediates are no longer published.
<name>_sort_fp_dq.vcfand the split SNP and INDEL VCFs had been landing inOutput/VCF/since 1.0; the called VCF and the annotated one are what a run keeps.
Removed¶
params.gff,params.dir.scripts,params.dir.output.temp, andrgTagsFilewithrgTagsPath.migrate_configreports each underNo longer used. There are no legacy fallbacks anywhere in the pipeline: a parameter that is gone is handled at migration and nowhere else.
Commits¶
- (922c3b3) Manual is mostly moved to github pages
- (f9b3437) Version check is added.
- (85fbe74) Merge checks added
- (b7f95eb) pipefails added
- (a52da7f) pipefail added, atomic move implemented on reference file
- (a57e01b) Copy check added to BuildSnpEff
- (78f85c1) Raw filename check is enforced
- (15832fe) RGTags column count check enforced
- (513d417) VerifyAll now publishes the results in the output folder.
- (163ec6e) Log management improved, older logs are now retained
- (012f0d1) Hyphen is now accepted separator for reads
- (dc1e5a1) Combined log of the last run is assembled as a single file
- (afd78b6) Temporary file management improved
- (57251ce) conda env check hardened, rm legacy files explained
- (8902645) atomic move hardened
- (e440fa5) better config migration rules implemented
- (2d26123) bump version is improved
- (7a782a0) pycache is untracked
- (14e847d) version enforcement clarification, version bump fix
- (269d64c) Minor fixes in dictionary counting
- (2318a89) Medium importance fix on snpeff config file, now can be set properly through parameters
- (5d516a8) Support for multicharacter mate tokens added
- (c6d8dda) script hardening for midstream failures
- (372ac66) NextFlow warning sweep
- (5ed5002) Fix for a bug introduced in the previous stage.
- (eb777cf) Test suite added
- (af0d04e) Multiple versions become installable going forward
- (c2e1bf7) Added features to list all installed versions and uninstall all
- (69c8dac) cutadapt min length is clarified, and guards added
- (1e1fdbf) Test suite is being implemented now testing 30 cases
- (81edc1d) Environment creation with new version control is fixed
- (bc3b212) A script for preparing a new version is added
- (1e10c77) projectDir is renamed as storageDir for clarity
- (9ffe284) snpEff improvement for multiple database support, verify env improvements
- (9d8a17b) mainDir and storageDir cannot be same anymore, guards added.
- (2ee0e27) classify_manifest moved to its own script, test suite efficiency improved
- (79efcd2) parameter control automated, override is still allowed
- (9b37dbc) Directory for install is now separate and checked
- (6516431) storage management improvements are being implemented
- (991addf) Install function now installs the wrapper and scripts and multiple versions can be installed
- (2b75a6c) completion checks started to be built
- (cb72804) Dictionary tests are added
- (070fb81) trim paths are fixed for storage management
- (12e12d8) fastqc files, aligned bam, and bai files storage improvements
- (de7e288) rest of the storage management is done
- (065fcc7) clean and reset reworked to address previous changes
- (f38267a) The first half of multi-run work is completed
- (8546b50) Multiple runs from a single data source is added
- (1afb0d7) Process redundancy is resolved
- (3a3a10a) Directory structure clarified
- (6cc24e8) target.dir is removed
- (5cde1a4) make dir is added as a bug fix
- (9f4e647) sharing is implemented
- (048bd33) Environment verification is now aligned with multi-run
- (7f30177) verify environment is now checking parameter changes midway, version change mid-run is blocked
- (360cfbd) dryrun and dryclean added
- (1cd315c) Multiple bug fixes
- (c28fbdc) metadata.csv replaced and expanded rgtags.csv
- (4bdd3aa) per sample poolsize and sensitivity added, metadata addition is complete now
- (f2ec822) docs management streamlined with a master manual.md
- (00da07d) gitignore, gitattributes, and github workflow changes to reflect docs management
- (982e0c4) init project added, uninstall improved
- (7ad02c1) major documentation and comment overhaul
- (cc00833) parameter renaming for clarity, samtools is now cleanBAM and bcftools is variantCall
- (c3a3191) automatic maxDepth calculation added
- (da95a4b) Major commit: Analysis layer arrived, install and uninstall repaired, docs fixed
- (4303166) analysis.nf added to payload items
- (7810c56) Analysis layer foundation is being worked: verify analysis, tests, and config templates are in
- (f5c6135) Analysis output control mechanism
- (c209cce) Minor fixes on analysis related changes
- (abe00fb) the module store landed
- (5fafa80) analysis libraries are being built to allow modular statistical tools design
- (7d65893) analysis wrapper is folded inside the main wrapper
- (b432794) Docs and comments pass
- (4a5f635) Two invocation launch for analysis layer
- (0c82ac9) main.nf verification for each module
- (bf0dcfd) Module development rules added
- (d9886c5) Modules manifest and management subcommands added
- (f4f1104) fixed atomic.mv concurrency issue
- (9567e2a) module analysisPlan fix
- (ebf08a2) analysis writing results safely, intermediate checks, granular move back
- (77fdbd7) atomic_mv fix for multiple failure scenarios, rsync dependency added
- (f4508f5) analysis complete command moves files to storage, the resume copies them back to do the analysis
- (1f1fbcf) clean now cleans staged files, default params are hardcoded for analysis
- (c164ec5) defaults.config is now frame.config and users should not change it
- (e2406fa) provenance, frame and modules versioning
- (9cb22fe) R script is emitted along with results now
- (4c3d394) citations for modules added, test suite fixes
- (e76774a) Comment cleanup
- (aa48373) snpEff reports are both copied now
- (9f9b7da) histogram ceiling is a parameter, archive gate enumerates, glob loops guarded, execution defaults reachable
- (a06c312) analysis lib nf files moved
- (4cb4ae7) the experimental design, module settings, and how a result says to read itself
- (2915346) time variable is now configurable, time series feature added
- (fabdb5c) diploidy, poolsize are recovered from metadata
- (c6544d7) Fixes on derived parameters
- (381541e) F1 basicstats, and the PDF report every analysis carries
- (d674e8f) Split the suite, and run only what a change touches
- (306dd26) phenotype variables pt_ and covariate variables cov_ are added to the metadata
- (f2a7774) Manual and development notes are added
- (e3eb887) allele frequencies, benchmark comparisons added, analysis versioning fixes made
- (15d4c34) experimental design and covariate readjustments
- (12633a6) association analysis, validation of association, tests, reorganization of analysis.config
- (4cb4cff) Comments housekeeping
- (d024dd1) mds landed
- (409aef7) Language corrections for drift
- (d92411e) module store landed, the depth profile bug fixed
- (77cb91e) docs and comments pass
- (fc3d2a1) realease prep, minor fixes
- (e384928) Repo structure is created
- (268f508) Manual check, better metadata
- (12a1eb6) manual updates
- (2b5236f) fastqc now limits the cores to 2
- (abaa496) prep version script now covers analysis
- (ff08f61) Environment upgrade
- (e539333) config migrate prepared, new guards enforce it
2.2.0 - 2026-08-16¶
This release changes results. vcffilter.minDP previously had no effect on the output at all; it now removes sites. Read the first entry under Changed before upgrading a project that has outputs you intend to keep — and expect step 0 to stop your next run, because the analysis parameters have changed. That is the guardrail working; the report names the folders to delete.
Alongside that: a documentation site, an installation check that fails an install rather than letting a half-built environment through, and a release process that publishes a verified download so nobody has to clone the repository to use the pipeline.
Changed¶
vcffilter.minDPnow filters, where before it did nothing. The depth filter wasvcftools --minDP, which expresses a failed genotype-level test by rewritingFORMAT/GTand nothing else — it never touchesADorDP, and it never removes a site. Because step 7's major-allele normalization sets everyGTto./.before that filter runs, and because frequency conversion readsADrather thanGT, the setting had no path to the output: running the old command with--minDP 20and with--minDP 50produced byte-identical frequency tables. It is nowbcftools view -e "FMT/DP<N", applied before the quality filter. The test is per site, not per sample: a site is removed if any sample falls below the depth, so the weakest library sets the threshold for the whole cohort. CheckOutput/Reports/Coverage/for your least-covered sample before trusting the default of20— on a run with one thin library it can remove most of the call set.params.vcftoolsis nowparams.vcffilter. The block never mapped to one tool and now genuinely does not: depth filtering is bcftools, quality filtering is vcftools../PoolSeqFlow migrate_configcarries your values across to the new names and reports them asRenamed this release.- Two bcftools parameters were the wrong way round.
baseQualMinsuppliedmpileup -q, which is the mapping quality minimum, andvarQualMinsupplied-Q, the base quality minimum. Both default to30, so no run changes behavior — but anyone who tuned one was tuning the other. - Citations point at the Zenodo concept DOI (10.5281/zenodo.19245611) rather than a version DOI. The badge previously pointed at the v1.0.0 record, which is frozen and therefore permanently flagged "a newer version is available". The concept DOI always resolves to the newest release. Papers should still cite the version DOI of the release they ran —
./PoolSeqFlow citeexplains which and why. install/install.shremoved. The wrapper'sinstallsubcommand creates the environment itself and never called it; the script also used a relative path toenvironment.ymland aconda activatewith no shell hook, so running it directly would not have worked either.cutadapt.min_lengthis still not applied, and the template now carries a commented-outoptionsline to switch it on deliberately rather than leaving the parameter looking active.
Added¶
- A documentation site at https://ozankiratli.github.io/PoolSeqFlow/, built with MkDocs Material and published from
mainby GitHub Actions. It goes well beyond the README: when Pool-seq fits and when it does not, why the pipeline replaces Nextflow's-resumeand what that costs, the full filter chain from alignment flags to frequency conversion with what each stage removes, and how to read the frequency tables. Broken internal links fail the build. ./PoolSeqFlow check— verifies an installation and reports what it finds. Every command the pipeline invokes, with the version each reports; every helper inbin/, present and executable, since they are called by bare name offnextflow.config'sPATHand a lost executable bit fails mid-run; and thatparameters.configparses. With a config present the tool list is read fromparams.softwarethroughnextflow config, so a command repointed at a system binary is checked as configured rather than as shipped. It also runs at the end ofinstalland fails the install if anything is missing — an environment that was created but is short a tool would otherwise surface partway through step 4, hours in../PoolSeqFlow cite— prints the citation for the copy you have, with its version filled in, and explains which DOI to use.- A release workflow. Tagging
v*publishes a curated tarball: the pipeline only, in a versioned directory, built withgit archiveso the executable bit on./PoolSeqFlowandbin/*comes from the git index rather than the runner's umask. What ships is decided byexport-ignorein.gitattributes, and the workflow asserts both directions — required files present, repository furniture absent — along with the executable bits, shell syntax and the version the extracted wrapper reports. It refuses to publish unless the tag, both version strings in./PoolSeqFlowand a changelog section all agree.SHA256SUMSis attached, andPoolSeqFlow.tar.gzcarries a stable name for scripted installs. config_migrate.shhandles renamed parameters. A rename was previously two unrelated events — oneDROPPED, oneNEW— and your tuned value silently reverted to the template default. Renames now carry the value across and report it asRenamed this release. If a rename also changes what the parameter means, adding it toreformatted()makes the template value win while still surfacing the change.
Fixed¶
- The sensitivity formula in
bin/filterFalsePositives.sh -hnow readss = 1 / (2 * [DIPLOIDY] * [POOLSIZE per SAMPLE]), matching whatparameters.configcomputes. The correction in 2.1.1 was itself wrong. Help text only; the value the pipeline passes was never affected.
Commits¶
- (2c29d35) Depth filtering fix, and minor config corrections.
- (78ac157) Typo fix, not a functional problem
- (b087cd5) site is added to gitignore
- (8595513) Renaming check is added to the migration script
- (0e63c74) dev files added
- (9687114) Check install status added
- (6b5a549) Release workflow added
- (7aae36e) check install added to workflow
- (aaec06c) Citation fixes
- (d423c8e) Website is finished
2.1.1 - 2026-08-15¶
A stability release. Nothing new to configure and no change to how a run is invoked — this closes the gaps where a result could be quietly wrong or quietly irreproducible. An existing parameters.config needs no changes.
Added¶
- Step 0 refuses to run when the analysis parameters changed since the existing outputs were produced. Completed steps are skipped by looking for output files, not by checking what produced them, so a changed
poolSizeor filter threshold would otherwise leave one output folder holding results from two different settings. The values behind a set of outputs are recorded in.poolseqflow_paramsand mirrored to a read-onlyOutput/run_parameters.txt. Path, resource and software parameters are excluded; anything added in a later release counts as analysis-affecting until decided otherwise. - Step 0 refuses to run when
RGTags.csvchanged after the file was consumed. The tags are written into the BAMs at step 4 and the row order is fixed into the VCF at step 6, and neither is re-derived once its output exists. The report separates a changed tag value (invalidatesReady/,VCF/,Frequencies/) from a reordering (invalidatesVCF/,Frequencies/only) and names the folders to delete; deleting them is what clears the check. Projects whose outputs predate this release adopt their current file as the baseline, with a note to confirm it against the BAM headers. - Sample columns follow
RGTags.csvrow order, so results come out arranged the way the samples were laid out rather than however they sort as strings. Where several rows share anSM, the merged column takes the position of the first of them. - Duplicate
IDdetection. A row is looked up byIDand only the first match is read, so a repeatedIDsilently discarded the later rows and gave that sample the wrong tags — producing a perfectly valid BAM that nothing downstream could flag. - CRLF repair for
RGTags.csv. A file saved from Excel on Windows carries a stray carriage return into the last tag of every row; it previously failed withInvalid tag 'PU', which names nothing useful. Step 0 now rewrites the file with Unix line endings, preserving permissions and ownership, and reportsRGTAGS LINE ENDING CHECK: FIXED. bin/atomic_mv.sh— moves that cross a filesystem boundary now stage through a.partfile and rename into place.
Changed¶
workDiris now undermainDir. It was a relative path, so the scratch/permanent split the pipeline documents was not actually in effect — work directories landed wherever the pipeline was launched from.- Variant calling receives its BAMs in a defined order.
collect()emitted them in task-completion order, so the sample column order of the VCF varied between runs on identical input; three consecutive runs gave three different orders. The sort keys on the sample id, because the file paths begin with Nextflow's work-directory hash and sorting those is no better than chance. - All 22 cross-filesystem moves are atomic. A plain
mvacross filesystems is a copy followed by an unlink, so a job killed mid-move left a truncated file under its final name — which the existence-based skip logic then accepted as a completed step. resetis behind a typedDELETE_MY_ANALYSISconfirmation and also clears.poolseqflow_paramsand.poolseqflow_rgtags, which would otherwise fail the next run's checks against outputs that no longer exist.cleanandresetresolve paths throughnextflow configrather than parsingparameters.configas text. Values are interpolated, so text matching returned the wrong path.RGTags.csv.templatenow shows the replicate andSM-merge pattern, withDScarrying a per-replicate descriptor instead of repeating the sample name.
Fixed¶
- The sensitivity formula in
bin/filterFalsePositives.sh -hwas missing a factor of two. It reads = 1 / ([DIPLOIDY] / [POOLSIZE per SAMPLE])and should reads = 1 / 2 * ([DIPLOIDY] / [POOLSIZE per SAMPLE]). Help text only — anyone who ran the script by hand and followed it would have passed the wrong-s.
Commits¶
- (6e18762) Minor fix in help for manual use
- (715f822) Parameter change detection added.
- (f73b33d) workDir and reset fixes
- (0f5666f) Output parameters to a file
- (408efb4) File move process improved
- (dc2ea72) Sample ordering in vcf fixed. NF orders samples first come first serve
- (4a9f89d) Sample ordering in vcf fixed. RGTags guardrails added.
2.1.0 - 2026-08-15¶
Resource allocation is now declared to Nextflow rather than only passed to the tools, and there is a helper for carrying an older configuration forward.
Added¶
./PoolSeqFlow migrate_config— rebuildsparameters.configfrom the current template, backs the original up, carries across every setting whose parameter still exists, and reports what it kept, what is new, what the pipeline now computes for itself, and what it dropped. It refuses to carry a value the template derives, so it cannot reintroduce a stalesnpEff.dbor a hand-set thread count. The report is a starting point: a parameter whose behavior changed while its value still looks ordinary will be carried across, so compare against the template afterwards.- Every process declares
cpus, so Nextflow schedules against real requirements instead of assuming one core per task. Previously threeAligntasks each using ~2.2 cores ran concurrently on an 8-core machine withcpus=1recorded for each. params.memory, feedingresourceLimitsalongsideparams.threads, so one place sizes a run.
Changed¶
- Tools now read
task.cpusrather than thread counts baked into option strings, so the number Nextflow reserves and the number the tool receives cannot diverge. Overridingcpusin a profile now changes the tool's behavior too. TrimReadsreserves Trim Galore's full footprint.--cores Nruns N+4 threads (measured:--cores 8peaks at 12 OS threads), so the process reservescores.trimTotaland maps back to the worker count. A request larger than the machine now fails withProcess requirement exceeds available CPUsinstead of silently oversubscribing.- JVM garbage-collection threads come from
task.cpus.-XX:ParallelGCThreadswas read from a config string, socpushad no effect on SnpEff or FastQC. resourceLimitsmoved toparams.threads/params.memory; it was hardcoded and would not follow a change tothreads.- Eight parameters removed after the rework left them unreferenced: the five per-tool
threadsvalues,fastqc.bundledOptions,java.garbageCollectandjava.options. Each looked like a knob that did nothing. TrimReadsno longer exports_JAVA_OPTIONS; Trim Galore 2.x is a native binary with a bundled FastQC and never starts a JVM.
Fixed¶
parameters.config.templatewas missing parameters the pipeline requires —annotate,snpEff.runOptions,rgTagsPath,diploidy— and carried a differentvcftools.minDPand different report directory names. A configuration built from it failed step 0 withRGTAGS VERIFICATION: STATUS=FAIL. The template is now generated from the reference configuration and resolves identically to it.- Step 7 created the wrong output directory.
SortRefAltByFrequencyranmkdir -pon the frequencies folder and then moved into the VCF folder, which only worked because step 6 had created it first. parameters.configcontained thedir { }block and the reference path assignments twice, byte-identical.
2.0.1 - 2026-08-12¶
Fixed¶
parameters.config.templatewas missing theparams.coresblock introduced in 2.0.0, so a configuration created from the template kept the old hardcoded per-tool thread counts instead of deriving them fromparams.threads. Existing runs were unaffected; the template now resolves identically to a 2.0.0 configuration at everythreadsvalue.
2.0.0 - 2026-08-12¶
Major upgrade to Nextflow 26 and Trim Galore 2.x. This release is not backwards compatible: an existing parameters.config will fail mid-run, and completed trimming and annotation outputs are regenerated on first use.
Breaking¶
- Requires Nextflow 26 (
26.04.6). Thecleanup { }block innextflow.configwas invalid and is rejected by the stricter config parser; the pipeline could not start on 26 before this release. - Requires Trim Galore 2.x (
2.3.0). The bundled FastQC engine and--basenamenaming are both assumed. parameters.configis no longer tracked in git. Copyparameters.config.templateand re-apply your settings — see Upgrading from an earlier release in the README. Carrying an older file over causes a later step to fail with a barenull: command not found.- Trimmed read filenames changed from
<sample>_R1_val_1.fq.gzto<sample>_val_1.fq.gz. Trimming is redone once on the first run after upgrading. - SnpEff database name is derived from the GFF filename instead of being set by hand, so an existing database directory is not found and is rebuilt.
params.fastqc.memoryis now a plain number of megabytes (2048). The previous"2G"was rejected by FastQC, which silently fell back to its 512 MB default.-resumeis no longer passed to Nextflow../PoolSeqFlow runalready resumes through its own filesystem checks;./PoolSeqFlow resumeremains as a deprecated alias.
Added¶
- Automatic core allocation: a
params.coresblock derives every tool's thread count fromparams.threads. Trim Galore is costed on its true footprint (--cores Nruns N+4 threads), andthreads = 1forces everything single-core. params.trim_galore.autodetect— whentrue, no adapter is passed and Trim Galore detects it; whenfalse, both adapter sequences are required.- Step 0 now validates trimming parameters, failing early if auto-detection is off and the adapters are missing or are not DNA sequences.
unzipadded toenvironment.ymland toparams.software, so step 0 verifies it. It was always required by the clipping step but never declared.- Trim Galore 2.x
*_trimming_report.jsonfiles are kept alongside the.txtreports. - README section on upgrading, covering the stale-configuration failure mode.
Fixed¶
- Trimming failed on standard Illumina filenames. Output patterns assumed the read number ended the filename, so
<sample>_R1_001.fastq.gzproducedMissing output file(s) *_R1_val_1.fq.gz. Output naming is now pinned with--basename. - The SnpEff database could never be built. Only the GFF was staged, so the build aborted with
Cannot find reference sequence.and produced no.binfiles. The reference FASTA is now copied in alongside it. - Clipping thresholds could be computed from truncated data. A zero base fraction aborted the AWK pass mid-pipeline; without
pipefailthe failure was swallowed and a wrong read-length limit was used silently. Zero divisors are skipped, bounds are validated, and the chosen parameters are logged. - Alignment and coverage reports paired BAMs with indexes by position rather than by sample; the two channels are now joined on
pair_id. SkipGFFCheckwas unparseable because of a duplicatedscript:label, soannotate = falsecould not run at all on Nextflow 26.- Resuming a completed run failed at the frequency step, which linked a bare filename and created a self-referential symlink.
parameters.config.templatedid not parse on Nextflow 26 — it used${mainDir}instead of${params.mainDir}in nine places.- README documented parameters that do not exist (
refGenome,refGFF,ploidy) and placed the data directory under the wrong root.
1.0.1 - 2026-06-02¶
Fixed¶
- Removed
conda update --allfrom the install script. Package versions are now fully governed byenvironment.yml, improving reproducibility and preventing unintended upgrades after installation.
1.0.0 - 2026-03-26 — Initial Public Release¶
Added¶
Core pipeline (Nextflow DSL2)
- End-to-end Pool-seq analysis workflow (poolseqflow.nf) with 9 modular steps
- Wrapper script (PoolSeqFlow) exposing install, run, resume, clean, and reset subcommands
Step 0 — Environment verification
- Pre-run checks for all required input files, folder structure, RGTags CSV format, and software dependencies
- Generates Reports/0_verify_environment.txt
Step 1 — Reference indexing
- Builds BWA, SAMtools (.fai), and SnpEff indices from a gzipped reference FASTA and GFF
Step 2 — Quality control and trimming - FastQC assessment of raw reads - Adapter trimming via Trim Galore with user-specified adapter sequences - Automated per-cycle base-composition analysis of FastQC reports - Intelligent hard-clipping via Cutadapt driven by A/T and G/C imbalance thresholds — no manual parameter tuning required
Step 3 — Alignment - Paired-end alignment to the reference genome using BWA-MEM
Step 4 — BAM post-processing
- Full SAMtools-based cleanup: name-sort → fixmate → coord-sort → markdup → addreplacerg → filter → index
- Configurable alignment filter flags (samFlags.filter, samFlags.required)
Step 5 — Alignment reporting
- Per-sample alignment statistics via bamtools stats
- Coverage summaries via samtools coverage
Step 6 — Variant calling
- Multi-sample SNP and indel calling with BCFtools mpileup + call in multiallelic mode
- Outputs VCFs with per-sample AD and DP FORMAT fields
Step 7 — VCF to allele frequency tables - Major-allele normalization: VCF re-encoded so the major allele is always REF - Multiallelic site support throughout variant calling and frequency conversion - Ploidy- and pool-size-aware minimum frequency filter: \(f_{\min} = 1 / (2 \times ploidy \times poolSize)\) - Depth and quality filtering - SNP / INDEL split - Export to tab-separated allele frequency tables
Step 8 — Variant annotation (optional)
- SnpEff-based functional annotation, toggled via params.annotate
Resume logic
- Custom filesystem-based resume strategy using symbolic links between mainDir (working directory) and projectDir (permanent storage)
- Completed steps are skipped based on presence of permanent output files — resilient to job timeouts, reboots, and work/ directory cleanups
- Supports HPC environments where compute nodes and storage are on separate filesystems
Configuration
- parameters.config for analysis parameters (mainDir, projectDir, poolSize, ploidy, adapter sequences, filter flags)
- nextflow.config for computational resources (CPUs, memory, executor)
- RGTags.csv template for sample read group metadata
- parameters.config.template for getting started
Environment
- Single conda environment (install/environment.yml) covering all dependencies
- Automated install and verification scripts (install/install.sh, install/test-install.sh)