Skip to content

Directory Layout

There are three directories, and keeping them apart is most of understanding the layout: the installation, which is a tool; mainDir, which is your project and where the work happens; and storageDir, which holds finished results.

The installation

~/.local/opt/PoolSeqFlow-<version>/
├── bin/                          # Run, never sourced; all executable, all on PATH
│   ├── atomic_mv.sh              # Cross-filesystem moves, staged and renamed
│   ├── cap_depth.awk             # Truncate a BAM to a depth ceiling
│   ├── check_install.sh          # Verifies an installation (PoolSeqFlow check install)
│   ├── check_project.sh          # Verifies a project (PoolSeqFlow check project)
│   ├── check_analysis_install.sh # The analysis layer, R packages included
│   ├── classify_manifest.sh      # Sorts a parameter change into added/changed/removed
│   ├── config_migrate.sh         # Backs migrate_config
│   ├── createDepthFile.sh        # Extract AD/DP columns from a VCF
│   ├── depth2freq.awk            # Convert allelic depths to frequencies
│   ├── depth_cutoff.py           # Choose a sample's depth ceiling from its histogram
│   ├── filterFalsePositives.sh   # Cross-sample support filter
│   ├── find_artifact.sh          # Locate an output across the storage tiers
│   ├── MajorAlleleToRef.py       # Re-encode VCF with the major allele as REF
│   ├── parse_metadata.py         # Read and validate metadata.csv
│   ├── parse_multirun.py         # Read and validate the run table
│   └── write_citations.py        # Writes CITATIONS.md and references.bib per run
├── lib/                          # Sourced by another script, never run; not on PATH
│   ├── tool_version.sh           # Asks each tool its version, one way per tool
│   └── wrapper_lib.sh            # Machinery shared by the wrapper and the checks
├── install/                      # The pinned environments, and nothing else
│   ├── environment.yml           # The pipeline's
│   └── environment-analysis.yml  # The analysis layer's
├── citations/                    # What the pipeline itself cites
│   ├── references.bib            # Authored; edit this one
│   └── citations.json            # Generated from it, and what a run reads
├── scripts/
│   ├── 0_verify_environment.nf   # The nine checks that gate everything else
│   ├── 1_build_dictionaries.nf   # BWA, SAMtools and SnpEff indices from your reference
│   ├── 2_trim_reads.nf           # Trim Galore, then composition-aware clipping
│   ├── 3_align.nf                # BWA-MEM
│   ├── 4_clean.nf                # Name-sort → fixmate → markdup → addRG → filter → index
│   ├── 5_reports.nf              # Alignment and coverage reports
│   ├── 6_variant_call.nf         # One joint bcftools mpileup and call
│   ├── 7_vcf2freq.nf             # Normalize, filter, split, convert to frequencies
│   ├── 8_annotate_variants.nf    # SnpEff, optional
│   ├── 9_completion.nf           # Promotion: moving finished artifacts to storageDir
│   ├── citations.nf              # Writes CITATIONS.md and references.bib for the run
│   ├── metadata.nf               # Reading metadata.csv, and the projections from it
│   ├── resolve_parameters.nf     # Computed parameters, and one parameter set per run
│   └── variants.nf               # Which runs share which work, and where it goes
├── analysis/                     # The analysis frame. No module is part of it
│   ├── lib/nf/                   # The workflow library a module imports
│   ├── lib/rmd/report.Rmd        # The PDF report every analysis carries
│   ├── modules.nf                # Finds and dispatches an installed module
│   ├── 0_verify_analysis.nf      # The builtin verify module
│   ├── complete.nf               # Promotion for analysis results
│   ├── frame.config
│   ├── frame.version             # What a module declares it needs
│   ├── analysis.config.template
│   ├── citations.json
│   ├── references.bib
│   └── modules/                  # THE MODULE STORE — created empty, and yours to fill
├── manual/                       # This manual
├── nextflow.config
├── parameters.config.template
├── metadata.csv.template
├── multi-run.csv.example
├── poolseqflow.nf                # Workflow entry point
├── analysis.nf                   # Entry point for the analysis layer
├── dryrun.nf                     # Entry point for the layout preview
└── PoolSeqFlow                   # CLI wrapper, pipeline and analysis layer alike

One copy serves any number of projects, and it is replaced wholesale when you upgrade — which is why nothing of yours belongs in it. bin/ is prepended to PATH by nextflow.config, which is how the helper scripts are callable by bare name inside process scripts.

analysis/modules/ is the one part of the installation you add to. No module arrives with a release, so a fresh copy has an empty store; analysis modules install <name> puts a module there along with the libraries it declares, which live together under analysis/modules/lib/. It is still part of the installation and not part of a project — replaced wholesale on upgrade, like everything else here, which is why upgrading means installing your modules again.

What you provide, on mainDir

/path/to/working/directory/  ← mainDir, and where you run from
├── parameters.config
├── metadata.csv
├── metadata.csv.example     ← left by init; reference, nothing reads it
├── Data/                    ← dataSource names this folder
│   ├── Sample1_R1.fq.gz
│   ├── Sample1_R2.fq.gz
│   └── …
└── Reference/
    ├── reference.fasta.gz
    └── reference.gff.gz     ← only if annotate = true

Either reference file may be gzipped or plain.

What appears on mainDir as the run proceeds

/path/to/working/directory/
├── Reference/Dictionaries/       # Built from your reference in step 1
│   ├── reference.fasta
│   ├── reference.fasta.{amb,ann,bwt,fai,pac,sa}
│   └── snpEff/
├── Utilized/                     # Outputs still to be read; mirrors Output/'s tree
└── work/                         # Nextflow's task directories

Utilized/ empties itself as the run proceeds — each artifact moves to storageDir once the last step that needed it has finished, so a completed run leaves it empty. Under a run table each run gets its own, Utilized_<RunID>, because the runs share mainDir and would otherwise write to one path.

work/ is emptied by cleanup = true after a successful run, and PoolSeqFlow clean removes what is left. The dictionaries stay: they are derived from your reference and rebuilding them costs time for no gain.

What the pipeline produces, on storageDir

/path/to/permanent/storage/
├── Logs/                         # Per-step .log and .err, mirrored from every task
└── Output/
    ├── .poolseqflow_version      # The release that produced these results
    ├── .poolseqflow_params       # The analysis parameters behind them
    ├── .parameters.config        # Your configuration, copied verbatim
    ├── .multirun.csv             # Your run table, copied verbatim, if you used one
    ├── .poolseqflow_metadata     # The analysis-affecting columns of metadata.csv
    ├── run_parameters.txt        # Readable mirror of .poolseqflow_params
    ├── Trimmed/<sample>/         # Clipped FASTQs
    ├── Unpaired/<sample>/        # Reads whose mate was discarded
    ├── Aligned/                  # Raw BWA output
    ├── Ready/                    # Cleaned, filtered, indexed BAMs
    ├── VCF/                      # Call sets
    ├── Frequencies/              # The result
    └── Reports/
        ├── Alignment/
        ├── Coverage/
        ├── Depth/                # Depth histogram and chosen ceiling, per sample
        ├── Fastqc/<sample>/
        ├── Trimming/<sample>/
        ├── 0_verify_environment.txt
        ├── snpeff_summary.{html,genes.txt}
        └── PoolSeqFlow_pipeline_{report,timeline,trace,dag}.*

That is the shape for a single run.

Under a run table

Output/ and Logs/ gain one level, and only divergence is named. Work every run shared goes under All_Runs/, work some of them shared under Shared_<N>/, and whatever a run did alone under its own RunID. Below each of those, the subtree is exactly the one above.

Take three runs against two references, one of them filtered harder:

RunID,referenceFile,gffFile,vcffilter.minDP
run_a1,ref_a.fasta.gz,ref_a.gff.gz,20
run_a2,ref_a.fasta.gz,ref_a.gff.gz,40
run_b,ref_b.fasta.gz,ref_b.gff.gz,20
/path/to/permanent/storage/Output/
├── .poolseqflow_version          # The invocation's records stay at the root,
├── .poolseqflow_params           #   describing the whole set of runs
├── .parameters.config
├── .multirun.csv
├── run_parameters.txt
│
├── All_Runs/                     # every run agreed on the reads and the trimming
│   ├── Trimmed/<sample>/
│   ├── Unpaired/<sample>/
│   └── Reports/
│       ├── Fastqc/<sample>/
│       ├── Trimming/<sample>/
│       └── PoolSeqFlow_pipeline_{report,timeline,trace,dag}.*
│
├── Shared_1/                     # run_a1 and run_a2 share reference A
│   ├── members.txt               #   -> "run_a1", "run_a2"
│   ├── Aligned/
│   ├── Ready/
│   ├── VCF/
│   └── Reports/{Alignment,Coverage}/
│
├── run_a1/                       # same reference, different minDP: only step 7 differs
│   └── Frequencies/
├── run_a2/
│   └── Frequencies/
│
└── run_b/                        # a reference of its own, so it shares nothing past trimming
    ├── Aligned/
    ├── Ready/
    ├── VCF/
    ├── Reports/{Alignment,Coverage}/
    └── Frequencies/

Logs/ mirrors that shape exactly, so a step's log sits beside the output it produced.

Three things are worth reading off it:

  • The trimming is done once, not three times. Every run reads the same FASTQ files with the same settings, so there is one set of trimmed reads under All_Runs/.
  • Shared_1 needs members.txt, because a group number says nothing about which runs are in it. Numbers are assigned in order of appearance in the table, so reordering rows can move them.
  • run_a1 and run_a2 hold only Frequencies/. Everything earlier was identical between them and lives in Shared_1; the filtered VCFs that differ are intermediates and do not survive.

.poolseqflow_metadata is the one record that does not sit at the root with the others. It describes the column order of a particular VCF, so it is kept beside that VCF — here, one in Shared_1/ and one in run_b/.

The Nextflow reports describe the whole invocation rather than any one run, so there is a single set under All_Runs/Reports/.

A run that sets its own storageDir gets none of this. It has a results tree to itself, shares nothing, and repeats every step alone.

What survives a completed run

Several steps delete their inputs once the next stage has consumed them, and the deletion follows the symlink — the permanent copy goes too (why). So the contents of Output/VCF/ mid-run and after a completed run are not the same.

With vcf.fileName = 'Test':

File Survives Notes
Test.vcf Yes Raw call set from step 6
Test_sort.vcf No Deleted by the false-positive filter
Test_sort_fp.vcf No Deleted by the depth/quality filter
Test_sort_fp_dq.vcf No Deleted by the SNP/INDEL split
Test_sort_fp_dq_snp.vcf No Deleted by frequency conversion
Test_sort_fp_dq_indel.vcf No Deleted by frequency conversion
Test_annotated.vcf Yes If annotate = true; annotates the raw call set
Test_snp_freq.tsv Yes In Frequencies/
Test_indel_freq.tsv Yes In Frequencies/

So Output/VCF/ holds exactly two files after a completed run: the raw call set, and the annotated one if you enabled annotation. Every VCF step 7 produces is an intermediate, including the fully filtered one — what that becomes is the frequency tables.

Likewise, Output/Trimmed/<sample>/ keeps only the _clipped.fq.gz files after a completed run — the intermediate _val_1/_val_2 files are removed once clipping has used them.

Output/Aligned/ and Output/Ready/ are both kept. Nothing deletes a BAM.

Sizing storage

Rough guidance for planning, per sample:

Directory Relative size Kept?
Data/ Your input Yours
Trimmed/ Slightly under input, after clipping Yes
Unpaired/ Small Yes
Aligned/ Comparable to trimmed input Yes
Ready/ Smaller — duplicates and filtered reads removed Yes
VCF/ Depends on variant density, not read count Partly
Frequencies/ Larger than the VCF it came from — one row per allele, not per site Yes

The two BAM directories dominate. If disk is tight, Output/Aligned/ is the safe thing to remove after a completed run: step 4 has consumed it, and only a re-run from alignment would need it back.

mainDir needs room for more than scratch. It holds your reads and reference permanently, the dictionaries built from them, and — while the run is going — every output that a later step still has to read. At peak that is most of a run's intermediates at once. work/ itself stays small, because it holds symlinks rather than copies.

The peak on mainDir falls as the run proceeds, since each artifact leaves for storageDir as soon as the last step needing it finishes. A completed run leaves Utilized/ empty.