Bibliography¶
The work PoolSeqFlow is built on, and the work it can be read against.
This is not the list to cite for your results. That one is written for you, per run, and names only what the run actually invoked — see Citing the tools it runs. This page is the reading behind the choices: where an estimator comes from, why there is more than one version of it in circulation, and which other software would give a different number from the same reads.
Everything below is compiled from the references.bib files in the installation, and holds exactly the entries this manual cites. A citation with no entry, or an entry filed under no heading, fails the build — so the page cannot fall behind the modules as they are added.
A pooled allele frequency carries two rounds of sampling, and the correction for that is where implementations differ from one another most.
There are two effective sample sizes, and they are one apart
n·d/(n + d − 1) and n·d/(n + d) are both in circulation for a pool of n chromosomes read to depth d. PoolSeqFlow uses the first, which is the form implied by Hivert et al. 2018's D₂ — their sum of (d + n − 1)/n is the sum of d/n_eff under that form and no other.
They agree closely at high depth and diverge exactly where diversity estimates are most fragile: shallow coverage and small pools. At n = 2, d = 2 the first gives 1.33 and the second 1.00. If a number from this pipeline disagrees with one from other software, check which form the other used before concluding the data disagree.
The second form is widely used and hard to attribute: Kolaczkowski et al. 2011 work the pooled sampling through with pool size and depth as separate parameters and define no combined quantity at all, and Gautier et al.'s "effective pool size" is a different measure again. PoolSeqFlow therefore states which form it uses rather than citing one for it.
The diversity statistic itself is Nei 1973, the correction it needs is argued in Ferretti et al. 2013, and the estimators most likely to be compared against these are Kofler et al. 2011a and Kofler et al. 2011b — built on Futschik & Schlötterer 2010, and applied at scale in Fabian et al. 2012. Pooled variant calling as a practice starts with Koboldt et al. 2009.
The statistics PoolSeqFlow computes¶
Gower 1966¶
- Gower, J. C. (1966). Some Distance Properties of Latent Root and Vector Methods Used in Multivariate Analysis. Biometrika 53(¾), 325–338. 10.2307/2333639
- The ordination this module draws: a matrix of squared distances double centered into a Gram matrix, whose leading eigenvectors are the coordinates. It is what R's own cmdscale cites and implements, and the source of the negative eigenvalues this module publishes rather than hides - they are what a distance matrix that no flat space holds exactly produces.
Nei 1972¶
- Nei, M. (1972). Genetic Distance between Populations. The American Naturalist 106(949), 283–292. 10.1086/282771
- The distance the analysis layer places pools by, D_m = (J_X + J_Y)/2 - J_XY, where J is the probability that two chromosomes carry the same allele. It is the MINIMUM distance of this paper and not the standard distance D, which is defined in the same one and is a log of a ratio; the minimum distance is linear in the J terms, so averaging over loci and averaging over sites are the same operation and no ratio-of-averages question arises. What is applied here beyond the paper is the sampling correction: each J is replaced by its unbiased estimator from a sample of n_eff chromosomes, and J_XY takes none because the two pools are sequenced independently. It is HERE rather than in a module because it is the nei_distance library, installed with whatever module declares it.
Nei 1973¶
- Nei, M. (1973). Analysis of Gene Diversity in Subdivided Populations. Proceedings of the National Academy of Sciences 70(12), 3321–3323. 10.1073/pnas.70.12.3321
- The diversity statistic this module reports, H = 1 - sum(p_i^2), summed over every allele at a site rather than over two. It is defined for any number of alleles, which is why a triallelic site needs no collapsing here.
Benjamini & Hochberg 1995¶
- Benjamini, Y. & Hochberg, Y. (1995). Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing. Journal of the Royal Statistical Society: Series B 57(1), 289–300. 10.1111/j.2517-6161.1995.tb02031.x
- The default correction across sites, applied to the permutation p and never to the per-allele one. The number of tests passed to it is the count of sites that produced a statistic, so a site whose alleles were all invariant is excluded rather than counted as a test that failed to reject.
Phipson & Smyth 2010¶
- Phipson, B. & Smyth, G. K. (2010). Permutation P-values Should Never Be Zero: Calculating Exact P-values When Permutations Are Randomly Drawn. Statistical Applications in Genetics and Molecular Biology 9(1), Article 39. 10.2202/1544-6115.1585
- Why a sampled permutation p is (1 + reached) / (1 + draws) rather than the raw share. The raw form can return zero, which no permutation p can be, and understates by about one over the number of draws. Only used where the set of rearrangements is too large to enumerate; an enumerated null already counts the identity and needs no correction.
Ferretti et al. 2013¶
- Ferretti, L., Ramos-Onsins, S. E. & Pérez-Enciso, M. (2013). Population genomics from pool sequencing. Molecular Ecology 22(22), 5561–5576. 10.1111/mec.12522
- Why a pooled estimate needs correcting at all: individuals are sampled into the pool and reads are sampled from the pool, so an allele frequency carries both, and the estimator has to account for the smaller of the two.
Hivert et al. 2018¶
- Hivert, V., Leblois, R., Petit, E. J., Gautier, M. & Vitalis, R. (2018). Measuring Genetic Differentiation from Pool-seq Data. Genetics 210(1), 315–330. 10.1534/genetics.118.300900
- The effective sample size the analysis layer weights by, n_eff = n*d / (n + d - 1) for a pool of n chromosomes read to depth d. The paper does not write it in that form - it defines D2 as the sum of (d + n - 1)/n, which is the sum of d / n_eff under this form and under no other. Its pools are parameterized by HAPLOID size, which is what leaves n_eff, and everything weighted by it, general over ploidy. It is HERE rather than in a module because it is the n_eff library, installed with whatever module declares it, and every module that weights anything calls it.
Estimating from pooled reads¶
Futschik & Schlötterer 2010¶
- Futschik, A. & Schlötterer, C. (2010). The Next Generation of Molecular Markers From Massively Parallel Sequencing of Pooled DNA Samples. Genetics 186(1), 207–218. 10.1534/genetics.110.114397
- The estimator theory PoPoolation is built on, and the first treatment of unbiased pi and theta_W from pooled reads. Read it before deciding that a pooled estimate can be computed the way an individually-genotyped one is.
Kolaczkowski et al. 2011¶
- Kolaczkowski, B., Kern, A. D., Holloway, A. K. & Begun, D. J. (2011). Genomic Differentiation Between Temperate and Tropical Australian Populations of Drosophila melanogaster. Genetics 187(1), 245–260. 10.1534/genetics.110.123059
- One of the first genome-scale Pool-seq studies, and where its sampling properties are worked through from first principles - n chromosomes drawn from the population and sequenced to depth m, kept as two parameters rather than compressed into one. Worth reading precisely because it does not take the shortcut every effective-sample-size heuristic since has taken.
Long et al. 2026¶
- Long, A. D., Hanson, K. M. & Macdonald, S. J. (2026). The illusion of polygenicity in pool-seq genetic mapping studies: insufficient power can mask simple genetic architectures. Genetics 233(1), iyag068. 10.1093/genetics/iyag068
- What a modestly powered pool-seq association study produces when it is read as though it were well powered: scattered genome-wide significance from a single causal variant, which looks like a polygenic architecture and is noise. It is cited here because this module's permutation floor is a structural answer to it - a design of few units cannot reach a small p however large the statistic - and because the limit it describes is the reason that floor is printed rather than assumed.
Other software for pooled sequencing¶
Koboldt et al. 2009¶
- Koboldt, D. C., Chen, K., Wylie, T., Larson, D. E., McLellan, M. D., Mardis, E. R., Weinstock, G. M., Wilson, R. K. & Ding, L. (2009). VarScan: variant detection in massively parallel sequencing of individual and pooled samples. Bioinformatics 25(17), 2283–2285. 10.1093/bioinformatics/btp373
- Pooled variant calling as a thing that can be done at all - SNPs and indels detected from pooled reads at a frequency threshold, which is the premise every tool on this page rests on. PoolSeqFlow calls with BCFtools and applies its own per-pool threshold.
Kofler et al. 2011a¶
- Kofler, R., Orozco-terWengel, P., De Maio, N., Pandey, R. V., Nolte, V., Futschik, A., Kosiol, C. & Schlötterer, C. (2011). PoPoolation: A Toolbox for Population Genetic Analysis of Next Generation Sequencing Data from Pooled Individuals. PLoS ONE 6(1), e15925. 10.1371/journal.pone.0015925
- pi, theta_W and Tajima's D from a single pool, computed from a pileup rather than from called variants. Numbers from PoPoolation and from here will not generally be equal, and multiallelic sites are the largest single reason: the classical estimators it implements are defined over two alleles, where PoolSeqFlow's diversity sums over every allele at the site.
Kofler et al. 2011b¶
- Kofler, R., Pandey, R. V. & Schlötterer, C. (2011). PoPoolation2: identifying differentiation between populations using sequencing of pooled DNA samples (Pool-Seq). Bioinformatics 27(24), 3435–3436. 10.1093/bioinformatics/btr589
- F_ST, Fisher's exact test and the Cochran-Mantel-Haenszel test between pools.
Where these methods have been used¶
Fabian et al. 2012¶
- Fabian, D. K., Kapun, M., Nolte, V., Kofler, R., Schmidt, P. S., Schlötterer, C. & Flatt, T. (2012). Genome-wide patterns of latitudinal differentiation among populations of Drosophila melanogaster from North America. Molecular Ecology 21(19), 4748–4769. 10.1111/j.1365-294X.2012.05731.x
- An early large Pool-seq study, and a worked example of what these estimators are for: pi, theta_W and Tajima's D from PoPoolation and F_ST from PoPoolation2, across populations on a cline. Useful for seeing how the numbers are reported and argued from, rather than only how they are computed.