Is your differential-expression result worth reporting?
Drop a bulk RNA-seq counts matrix and a sample sheet. Your browser runs the pydeseq2 pipeline on it - size factors, dispersions, Wald tests, Cook's and independent filtering, adjusted p-values - free, nothing uploaded. A paid run then reviews what the result supports or writes the pydeseq2 script that reproduces it.
Examples are synthetic experiments with a known truth, and each has a saved model run - the whole page, free.
Your recent runs
What this does, and what it does not
The page follows pydeseq2 0.5.4 and the pydeseq2 agent skill's run_deseq2_analysis.py
step for step: genes with fewer than min_counts reads in total are dropped; size
factors are the median of ratios to the geometric mean over genes with no zero count;
dispersions are estimated gene by gene with the Cox-Reid adjustment, a parametric trend
a0 + a1/mean is fitted with pydeseq2's outlier loop, and the MAP dispersions shrink
toward it with a log-normal prior; the negative binomial GLM is fitted by IRLS; Cook's distances
mark extreme counts, which are replaced and refitted when a design cell has seven or more
replicates; the Wald test compares the test level with the reference; Cook's filtering and
independent filtering remove genes before the Benjamini-Hochberg adjustment; and, when chosen,
the tested coefficient is shrunk with pydeseq2's adaptive Cauchy prior. pydeseq2 minimises with
scipy's L-BFGS-B, and so does this page, through a JavaScript port that reproduced scipy bit for
bit on 123 test problems.
Checked against pydeseq2 0.5.4 on 120 random experiments (49,206 genes) and on the three examples through the skill's own script: size factors and baseMean agree to 1e-15, 99.4% of gene-wise dispersions to 1e-6, and the number of significant genes was identical in 119 of 120 experiments (off by one in the other) and in all three examples; 96.2% of p-values agree within 1%. The rest trace to genes whose dispersion search starts at pydeseq2's 1e-8 floor, where its likelihood is a difference of numbers near 1e10 and rounding decides where the search stops - one such gene can move the dispersion trend, and every p-value with it, by a few percent. For the same reason the page computes exp and log itself rather than trusting the browser's, so every browser gives the same answer.
A statistical call is not a biological explanation: the page and the paid run see gene ids and numbers, never what a gene does. Genes without a p-value (Cook's outliers) or without an adjusted p-value (independent filtering) are not evidence of no change. The paid run reads only what the browser computed and your notes, is told never to compute a new number, and the page checks every number it writes. Derived from the agent skill @k-dense-ai/pydeseq2 (k-dense-ai/scientific-agent-skills; see the notice). No pydeseq2 code or data is included.