← SkillSafe / DESeq Desk

Is your differential-expression result worth reporting?

Drop a bulk RNA-seq counts matrix and a sample sheet. Your browser runs the pydeseq2 pipeline on it - size factors, dispersions, Wald tests, Cook's and independent filtering, adjusted p-values - free, nothing uploaded. A paid run then reviews what the result supports or writes the pydeseq2 script that reproduces it.

Examples are synthetic experiments with a known truth, and each has a saved model run - the whole page, free.

Drop counts.csv here, or
Drop samples.csv here, or

CSV or TSV. The first column of the counts file is the gene id; featureCounts output works (its annotation columns are skipped, and columns named by BAM path are matched to the sheet's ids); htseq-count and STAR summary rows (__no_feature, N_unmapped...) are left out; a file with samples in rows is recognised from the sample ids. Counts must be raw integers - pydeseq2 refuses normalised or fractional values. Nothing leaves your browser.

Or paste the two tables
Or generate a synthetic experiment with a known truth
Negative binomial counts around log-normal means with the trend 0.03 + 1.2/mean and random library sizes.
Design and settings (defaults are the skill's run_deseq2_analysis.py)

The reference level is the baseline of the contrast (positive log2 fold change = higher in the test level). With a batch column the design is ~batch + condition, adjustment first, as the skill recommends.

Load a counts matrix and a sample sheet, or generate an experiment.
Run DESeq2 first to price the review.

Your recent runs

What this does, and what it does not

The page follows pydeseq2 0.5.4 and the pydeseq2 agent skill's run_deseq2_analysis.py step for step: genes with fewer than min_counts reads in total are dropped; size factors are the median of ratios to the geometric mean over genes with no zero count; dispersions are estimated gene by gene with the Cox-Reid adjustment, a parametric trend a0 + a1/mean is fitted with pydeseq2's outlier loop, and the MAP dispersions shrink toward it with a log-normal prior; the negative binomial GLM is fitted by IRLS; Cook's distances mark extreme counts, which are replaced and refitted when a design cell has seven or more replicates; the Wald test compares the test level with the reference; Cook's filtering and independent filtering remove genes before the Benjamini-Hochberg adjustment; and, when chosen, the tested coefficient is shrunk with pydeseq2's adaptive Cauchy prior. pydeseq2 minimises with scipy's L-BFGS-B, and so does this page, through a JavaScript port that reproduced scipy bit for bit on 123 test problems.

Checked against pydeseq2 0.5.4 on 120 random experiments (49,206 genes) and on the three examples through the skill's own script: size factors and baseMean agree to 1e-15, 99.4% of gene-wise dispersions to 1e-6, and the number of significant genes was identical in 119 of 120 experiments (off by one in the other) and in all three examples; 96.2% of p-values agree within 1%. The rest trace to genes whose dispersion search starts at pydeseq2's 1e-8 floor, where its likelihood is a difference of numbers near 1e10 and rounding decides where the search stops - one such gene can move the dispersion trend, and every p-value with it, by a few percent. For the same reason the page computes exp and log itself rather than trusting the browser's, so every browser gives the same answer.

A statistical call is not a biological explanation: the page and the paid run see gene ids and numbers, never what a gene does. Genes without a p-value (Cook's outliers) or without an adjusted p-value (independent filtering) are not evidence of no change. The paid run reads only what the browser computed and your notes, is told never to compute a new number, and the page checks every number it writes. Derived from the agent skill @k-dense-ai/pydeseq2 (k-dense-ai/scientific-agent-skills; see the notice). No pydeseq2 code or data is included.