The statistical, organizational, and infrastructure principles behind rigorous product experimentation in modern SaaS.
8 domains · 27 rules.
HYPOTHESIZE · INSTRUMENT · ANALYZE · DECIDE
An experiment does not prove your idea was right — it updates your belief about how users behave. The goal is not to validate intuition. The goal is faster learning under uncertainty, including the learning that your intuition was wrong.
The Because clause is the most important part of any hypothesis. Without it, a flat result teaches you nothing. With it, a flat result tells you the mechanism you believed in was wrong — which is a more valuable learning than a small positive effect with no explanation.
Statistical validity is not optional — it is the only thing separating a controlled experiment from a coincidence. The four errors below are not edge cases. They are the default behavior of teams without explicit statistical discipline.
Looking at results daily and stopping when p < 0.05 inflates your false positive rate dramatically. You are running multiple significance tests on the same data. Use sequential testing methods (always-valid p-values, mSPRT) or fix your sample size at launch and commit to analyzing at that point only.
Statistical power (typically 80%) is the probability of detecting a true effect if one exists. A test at 50% power misses half of real effects. Power is set by sample size, MDE, and baseline variance. An underpowered experiment that shows no effect has told you nothing — you couldn't have detected it regardless.
Testing 20 metrics and declaring a win when any one reaches p < 0.05 gives roughly a 64% chance of a false positive. Apply Bonferroni correction or FDR control (Benjamini-Hochberg) when testing multiple metrics. Designate one primary metric — all others are exploratory until confirmed in a follow-up experiment.
A p-value tells you whether an effect is distinguishable from zero. A confidence interval tells you the plausible range of the effect size. p = 0.04 with CI [+0.01%, +2.3%] and p = 0.04 with CI [+0.8%, +1.4%] are completely different decisions. Always report both.
Get the randomization unit wrong and no amount of statistical rigor downstream will save you. This is the decision most Growth teams make implicitly — and most often get wrong in B2B SaaS, where users are not independent of each other.
user_id. If your metric is "revenue per account," randomize by account_id. Mismatching the unit inflates variance and produces unreliable confidence intervals — you will see more false positives and larger apparent effects than are real.hash(user_id + experiment_id + salt) mod 100 maps each user to a stable bucket permanently. A user re-evaluated randomly on each page load is simultaneously in both control and treatment, rendering all metrics meaningless. Assignment must be stable across devices, sessions, and time.A metric that is valid but insensitive will never move. A metric that is sensitive but invalid will move in the wrong direction while the business suffers. Metric design is the hardest and most important part of building an experimentation program.
The most dangerous experiment infrastructure failure is one that produces numbers that look plausible but are wrong. Missing exposure logs, overlapping experiments, and assignment / logging coupling are the three failure modes that invalidate results silently.
Statistical significance tells you the direction is probably real. It does not tell you the effect matters, is durable, applies to all segments, or outweighs the cost of shipping. The analysis phase is where rigorous experiments are turned into bad decisions — or good ones.
Stopping an experiment early because results look promising inflates false positive rates — you are peeking. The required sample size is calculated before launch and documented. "It looks like it's winning" is not a stopping criterion. Neither is pressure from a product manager or an upcoming board meeting.
A statistically significant +0.1% lift on a 2% baseline conversion is not a shipping decision — it is noise that cleared a threshold. Always ask: if this effect is real and durable, does it move a business needle we care about? Significance answers "is it real?" — effect size answers "does it matter?"
Running 30 segments after results are in — by country, browser, cohort, plan tier — and declaring a win because one segment shows p < 0.05 is p-hacking. Segmentation generates hypotheses for the next experiment. It does not validate winners. Pre-specify any segments of interest in the experiment brief before launch.
Users often over-engage with anything new. An experiment showing +15% engagement in week 1 that decays to +2% by week 4 is a novelty effect, not a feature win. For engagement-sensitive metrics, run the experiment long enough to see the curve flatten. Shipping novelty as product value is one of the most common false wins in SaaS growth.
Each experiment is a small bet with bounded downside and potentially compounding upside. The product is not a single feature decision — it is the integral of all the experiment results that shipped over time. Velocity without rigor is noise; rigor without velocity is too slow.
Qualitative user research, session recording, funnel analysis, and cohort comparison answer many questions faster and cheaper than a controlled experiment. A/B tests are for establishing causality when confounders matter. Use the cheapest method that answers the question with sufficient confidence.
Every experiment — wins, losses, and nulls — is documented with hypothesis, methodology, result, and decision taken. This institutional knowledge compounds: future experiments reference past learnings, teams avoid re-running failed ideas, and every new engineer inherits the collective knowledge of every experiment ever run. A dashboard without a searchable repository is amnesia with good charts.
The bottleneck in most experimentation programs is not statistical — it is operational: weeks spent on instrumentation, flag setup, stakeholder alignment, and data pipeline delays. Every day of cycle time reduction is a permanent increase in experiment velocity. Treat experiment infrastructure with the same engineering investment as production infrastructure.
Experiments are bets, not proof.
Define the metric before you look at results.
The peeking problem is not a choice — it is a bias.
In B2B, the account is the unit. Not the user.
Null results are the most honest signal you will ever get.