Architecture Reference // 2026

Experimentation

The statistical, organizational, and infrastructure principles behind rigorous product experimentation in modern SaaS.
8 domains  ·  27 rules.

HYP
→
RUN
→
ANA
→
ACT

HYPOTHESIZE  ·  INSTRUMENT  ·  ANALYZE  ·  DECIDE

01 // Foundation

Experiments are bets, not tests

An experiment does not prove your idea was right — it updates your belief about how users behave. The goal is not to validate intuition. The goal is faster learning under uncertainty, including the learning that your intuition was wrong.

Null results are results — treat them as such
A well-powered experiment that shows no significant effect tells you something important: your intervention didn't move the metric. That is a decision. Ship it, kill it, or redesign based on the mechanism. A null result left running indefinitely is not caution — it is avoidance.
The HIPPO problem kills experimentation culture
When the Highest Paid Person's Opinion overrides experiment results, you have data theater — the appearance of a data-driven culture without the substance. Experiment results must be binding within pre-defined decision criteria, or the entire system loses credibility and engineers stop believing in the process.
Experiment velocity is a compounding competitive advantage
A team running 200 experiments per year learns faster than a team running 20 with the same rigor. The fastest-learning teams are not more reckless — they have invested in reducing cycle time: hypothesis formulation, flag setup, instrumentation, and decision-making. Reduce cycle time, not statistical standards.
02 // Hypothesis Design

A hypothesis without a mechanism is a wish

The Because clause is the most important part of any hypothesis. Without it, a flat result teaches you nothing. With it, a flat result tells you the mechanism you believed in was wrong — which is a more valuable learning than a small positive effect with no explanation.

01
Structure every hypothesis: If / Then / Because
"If we add a progress bar to onboarding, then 7-day retention will increase by 8%, because users who complete onboarding have 3x better 90-day retention and the progress bar removes ambiguity about next steps." The Because is the falsifiable mechanism. Without it, you can't learn from a flat result.
02
Define the primary metric and ship criteria before running
The primary metric, the minimum detectable effect, and the decision threshold are written into the experiment brief before any data is collected. Defining success after seeing results is HARKing — Hypothesizing After Results Known. It produces the answer you wanted, not the answer that is true.
03
Identify guardrail metrics upfront
Guardrail metrics are metrics you must not hurt, even if the primary metric improves. A feature that increases sign-ups 20% while increasing support volume 40% is not a win. Guardrails make trade-offs explicit before the experiment runs — not visible after you declare a winner.
Hypothesis template
IF [change]
THEN [metric] WILL [direction] BY [MDE]
BECAUSE [mechanism]
GUARDRAILS [metrics that must not worsen]
SHIP IF [decision criteria]
03 // Statistical Foundations

Most teams get the statistics wrong in the same four ways

Statistical validity is not optional — it is the only thing separating a controlled experiment from a coincidence. The four errors below are not edge cases. They are the default behavior of teams without explicit statistical discipline.

The peeking problem — the most common error

Looking at results daily and stopping when p < 0.05 inflates your false positive rate dramatically. You are running multiple significance tests on the same data. Use sequential testing methods (always-valid p-values, mSPRT) or fix your sample size at launch and commit to analyzing at that point only.

Power analysis before you run, not after

Statistical power (typically 80%) is the probability of detecting a true effect if one exists. A test at 50% power misses half of real effects. Power is set by sample size, MDE, and baseline variance. An underpowered experiment that shows no effect has told you nothing — you couldn't have detected it regardless.

Multiple comparisons inflate false positive rates

Testing 20 metrics and declaring a win when any one reaches p < 0.05 gives roughly a 64% chance of a false positive. Apply Bonferroni correction or FDR control (Benjamini-Hochberg) when testing multiple metrics. Designate one primary metric — all others are exploratory until confirmed in a follow-up experiment.

Confidence intervals, not just p-values

A p-value tells you whether an effect is distinguishable from zero. A confidence interval tells you the plausible range of the effect size. p = 0.04 with CI [+0.01%, +2.3%] and p = 0.04 with CI [+0.8%, +1.4%] are completely different decisions. Always report both.

04 // Randomization

The unit of randomization is the most consequential design decision

Get the randomization unit wrong and no amount of statistical rigor downstream will save you. This is the decision most Growth teams make implicitly — and most often get wrong in B2B SaaS, where users are not independent of each other.

In B2B SaaS, randomize at the account level — not the user level
Users within the same account share workspaces, collaborate, and talk to each other. Randomizing by user within a multi-seat account contaminates the assignment — control-group users learn about the treatment from colleagues. This is a SUTVA violation and invalidates the experiment. Cluster randomize at the account or team level, even though it requires significantly larger sample sizes.
The unit of randomization must match the unit of the metric
If your metric is "conversion per user," randomize by user_id. If your metric is "revenue per account," randomize by account_id. Mismatching the unit inflates variance and produces unreliable confidence intervals — you will see more false positives and larger apparent effects than are real.
Sticky, deterministic assignment via hashing
Use hash-based assignment: hash(user_id + experiment_id + salt) mod 100 maps each user to a stable bucket permanently. A user re-evaluated randomly on each page load is simultaneously in both control and treatment, rendering all metrics meaningless. Assignment must be stable across devices, sessions, and time.
Always run a Sample Ratio Mismatch check first
SRM occurs when the actual split between control and treatment doesn't match the intended split — you targeted 50/50 but got 52/48. This indicates a bug in assignment, logging, or analysis-time filtering. It invalidates the results entirely. SRM detection is a prerequisite check before any metric analysis, not an optional audit.
05 // Metric Design

Measuring the wrong thing with precision is still measuring the wrong thing

A metric that is valid but insensitive will never move. A metric that is sensitive but invalid will move in the wrong direction while the business suffers. Metric design is the hardest and most important part of building an experimentation program.

Validate surrogate metrics against the North Star before using them
Using "email open rate" as a proxy for retention, or "time on page" as a proxy for satisfaction, is only valid if you have empirical evidence they correlate — measured from your own product, not from published benchmarks. A surrogate metric that is not predictive of your North Star produces experiment wins in your dashboard while the business goes sideways.
Distinguish North Star, primary, guardrail, and diagnostic metrics
North Star is the long-term business outcome (revenue, retention). Primary is the metric your experiment directly moves (activation rate, feature adoption). Guardrail is what you must not hurt (support volume, latency, NPS). Diagnostic is the mechanism signal (click-through rate, funnel step completion). One primary metric per experiment — not four.
Ratio metrics require the delta method, not a naive t-test
Metrics like "revenue per user" or "errors per session" have both numerator and denominator variance. A naive t-test on the ratio underestimates variance significantly and produces false positives. Use the delta method or bootstrap sampling for confidence intervals on ratio metrics. Most experimentation platforms handle this correctly — if yours doesn't, it's a platform risk.
06 // Infrastructure

Bad instrumentation produces correct-looking wrong results

The most dangerous experiment infrastructure failure is one that produces numbers that look plausible but are wrong. Missing exposure logs, overlapping experiments, and assignment / logging coupling are the three failure modes that invalidate results silently.

01
Log exposures — not just conversions
An experiment that only logs conversions (the numerator) loses its denominator — you don't know how many users were assigned to the experiment without converting. Log an exposure event at the exact moment a user is assigned and first sees the treatment. This is the single most common infrastructure gap in early-stage experimentation programs.
02
Experiments must be mutually exclusive by default
A user in two overlapping experiments is being exposed to an interaction effect you didn't design. Use experiment namespaces or traffic layers to enforce mutual exclusivity by default. Factorial designs — where interaction effects are intentional — are the explicit opt-in, not the default. Most teams discover this gap the first time two experiments run simultaneously on the same feature surface.
03
Use holdbacks to measure long-term and compound effects
A holdback group — 1–5% of users kept on the old experience even after 100% rollout — lets you measure the true long-term impact of a feature after the novelty effect has decayed. It also provides a clean baseline for measuring the compound effect of multiple sequential experiments on the same user population.
07 // Analysis & Decisions

A statistically significant result is not a shipping decision

Statistical significance tells you the direction is probably real. It does not tell you the effect matters, is durable, applies to all segments, or outweighs the cost of shipping. The analysis phase is where rigorous experiments are turned into bad decisions — or good ones.

Never analyze before the pre-specified sample size

Stopping an experiment early because results look promising inflates false positive rates — you are peeking. The required sample size is calculated before launch and documented. "It looks like it's winning" is not a stopping criterion. Neither is pressure from a product manager or an upcoming board meeting.

Effect size matters more than significance

A statistically significant +0.1% lift on a 2% baseline conversion is not a shipping decision — it is noise that cleared a threshold. Always ask: if this effect is real and durable, does it move a business needle we care about? Significance answers "is it real?" — effect size answers "does it matter?"

Post-hoc segmentation is exploratory, not confirmatory

Running 30 segments after results are in — by country, browser, cohort, plan tier — and declaring a win because one segment shows p < 0.05 is p-hacking. Segmentation generates hypotheses for the next experiment. It does not validate winners. Pre-specify any segments of interest in the experiment brief before launch.

Check for the novelty effect before shipping

Users often over-engage with anything new. An experiment showing +15% engagement in week 1 that decays to +2% by week 4 is a novelty effect, not a feature win. For engagement-sensitive metrics, run the experiment long enough to see the curve flatten. Shipping novelty as product value is one of the most common false wins in SaaS growth.

08 // Culture & Velocity

The companies that win run the most experiments per unit of time

Each experiment is a small bet with bounded downside and potentially compounding upside. The product is not a single feature decision — it is the integral of all the experiment results that shipped over time. Velocity without rigor is noise; rigor without velocity is too slow.

1

Not every question needs an A/B test

Qualitative user research, session recording, funnel analysis, and cohort comparison answer many questions faster and cheaper than a controlled experiment. A/B tests are for establishing causality when confounders matter. Use the cheapest method that answers the question with sufficient confidence.

2

Build an experiment repository, not just a dashboard

Every experiment — wins, losses, and nulls — is documented with hypothesis, methodology, result, and decision taken. This institutional knowledge compounds: future experiments reference past learnings, teams avoid re-running failed ideas, and every new engineer inherits the collective knowledge of every experiment ever run. A dashboard without a searchable repository is amnesia with good charts.

3

Invest in reducing cycle time, not just sample size

The bottleneck in most experimentation programs is not statistical — it is operational: weeks spent on instrumentation, flag setup, stakeholder alignment, and data pipeline delays. Every day of cycle time reduction is a permanent increase in experiment velocity. Treat experiment infrastructure with the same engineering investment as production infrastructure.

Experiment repository — every entry contains
hypothesis (IF/THEN/BECAUSE)
primary metric + MDE
result + confidence interval
decision taken
guardrail results
follow-up experiments spawned
owner

The mental model

Experiments are bets, not proof.
Define the metric before you look at results.
The peeking problem is not a choice — it is a bias.
In B2B, the account is the unit. Not the user.
Null results are the most honest signal you will ever get.

Daniel Brasileiro