Skip to main content

A/B testing framework

A/B Testing Framework: From Research to a Trustworthy Decision

A repeatable process for teams that want more than a backlog and a significance screenshot: a decision system that preserves the evidence behind each outcome.

By Atticus Li · Updated July 28, 2026

The direct answer

A strong experimentation framework separates four jobs: diagnosing a problem, researching a mechanism, estimating a causal effect, and deciding what to do next. They reinforce each other, but they are not interchangeable proof.

01

Diagnose the problem

Use funnel, event, cohort, error, and segment data to locate a decision-relevant problem. Check the event definition and instrumentation before treating a drop as customer friction.

02

Explain the mechanism

Use interviews, usability sessions, support themes, surveys, session review, or other named research methods to explain who is affected and why. Record the method, audience, and limitations.

03

Write the decision rule

Before exposure, document the hypothesis, control, variant, primary metric, guardrails, baseline, MDE, sample plan, allocation, analysis method, exclusions, and stopping rule.

04

Run with an execution log

Check exposure and instrumentation, calculate SRM against planned allocation, and log traffic, product, campaign, targeting, and implementation changes while the test runs.

05

Analyze both uncertainty and practical value

Report arm-level counts, effect estimate, interval or decision metric, primary result, guardrails, and limitations. Do not replace this with a winner label.

06

Make and preserve the decision

State what shipped, stopped, or needs another test; explain why; tag the evidence tier; and save the learning so the next planning conversation starts with evidence.

What good documentation contains

Before launchDuring the runAfter analysis
Problem evidence, hypothesis, treatments, metrics, baseline, MDE, sample, allocation, analysis plan, stopping rule.Actual exposure, SRM, instrumentation QA, ramps, campaigns, releases, outages, targeting changes, concurrent experiments.Arm-level data, effect and uncertainty, guardrails, practical significance, decision, learning, limitations, and evidence tier.

Choose the analysis family before you look

A fixed-horizon frequentist plan, a pre-defined group-sequential design, an anytime-valid procedure, and a Bayesian sequential decision rule can all be defensible. The failure is changing the rule after the result is visible. Save the method, thresholds, and stopping rule with the experiment so future reviewers can reproduce the decision logic.

Put the framework into practice

Use the free record to turn the process into a usable planning and reporting habit.

Open the A/B test template