Field note · causal inference
A randomized trial can answer a well-defined causal question. Before proposing one, check whether you can run it and what the result would actually tell you.
01
I prefer the full term, randomized controlled trial, because it reminds us what the design requires: random assignment, a control group, and a specific change to evaluate.
The question here is whether the exposure changed the conversion rate, allowing for chance variation between the groups.
02
When a decision is uncertain, asking for a test is understandable. The difficulty is that some decisions do not give you a workable treatment and control group.
“Can we just A/B test it?” — asked about a decision that cannot be A/B tested
03
Sending half your users somewhere different is only part of the design. Decide what you will measure and how you will evaluate it before looking at the results.
An unreliable test
A trial
MDE is the minimum detectable effect: the smallest lift worth shipping, say +1.5 points on a 5% baseline. Choose it first. It sets your sample size, and it is what separates real from worth it: an effect can be statistically significant and still too small to justify the work.
04
Wall 1
You cannot randomize the entire decision to launch a new product into two versions of the same market. You may be able to test parts of the launch, but those tests answer narrower questions.
05
Wall 2
Purchase conversion may be the outcome you care about, but a low purchase rate can make the required sample impractical.
Take 2,000 visitors a week and a 2% purchase rate. Detecting a 10% lift (an MDE of 0.2 points) needs about 78,400 users per group.
Users needed per group, for a 10% lift
Eighteen months. The feature gets rewritten twice, and whoever asked for the test has changed teams.
06
So narrow the question. Make the conversion event an earlier step in the journey, one with the traffic to support a real test. That is the second bar: six weeks instead of eighteen months.
A user journey
This test estimates the effect on checkout rate. It does not establish the effect on purchases or repeat business; those need additional evidence.
State that limit alongside the result. Otherwise, an improvement in checkout rate can easily be repeated as evidence of higher revenue.
07
Depending on the setting, these designs may help. Some can use randomization; others rely on assumptions about the comparison group or what would have happened over time.
Discuss the evaluation before launch. You may need to preserve a comparison group, schedule a rollout, or collect a baseline that cannot be reconstructed later.
The job
Say what the test can answer.
If the test cannot support the claim, say so before the result goes into the presentation. Explain what you did learn and what remains uncertain.