Nine Months Later, the Real Gain Was 7.0 Points
Samuel Martin paired a 30-day A/B test for leading indicators with a nine-month difference-in-differences study for outcomes, so an AI chat feature would be judged on whether people stayed rather than on whether they tried it. The nine-month answer was smaller than the launch enthusiasm and durable: treatment cohorts gained 7.0 percentage points more 30-day retention than control, with a 95% interval of 3.4 to 10.6 points.
Published 2026-07-25. Last updated 2026-07-25. Roughly an eight minute read.
In one paragraph
Samuel Martin designed the measurement for an AI chat feature as two studies rather than one. A 30-day randomized test carried the leading indicators, which is all a 30-day window can honestly support. A nine-month difference-in-differences study on monthly cohorts carried the outcome. Treatment cohorts moved from 37.0% to 47.0% 30-day retention while control moved from 36.0% to 39.0%, a difference in differences of 7.0 percentage points with a 95% interval of 3.4 to 10.6. The three pre-launch months, in which the cohorts moved together, are the part of the design that makes the number believable.
- +7.0 pts
- difference in differences, 30-day retention
- +10.0 pts
- treatment, 37.0% to 47.0%
- +3.0 pts
- control, 36.0% to 39.0%
- 3 + 9
- pre-period and post-period months
- 30 days
- paired A/B test for leading indicators
01 / The spine
Four steps, in order.
An AI chat feature had shipped to a subset of users and the launch review was scheduled for the following month. The organization wanted to know whether to fund the roadmap behind it. A month is long enough to measure engagement and far too short to measure retention, and the review was going to happen anyway.
The default plan was a single 30-day test read as an outcome study. That design fails in a predictable direction. Early engagement with a novel feature is inflated by novelty and by the self-selection of users who seek new things out, so a 30-day number is a reliable measure of trial and a poor proxy for retention.
Split the question in two and said so in advance. The 30-day randomized test was scoped to leading indicators only and reported as such. A difference-in-differences design on monthly cohorts carried the outcome, using three pre-launch months to establish that treatment and control moved together before the feature existed and nine post-launch months to measure where they diverged.
The nine-month estimate was 7.0 percentage points of additional 30-day retention, with an interval that clears zero comfortably. The roadmap was funded on that number rather than on the launch figures, and the pairing became the default shape for feature measurement: leading indicators fast, outcomes slow, neither one pretending to be the other.
02 / The result
Treatment gained 7.0 points more than control over the nine months after launch.
30-day retention by monthly cohort. The 95% interval is stated on the difference-in-differences estimate, not on either series.
The treatment cohort gained 7.0 percentage points more than control over the nine months after launch, 95% interval 3.4 to 10.6 points, and the two cohorts moved together for the three months before it, which is the parallel-trends assumption the design rests on. Difference in differences on monthly cohorts, three pre-period months and nine post-period months, paired with a 30-day A/B test for leading indicators.
The vertical axis is framed from 30 to 50 percent rather than zero: the treatment series moves 10 points on a base near 37, and a zero baseline hides it. Lines may be framed, bars may not. The interval is placed on the difference in differences rather than drawn as a band on either series, because the estimate depends on all four cohort-period values and a band on one series would imply the other was measured without error. Figures redrawn on synthetic data. No client named. Method and reasoning are exact.
The three months to the left of the marker do more work than the nine to the right. They are the evidence for parallel trends, and without them the divergence after launch is just two lines that happen to differ. A reader who wants to attack this figure should attack the pre-period, which is why it is drawn at the same weight as the rest.
“A 30-day test tells you whether people tried it. Reading that number as retention is how a team talks itself into a feature, and it is a mistake you make on purpose.”
03 / The method, in full
The depth is on this page, not behind a link.
Each heading states its own conclusion, so nothing below requires opening to be understood. Open one when you want the mechanics.
The two cohorts moved together for three months before launch, which is the assumption the whole design rests on.
Difference in differences does not require the two groups to be at the same level. It requires them to have been moving in parallel, so that the control series is a credible answer to what the treatment series would have done untreated. Treatment sits about a point above control before launch and the two move together, which is exactly the shape the design needs.
Three pre-period months is thin. Samuel Martin reported it as thin rather than presenting it as settled, and treated the pre-period as the figure's weakest point rather than its strongest. A seasonal pattern with a period longer than three months would not be visible here.
The alternative was a synthetic control built from unexposed segments, which is more robust and much harder to explain to the people approving the roadmap. The parallel-trends check is legible in one glance at the chart, and legibility was worth more than the marginal robustness here.
The 30-day test and the nine-month study answer different questions, and running only one of them fails in a predictable direction.
The 30-day randomized test measured leading indicators: whether users engaged with the feature, how often, and whether engagement held across the month. That is a genuine question and the test answers it well. It cannot answer whether the feature keeps anyone.
Novelty and self-selection both inflate early engagement with a new surface. Users who adopt an AI feature in its first month are not a random sample of users, even inside a randomized arm, because assignment controls exposure and not curiosity.
Reporting the two studies as a pair, with the 30-day results labelled leading indicators in the document itself, was the part that changed behavior. The failure mode is not a bad 30-day number. It is a good 30-day number quoted six months later as though it had measured retention.
The interval sits on the difference in differences, not on either series, because the estimate depends on all four values.
A shaded band on the treatment series alone is the conventional treatment and it misleads twice. It implies the control series is measured without error, and it draws uncertainty on a quantity that is not the finding.
The finding is the difference of two differences, built from four cohort-period values, and it carries its own standard error. Stating that interval next to the estimate puts the uncertainty on the number a reader will quote.
The axis is framed from 30 to 50 percent rather than from zero, and that choice is disclosed in the caption. A 10-point move on a base near 37 disappears against a zero baseline. Lines may be framed this way because they encode position; bars may not, because they encode length from zero.
04 / Limits
What this does not show, and what I would do differently.
Limits
This measures 30-day retention, not revenue, and not long-run retention. A cohort that stays 30 days longer is worth something, and this study does not say how much.
Cohort composition drifts over nine months. Acquisition mix changed across 2025, and while that pressure applies to both arms, difference in differences only removes it to the extent that it applies to both equally.
What I would do differently
Extend the pre-period before launch rather than after. Three months was what the calendar allowed, and six would have made the parallel-trends claim substantially harder to argue with for no additional analytical work.
Pre-register the outcome window. Nine months was chosen because it was the horizon leadership cared about, which is defensible, and it would look considerably better declared in advance than selected once the series were in hand.