Measuring a Feature Without Fooling Yourself
Product measurement fails in a predictable direction: the launch number is flattering, the retention number arrives too late to matter, and nobody separates the two out loud. Samuel Martin designs the pair up front and labels which is which.
This surface re-sequences the same four case studies and swaps the sentence above it. No proof point here is unique to this page, and none of the four surfaces is the real one.
Causal inference first, then the event layer it depends on.
Nine months later, the real gain was 7.0 points
A 30-day randomized test scoped explicitly to leading indicators, paired with a nine-month difference-in-differences study on monthly cohorts. Treatment gained 10.0 points and control gained 3.0, so the difference in differences is 7.0. The three pre-launch parallel months are the load-bearing part.
Financial reporting went from weeks to minutes
Causal work on product events needs an event layer that can be trusted. Making each join emit a survival count turned an opaque pipeline into four readable numbers and located a 38% loss the organization had been attributing to the wrong step.
A $99 experiment that made outbound a revenue product
A clean randomized trial when randomization is available: 4,812 contacts randomized at the contact level, intent-to-treat analysis, 5.5% against 1.6%, confidence intervals that do not overlap.
One average became four peer groups
Cohort comparison done properly. Gower’s distance on mixed-type attributes rather than Euclidean distance on dummy variables, and medoids rather than means, so every group has a real member at its center.
What Samuel Martin would own here.
The experiment and quasi-experiment design, the event taxonomy underneath it, and the reporting discipline that keeps a leading indicator from being quoted as an outcome six months later.
Figures across this site are redrawn on synthetic data, no client is named, and the method and reasoning are exact.