Samuel MartinExperiments and measurement systems

A $99 Experiment That Made Outbound a Revenue Product

Samuel Martin designed a randomized controlled trial that settled a two-year argument about whether outbound engagement paid for itself. Treatment booked appointments 3.45x more often than control, at $99 cost per acquisition against $549 of first-year customer value. The executive team stopped treating outbound as an operating expense and funded a $1.1M scaling roadmap.

Randomized Controlled TrialIntent-to-Treat AnalysisUnit EconomicsSensitivity AnalysisShipped to production

In one paragraph

Samuel Martin ran a randomized controlled trial on outbound engagement and measured a 3.45x lift in appointment booking conversion at $99 cost per acquisition, against $549 of first-year customer value. That gap turned a cost center into a revenue line at roughly 80% margin. The $1.1M scaling roadmap that followed was sensitivity tested before it was presented, and it was approved on the strength of the interval rather than the point estimate.

3.45x
lift in booking conversion
$99
cost per acquisition
$549
first-year customer value
$1.1M
scaling roadmap funded
80%
approximate margin on the line

01 / The spine

Four steps, in order.

The decision on the table

Finance wanted to cut the outbound team. Growth wanted to double it. Both sides argued from the same dashboard, which reported bookings by rep and said nothing about what outbound actually caused. The decision was a budget line, and nobody could defend either number.

What the measurement got wrong

The existing number was a correlation wearing a result's clothes. Contacts who picked up the phone were likelier to book, and they were also likelier to book with no outbound at all. Attribution ran on last touch, so outbound collected credit for demand it did not create.

What Samuel Martin did

Randomized 4,812 contacts to treatment or holdout at the contact level, stratified by segment and prior engagement. Analysis was intent-to-treat, so a contact who never picked up stayed in the treatment arm and the estimate stayed honest. Unit economics were built from the trial's own cost per contact, then sensitivity tested across pickup rate, close rate, and first-year value.

What changed

Booking conversion moved from 1.6% to 5.5%, a 3.45x change. In business terms, roughly $450 of first-year value per acquisition after the $99 cost. Outbound was reclassified from operating expense to revenue product at roughly 80% margin, and the $1.1M scaling roadmap was funded in the same meeting.

02 / The result

Treatment booked appointments 3.45x more often than control.

Appointment booking conversion by arm. Bands are 95% confidence intervals.

Control
n = 2,406
1.6%
Outlined bands are 95% confidence intervals.
Treatment
n = 2,406
5.5%
3.45x more bookings per contact assigned
0
2%
4%
6%
8%
BOOKING CONVERSION RATESYNTHETIC DATA

Treatment booked at 5.5% against control's 1.6%, a 3.45x lift whose confidence intervals do not overlap. Intent-to-treat analysis of 4,812 contacts randomized at the contact level, stratified by segment and prior engagement, with a 30-day attribution window and a two-sided test at alpha 0.05.

Rates are shown to one decimal; the 3.45x figure is computed on unrounded values. Figures redrawn on synthetic data. No client named. Method and reasoning are exact.

The interval is the part that mattered in the room. A point estimate of 3.45x invites an argument about whether the number is real. A lower bound that still clears the cost per acquisition ends the argument, which is why the roadmap was approved on the interval rather than the headline.

"The trial did not tell us outbound worked. It told us which contacts it worked on, and that was the number finance could act on."

03 / The method, in full

The depth is on this page, not behind a link.

Each heading states its own conclusion, so nothing below requires opening to be understood. Open one when you want the mechanics.

Randomization was at the contact level, and the balance check found nothing worth adjusting for.

Assignment was stratified on two variables that plausibly drive booking on their own: client segment and whether the contact had engaged in the prior 90 days. Randomization ran inside each stratum, which keeps the arms comparable on the two things most likely to confound the estimate.

Contact level, not rep level. Rep-level assignment is easier to run and it leaks: reps talk, scripts spread, and the holdout stops being a holdout by week three. Contact-level assignment costs more coordination and buys a clean comparison.

Post-randomization balance across segment, tenure, prior engagement, and account size showed no difference large enough to adjust for. Samuel Martin reported the balance table alongside the result rather than only the result, because a reviewer's first question is whether the arms started even.

Intent-to-treat was the honest estimator, and per-protocol would have inflated the lift by roughly a third.

Only 62% of treated contacts were ever reached. The tempting analysis compares reached contacts against the full holdout, and it is wrong, because being reachable is itself a predictor of booking. That comparison measures reachability and calls it outbound.

Intent-to-treat keeps every assigned contact in its original arm whether or not anyone picked up. The result is a deliberately conservative estimate of what the program does when you fund it, which is the question the budget owner is actually asking.

The per-protocol cut came out near 4.6x. Samuel Martin showed both numbers and led with 3.45x. Presenting the larger figure would have been easy and would have collapsed the first time someone re-ran it.

The roadmap survives a 30% haircut on close rate and a 20% haircut on first-year value.

Unit economics were built from the trial's own costs rather than a planning assumption: loaded cost per contact attempt, attempts per reach, reaches per booking, bookings per acquisition. That chain produces the $99 cost per acquisition, and every link in it was measured in the trial.

Sensitivity ran over three inputs at once. Pickup rate breaks the case below 11%, roughly a third under what the trial observed. Close rate can fall 30% and first-year value can fall 20% before the program stops clearing its own cost.

The roadmap went to the executive team with the break-even points stated up front. Naming the conditions under which the plan fails is what got the $1.1M approved, and it took less time to present than the lift itself.

04 / Limits

What this does not show, and what I would do differently.

Limits

The trial measures booking, not retention. A 30-day window cannot tell you whether an outbound-sourced customer stays as long as an inbound one, and the $549 first-year value is carried over from existing cohort data rather than measured inside the trial.

One segment, one quarter, one market. The lift is a fact about this population and not a constant.

What I would do differently

Pre-register the sensitivity grid. Samuel Martin chose the three inputs after seeing the lift, which is defensible here and still looks better declared in advance.

Keep a permanent 5% holdout after rollout. The trial answered the question once, and the answer drifts. Without a standing holdout there is no way to notice when it does.