Samuel MartinExperiments and measurement systems

Four figure bodies, one frame.

Every figure on this site renders through FigureFrame, which owns the card, the title, the light field, the caption, the accessible name, and the print selector. Each body owns only its marks. Changing S-35, S-36, or S-17 means editing one file.

The frame also publishes the in-frame type roles and two non-data ink tokens. A body cannot author a baseline, a tick, or a gridline at data weight, because the frame paints anything marked as non-data itself.

Body 1 / Two-group comparison with uncertainty · Outbound RCT

Treatment booked appointments 3.45x more often than control.

Appointment booking conversion by arm. Bands are 95% confidence intervals.

Control
n = 2,406
1.6%
Outlined bands are 95% confidence intervals.
Treatment
n = 2,406
5.5%
3.45x more bookings per contact assigned
0
2%
4%
6%
8%
BOOKING CONVERSION RATESYNTHETIC DATA

Treatment booked at 5.5% against control's 1.6%, a 3.45x lift whose confidence intervals do not overlap. Intent-to-treat analysis of 4,812 contacts randomized at the contact level, stratified by segment and prior engagement, with a 30-day attribution window and a two-sided test at alpha 0.05.

Rates are shown to one decimal; the 3.45x figure is computed on unrounded values. Figures redrawn on synthetic data. No client named. Method and reasoning are exact.

Body 2 / Stage bars, not a tapered funnel · Attribution warehouse

Two of every five conversion events lost their source at the CRM join.

Conversion events surviving each join, one quarter, four source systems. Common zero baseline.

Ad platform events
482,000
Everything the four ad platforms reported.
Matched to a session
411,000
15% lost, mostly cross-device.
Joined to a CRM lead
254,000
38% lost here, the single largest drop in the pipeline.
Reconciled to finance
249,000
2% lost, which is the step everyone assumed was the problem.
0
125,000
250,000
375,000
500,000
CONVERSION EVENTS SURVIVING EACH JOINSYNTHETIC DATA

Two of every five conversion events lost their source at the CRM join, and the finance reconciliation everyone blamed lost almost nothing. Counts are conversion events over one quarter across four source systems, measured after the unified warehouse was in place.

An event is counted at a stage only if it carries the keys the next join requires, so the bars are survival counts and not a tapered funnel. Figures redrawn on synthetic data. No client named. Method and reasoning are exact.

Body 3 / Time series with intervention marker · AI chat measurement

Treatment gained 7.0 points more than control over the nine months after launch.

30-day retention by monthly cohort. The 95% interval is stated on the difference-in-differences estimate, not on either series.

50%
45%
40%
35%
30%
AI chat launches, Apr 2025
Treatment cohort
Control cohort
Difference in differences: +7.0 points (95% CI 3.4 to 10.6)
Jan 2025
Jun 2025
Dec 2025
Treatment cohort
Control cohort
AI chat launches, Apr 2025
Difference in differences: +7.0 points (95% CI 3.4 to 10.6)
30-DAY RETENTION BY MONTHLY COHORTSYNTHETIC DATA

The treatment cohort gained 7.0 percentage points more than control over the nine months after launch, 95% interval 3.4 to 10.6 points, and the two cohorts moved together for the three months before it, which is the parallel-trends assumption the design rests on. Difference in differences on monthly cohorts, three pre-period months and nine post-period months, paired with a 30-day A/B test for leading indicators.

The vertical axis is framed from 30 to 50 percent rather than zero: the treatment series moves 10 points on a base near 37, and a zero baseline hides it. Lines may be framed, bars may not. The interval is placed on the difference in differences rather than drawn as a band on either series, because the estimate depends on all four cohort-period values and a band on one series would imply the other was measured without error. Figures redrawn on synthetic data. No client named. Method and reasoning are exact.

Body 4 / Small-multiple grid · Peer benchmarking

Against its real peer cluster the client sits at the median, not below average.

90-day activation rate, eight clients per cluster, one shared scale across all four panels.

High-touch enterprise
8 clients
60%
30%
0
Self-serve mid-market
8 clients
This client
Seasonal
8 clients
Book average 33%
Regulated
8 clients
The labelled dashed line is the book-of-business average, 33%, drawn in every panel. All four panels share one 0 to 60% scale, so bar heights are comparable across panels.
Mid-pack in its own cluster, below an average that mixes four different businesses.
90-DAY ACTIVATION RATE, SHARED SCALESYNTHETIC DATA

The client sits at the median of its own peer cluster while falling below the book-of-business average, which is why the average was the wrong comparison. Clusters were built on mixed-type client attributes with Gower's distance and partitioning around medoids, with k=4 chosen on silhouette width.

All four panels share one 0 to 60% scale, so heights are comparable across panels. Figures redrawn on synthetic data. No client named. Method and reasoning are exact.