The Sample Size of One
You changed the pricing page and MRR went up. Was it you, or was it March? What 520 controlled measurements on synthetic companies with known planted effects say about the numbers founders trust.
You changed the pricing page on 3 March. MRR is up 9% since. Was it the pricing page, or was it March? Issue 001 told you what to do. This one asks how you would know it worked. We planted effects of known size in synthetic companies and measured them with the standard causal estimator. On a compounding metric like MRR, a single company's own history recovered a planted 10% lift 0 times out of 10. Add one comparable outside series and a planted 5% lift goes from 0.00 detection to 0.70. The problem is not your analysis. It is that one company is not enough companies.
The noise
Start with what a metric does when nothing happens. We ran the shipped trend test, Mann-Kendall with the standard autocorrelation correction, over series with no effect planted. Nominal false-positive rate: 5%. 2,500 to 3,000 trials per cell.
| Points in window | Persistence 0.0 | Persistence 0.5 | Persistence 0.8 |
|---|---|---|---|
| 24 (monthly, two years) | 8.4% | 23.0% | 34.7% |
| 90 (daily, one quarter) | 9.1% | 17.3% | 27.6% |
| 180 (daily, six months) | 10.0% | 14.7% | 25.0% |
Persistence is how much of yesterday a metric carries into today. A daily rate sits near 0.5. A level that accumulates, cash or subscribers or MRR, sits near 0.8. At that end the test invents a trend in a quarter of quiet windows, and more data barely helps: 90 daily points to 180 moves it from 27.6% to 25.0%.
On compound growth it stops being a rate. The detector fires on 100% of clean compound-growth series (n=200), because the baseline itself trends and "the line went up" was never the question.
Now count your metrics. A 13-metric batch with 12 quiet metrics produced 57 signals. False-discovery correction removed none of them, because each had already cleared p of 0.01 before the correction ran. On the measured per-metric rate, 12 quiet metrics is 12 x 0.355 = 4.3 false trends per pass. Weekly.
The experiment
Synthetic companies. Effects of known size planted on known dates. Half the runs quiet. The real production estimator, Bayesian structural time-series, scored against the truth we planted. 360 fits in the core run, 160 in the control-series run.
A single company detects almost nothing on a compounding metric. In the core run, a planted +10% lift was recovered 0 times out of 10. The control-series run put both arms on the same response series:
| Planted effect | Alone | With a control series |
|---|---|---|
| +5% | 0 of 20 | 14 of 20 |
| +10% | 2 of 20 | 18 of 20 |
| +20% | 3 of 20 | 20 of 20 |
Same response series in both arms. What changed is the counterfactual: instead of extrapolating the company's own slope, the model watched what comparable companies did over the same weeks.
Two things that complicate the good news, both measured:
- The control version false-flags 6 of 20 quiet periods, against 0 of 20 for the single-company arm. Control series buy identification and hand back a calibration problem.
- When the single-company path does fire on a compounding metric, it oversizes: 2.53x the planted truth at a planted 5%, 1.27x to 1.40x at 10% to 30%.
The counterexample matters. On a mean-reverting daily rate, one company's history detected every planted effect from 5% to 30%, 10 of 10, sized correctly. Shape of the metric, not quality of the analysis. Even there the interval is not trustworthy: 95% intervals covered the true null 59.2% of the time.
The ceiling
Monthly data has a wall that no method choice gets around. At 24 points on a sticky series, the shipped test false-flags 34.7% of quiet windows. The one correction that holds its size drops detection power to 12.3%. That is the whole menu.
It is not a data-volume problem. 24 monthly points detect an 8% per year drift; 90 daily points need 30% per year. Span buys detection, not sample count. What monthly data lacks is independent information: 24 points of a persistent process genuinely look like a trend.
One layer is pure arithmetic. Over a baseline of k points the largest attainable z-score is (k-1)/sqrt(k):
| Baseline points | Max attainable z |
|---|---|
| 6 | 2.041 |
| 10 | 2.846 |
| 11 | 3.015 |
| 24 | 4.695 |
Against a 3.0 threshold, 10 or fewer baseline points can never fire, for any data whatsoever. And the causal estimator needs 60 observations inside a 180-day pre-period; monthly supplies 5. Because the lookback is capped at six months, that is unreachable at any company age. Cash, runway, burn and ledger revenue are the numbers you care most about and the ones your instruments can say least about.
The playbook
- Separate detection from attribution. Detection holds 100% power on a planted trend (n=200). Attribution of a planted +10% on a compounding metric recovered 0 of 10. "Signups fell 18%" is defensible. "Signups fell because of onboarding" is a different claim needing a counterfactual you do not have.
- Demand magnitude, not significance. 49 of 120 quiet windows produced a significant effect, median size 0.58%. A 1% floor removed 48 of the 49 and cost no power. Write your floor before you look.
- Aim decisions at daily-observable outcomes. The estimator accepts a daily metric with 180 pre-period points and permanently rejects a monthly one with 5. Keep the monthly number as the scoreboard, not the evidence.
- Date every decision. Overlapping pre and post windows cannot be separated, only flagged and discounted. And bad early readings persist: one wrong -20% before twelve correct +10% readings left the running estimate at +7.34%.
- Your control series lives outside your company. 0.00 to 0.70 at a planted 5%, with the 30% false-positive caveat above. Benchmark peers for direction, never for attribution.
Can you measure it?
Enter your metric type, the effect size you are hoping for, and your window. You get one of three verdicts mapped to a measured cell, with the rate and sample shown: detectable, detectable but not attributable, or statistically invisible at this window. Outside our measured range it says so instead of guessing.
Everything above is synthetic: the estimator failing to find effects we planted ourselves, not evidence about any real company's numbers. Limits, open problems and the file behind every figure are in the Issue 002 appendix.
The control series a single company lacks is what BayesBrain is building across companies. Scan your stack.
Take the founder survey (~5 min) or subscribe by email.
Be part of Issue 003's dataset.
The founder survey takes about 5 minutes. Your responses feed Issue 003 onward, and you get every issue in your inbox the morning it ships.
The control series a single company lacks is what BayesBrain is building across companies. Scan your stack.