← Benjamin J. Sunter 5 minutes
Measurement

Did it work?

Benjamin J. Sunter · Oct 2026

You change the price, the intake form, the bid strategy. Three weeks later someone asks whether it worked, and the answer in the room comes from whoever sounds most sure about the chart.

I built a small tool for that moment. Paste the number you watch, click the period when you changed something, and read one sentence. Nothing you paste leaves the page.

Try it with your own column →

Three sentences it gave on the example data that ships with it:

How it decides

It fits a line to the periods before the change, projects that line forward, and measures the gap between the projection and what happened. The range around the gap allows for each week leaning on the one before it. Plain regression ranges on business series come out too narrow and call noise a result.

Then it tries to explain the gap away, using only the data you pasted. Was the number already heading this way? Does the stretch after the change line up with the same stretch a year earlier? Does the result rest on one or two unusual periods? Did the periods right before the change dip or spike, so the move is a return to normal? If you pasted a comparison group, did it move the same way? The strongest of those goes in the sentence when it accounts for the move.

Last, it pretends the change happened on every other eligible stretch of your history and counts how often a gap this size shows up. "A change this size showed up at 0 of 10 other dates" means something. "9 of 10" means the move is ordinary for this series.

The words follow the evidence. A before and after comparison on its own never gets a causal word. A projection with the checks agreeing gets "likely". "Caused" needs a comparison group that didn't get the change. If you didn't say what size would matter, the sentence says that instead of deciding for you.

What it won't do

It won't say "no effect". It won't invent a threshold. It won't give a verdict on fewer than six periods after the change. It can't see a change in mix: a bid change that alters which impressions you win can raise click-through with no ad performing any better. It can't separate three changes made in the same month. The drawer under every sentence lists what was checked and what wasn't.

How I checked it

1,198 tests, and 42 synthetic datasets with known truths, each pinned to the verdict it has to produce. The false-alarm rate was measured on 2,000 series per setting with no real change in them, against a 5% target.

3.0%white noise
6.3%autocorrelation 0.3
7.0%autocorrelation 0.6
6.8%autocorrelation 0.7

Then twelve pastes shaped like real exports, with the truth built in. Nine came out as they should. Three real effects on short series were read as "can't tell" or "partly". None of the twelve produced a false yes.

PasteWhat was trueWhat it said
Weekly signups, 40 weeks, Stripe dates12% drop after a price rise; wanted a riseNo. Likely lowered 16%, the opposite of what you wanted
Daily sessions from a GA4 export, 16 weeks18% riseCan't tell. Range -3% to 39%
Resolution hours, 30 weeks, US dates30% drop; bar was 25%Partly. 25%, range doesn't settle the bar
Conversion rate in percent, 32 weeksNo changeCan't tell
Monthly bookings, 36 months, Q4 seasonal, "$410,000"No change; change date at Q4 startCan't tell
Weekly active users on a growth trendNo change, trend onlyCan't tell. The trend it was on explains it
Click-through, one site changed and one not, both hit by a market dip15% rise on the changed site onlyCaused a 14% rise, range 7% to 21%
Monthly churn in percent, 26 monthsDrop from 4.1% to 3.3%Can't tell. Range -30% to 2%
Eight weeks of ordersToo little dataToo early. 3 weeks after, 6 needed
Excel export, newest first, with a Grand Total row20% risePartly. Likely raised 34%; no bar given
MRR with timestamps, "$84,000.00"8% rise; bar was 15%No. Whole range under the bar
A promo spike the week before, then normalNo lasting effectCan't tell

What I got wrong

  1. The slope. On a short series a chance slope before the change, extrapolated over the periods after it, moves the estimate several points. A true 20% rise read as 34%, with a range that covered 20. I tried projecting flat whenever the slope wasn't established. That cut the error on clean steps by a third and raised false alarms on series with a weak real trend from 4% to 35%, because a trend the data can't establish then gets charged to the change. The slope stays. On eight weeks either side, a real 18% rise reads "can't tell". That is the price, and the range on screen carries it.
  2. Commas. A GA4 export writes "Jan 1, 2026" and "1,423" on the same line, unquoted. The parser split that into four cells and refused the paste. It rejoins them now.
  3. Rows that share a date. Three publishers per week came in as three rows per date and were refused. They're added up now, with a note saying so. I kept the interface to one paste and left the filtering to the spreadsheet, where people already do it.

Is it worth anything

A chatbot will run a before and after on a pasted column on demand, and that part of this tool is old. What it has that a chat doesn't: the same method and the same words every time, wording that can't overclaim by construction, a measured error rate, and data that never leaves the sheet.

Log a change and it keeps the verdict, what you expected, and the date of the last check. Log a few and it tells you how your expected effects compare with the measured ones.

That log is the only part that gets more useful with time. Whether anyone wants it installed in their sheet is a question I haven't answered. The engine and the one-page version are done, and the add-on stays unbuilt until someone asks for it.