Lesson 20 of 22

The wall7 minWe were wrong

We said photo decks beat video 10×. We were wrong.

Two numbers we published, how both fell apart, and what we should have measured.

For a while we believed, and said out loud, that photo decks outperformed video by roughly ten times. It shaped what we made. It was wrong in two independent ways at once, which is the interesting part.

Error one: the accounts weren't the same

The decks and the videos hadn't run on the same accounts. Different accounts have different audiences, follower counts, warmup histories — and, as we found out separately, wildly different audience geography. Comparing across them measures the accounts at least as much as the format.

The most ordinary confound there is, and we walked into it because the comparison was assembled after the fact from data we already had rather than designed as a test.

Error two: mean versus median

Short-form performance is violently long-tailed. A handful of posts do enormous numbers and the rest cluster low. Take a mean across that and you're mostly reporting whether a breakout happened to land in your sample.

The correction

On medians — the number that describes a typical post — the two formats came out at 113 and 86. A real difference. Not a 10× difference, and not enough to reorganise production around.

The second number that fell

We'd also published that about 86% of viewers quit on the cover slide of a deck. Also wrong, and it fails for a subtler reason worth understanding, because a lot of people quote numbers like it.

You can estimate a survival curve across deck lengths. What you cannot do is recover the per-slide drop-off rate from that data. Models with very different per-slide behaviour fit the observed numbers identically once you allow for decks differing in quality. The data doesn't distinguish them.

So "86% quit on the cover" was never a measurement. It was one of several hypotheses that fit equally well, reported as though the data had picked it. Answering it properly needs a slide-swap experiment — same deck, same account, cover varied.

What we do now

  • Compare formats on the same accounts, or don't compare them.
  • Report medians on anything long-tailed. Means describe the outlier.
  • Before quoting a rate, ask whether the data could have come out differently under a competing explanation. If not, it isn't evidence.
  • Never let a number estimated one way get quoted as though it were measured another way.

Takeaway

Same accounts or no comparison. Medians, not means. And check a competing explanation could have been ruled out.

Questions this raises

Why do means overstate short-form performance so badly?

The distribution is long-tailed — a few breakouts sit far above a low-clustered body. The mean tracks whether a breakout landed in your sample rather than what a typical post does, which on content-test sample sizes is mostly noise.

Every law on one page

The compiled reference — framing laws, the negative prompt block and the pre-ship checklist. One email, no sequence you can't leave.

The pack, and a note when a new teardown goes up. Unsubscribe whenever.