--:--:--
← ALL NOTES
Sep 08, 20264 MIN READRESEARCHSTATISTICSMETHOD

Compute the placebo floor first

Any estimator that extrapolates a trend will return something. Running it on years where nothing happened is the cheapest test we know, and the one that most often kills the claim.

Here is a shape of analysis that feels rigorous and is not. Fit a trend on old data. Extrapolate it to a target year. Subtract. Call the remainder excess and write it up. The method never says no, because subtraction always returns a number, and the number always looks like a finding.

The fix is boring and it is the first thing we run now: point the entire pipeline at a year when the thing you are looking for had not happened yet, and see what it hands back.

What the floor cost us to learn

In the Korean abstract study we targeted 2020, 2021 and 2022 with the same estimator, the same base window shape, the same everything. The floors came back at 0.1 to 2.2 points for the single-word statistic and at most 2.9 for the split-half set statistic.

Then we looked at our own 2023 number. It sat inside that band. So the paper reports nothing for 2023, and says so in the abstract. That is the entire value of the exercise: it took a year off the timeline that we would otherwise have claimed.

The floor is not one number

It moves with two things, and both are choices you make rather than properties of the data.

Extrapolation distance. Predicting one year ahead from five base years is a different act from predicting four years ahead from two. We ran the grid rather than picking one cell, and the floor grows monotonically along it. Any single reported floor without the grid behind it is a floor for one configuration, quoted as if it were a property of the method.

The selection rule. A floor computed over every word is dominated by high-frequency function verbs that drift for reasons no one is claiming. Requiring a marker to be at least 1.5 times its expected rate before it counts drops the floor by roughly a factor of three at every horizon, because it removes exactly the words whose drift is uninteresting.

Clustering eats intervals

The second cheap test is to ask what your resampling unit is. Abstracts within one journal share authors, subject, house style and an editor. Resampling documents treats them as independent and returns intervals about 40% narrower than resampling journals. Nothing about the point estimate changes. Only the confidence you are entitled to.

Selection and measurement must not share a corpus

If you pick your marker words on the same data you then measure them on, you have measured your own selection. Splitting journals by a hash, choosing markers on half A and measuring on half B, costs a few lines and made the difference between a statistic we could publish and one we could not. Run the same split on a placebo year and the set bound comes back near zero, which is what tells you the selection bias is small rather than assuming it.

Why this is a blog post and not a footnote

Every one of these is a way of being told no by your own data before a reviewer tells you. None of them is expensive. The placebo grid in the study is fifteen cells and runs in minutes. We would rather spend that and lose a year off a headline than defend a number that was never above the noise.

WRITTEN FROM THE INTFRAME ENGINE ROOM

WORK WITH US →