The ceiling

The knowledge layer

The band across the top of the map, drawn as a band because every box touches it. It is not a stage and it is not a method: it is where the map keeps what it learned.

It is the only object in the whole system that gets more valuable the longer a team runs it. Everything else on the map answers a question and stops; the ceiling is where the answers accumulate into something the next question can stand on. It has two halves, and they are different kinds of object.

The knowledge repo

Prose. One searchable entry per question, holding the design, the number, the decision, and what you would do differently, with the losses written up as carefully as the wins. The repo is what a person reads before proposing anything.

The unit is the question, not the ticket. Tickets get closed and archived by the board; a question outlives the feature that prompted it, and eight months later somebody proposes the same test. The entry is what turns that moment from a re-run into a lookup.

What an entry has to contain:

  • The question, written so that two different answers were imaginable. If you cannot picture being surprised, it was not a question.
  • The decision it unblocked: what changed on the answer, and who made the call. A named person, a named choice.
  • The method: which bucket it landed in from the routing, and the design. Half of a roadmap's questions turn out not to be randomisable, and the entry is where that gets found out once instead of quarterly.
  • The number, shrunk, with its interval, or the honest statement that no defensible number existed.
  • The belief: the sentence that outlives the report. Not "variant B won, +2.1% on checkout completion", which dies with the ticket, but "this product responds to friction removal at the payment step, about two points, and it held for six weeks", which the next three questions start from.

The prior store

0 10 20 30 40 -0.8 -0.4 +0.0 +0.4 +0.8 +1.2 +1.6 effect each test produced, percentage points results on this metric, out of a hundred mean +0.07pp: the honest prior an MDE of 1.5pp, which nothing here has ever produced largest ever: +0.76pp
Illustrative: a hundred results on one metric. The mean is the honest prior; the MDE somebody wished for sits to the right of every effect the metric has ever produced.

A table. For each metric, the distribution of effects the last hundred tests actually produced. The store is what sizes the next test, and what the last result gets shrunk toward.

The store earns its keep twice per experiment:

  1. At design time, the minimum detectable effect stops being a wish. The store says what this metric has actually moved by; a test powered for a lift nobody at this company has ever produced was decided before it launched. Effects are small and most ideas do nothing: seventy to ninety percent of experiments are killed or neutral everywhere it has been measured, so the honest prior mean is roughly zero, and only a store that records the losses keeps it there.
  2. At read time, raw winners are inflated: a result reported because it crossed a threshold is, in expectation, an overstatement, and the overstatement grows as power falls. Empirical Bayes shrinkage toward the store's distribution is the correction, and the shrunk number is what gets written back, so the store deflates rather than inflates.