What we do

You shipped forty changes last quarter. Which ones made you money?

Most teams cannot answer that, and the ones who think they can are usually reading a dashboard that was never built to answer it. We fix that in the only way that survives a hard question, which is by measuring it properly before the money is spent.

01 / 05

Three ways this quietly costs you money.

None of these are our opinion. Each one is a published figure or a number we computed ourselves, and each one has a link under it.

2 in 3
shipped ideas did not improve the metric they were built to improve. You paid to build all three.
Kohavi & Longbotham, 2015
27%
of tests with nothing in them still produce a winner when somebody checks the dashboard daily. We measured that over 200,000 runs.
Our simulation
40×
gap between what attribution claimed a channel returned and what the experiment on the same spend actually found.
Blake, Nosko & Tadelis, 2015
30 mo
to detect a 5% checkout lift at a thousand checkout starts a month. Most shops run that test anyway and act on the result.
Our calculation

Put together, the damage is not that a change failed. It is that nobody could tell, so the failed change stayed in the product and the roadmap was built on top of it.

02 / 05

Three ways we work with you.

Pick the one that matches where the pain is. Most clients start with the second and move to the first once the numbers can be trusted.

01 · Test programme

We run the queue, and we tell you when to stop

You get a decision every fortnight instead of an argument every fortnight.

For you if

  • You have traffic and no shortage of ideas
  • Tests get called by whoever is loudest
  • Nobody can name last quarter's winners

What lands on your desk

  • A sized backlog, with a date against each test
  • A pre registered metric and stopping day
  • One page per readout, in plain language
  • A written no when the result does not hold

What it removes

  • Tests that could never have answered anything
  • Winners called on day four
  • Rollouts that quietly cost you money
02 · Measurement build

We fix the numbers everything else is standing on

The same question gets the same answer twice, and you stop paying people to argue about whose export is right.

For you if

  • Two dashboards disagree and nobody knows which is wrong
  • Every number needs an analyst to explain it
  • Tracking was built by whoever was free that sprint

What lands on your desk

  • A tracking plan tied to real business questions
  • Event schema, naming and version control
  • Warehouse tables you can query without us
  • Dashboards that show intervals, not just points

What it removes

  • Disputed results that were a tracking bug all along
  • The monthly reconciliation meeting
  • Dependence on one person who remembers how it works
03 · Analysis review

We are the second opinion before you roll it out

One page you can hand to your CFO, or one page that stops a bad rollout before it reaches production.

For you if

  • Somebody wants to ship a 12% lift on Monday
  • The win came from one segment, found afterwards
  • The result feels too good and nobody wants to say it

What lands on your desk

  • Re analysis from the raw assignment data
  • Peeking, novelty and sample ratio checks
  • Segment claims tested for multiple comparisons
  • A verdict in one page, with the workings attached

What it removes

  • A number in the board deck that will not survive a quarter
  • Engineering time spent scaling a false winner
  • The habit of never checking
03 / 05

What the first month looks like.

No discovery phase that bills for six weeks and produces a slide deck.

Week 1

We size the questions you already have. Some of them turn out to be unanswerable at your traffic, and you find that out on day four rather than in month seven.

Week 2

We check what your tracking is actually recording. This is where most disputed results are born, and it is usually a two day fix.

Weeks 3 and 4

First test goes live with its stopping day already written down, or the first measurement fix ships. You get the readout on the agreed date, not when it looks good.

After that

A standing queue, a fixed readout day, and a running record of what was tried and what it did. You keep all of it, including the raw data.

04 / 05

What we will not do, so you know what you are buying.

We size before you commit

If your traffic cannot answer the question, we say so in the first week and we do not take the project. You are not paying us to run a test that was never going to resolve.

We say no in writing

When a result does not hold, you get that in the same one page format as a win. A consultant who only ever finds wins is not measuring anything.

You own everything

Raw assignment data, warehouse tables, the analysis scripts. If you stop working with us, nothing walks out of the door with us.

We publish our workings

Every method we use on your data is written up on this site, with the numbers and the seed, so you can check us. That is the whole of our field notes section.

This page describes how we work and it is not advice for your specific situation. Before you act on anything here, check it against your own numbers and your own analyst.

05 / 05

Start with the question, not the contract.

Send us the change you are planning and the traffic you get. We come back with whether the test can answer it, how long it takes, and what it would cost to find out. That part is free and it takes us about an hour.

Read before you write to us
Why a daily look turns 5% into 27% The checkout test most shops cannot run Attribution said 4,173%, the experiment said minus 63%