Experiments · a layer of Growth Engineering
Nothing compounds. Month twelve starts like month one.
Experiments change that: simulate to remove the weak ideas, prove the rest in the market, keep what wins.
Most tests lose, and most teams cannot run them
Across 127,000 experiments, Optimizely found only 12% won on the primary metric. Most B2B sites do not have the traffic to reach a reliable answer at all, so teams either guess or run tests that can never conclude.
The ground moves under the measurement
Ask an AI engine the same question twice and the recommendation list rarely repeats, with agreement under one in a hundred in SparkToro's 2026 research. Any tool selling you a single AI rank is selling noise, and decisions built on it are guesses wearing a number.
Ideas as a set, not one at a time
Candidates are generated from what the system knows about your market, then cut by checks for distinctness and evidence before a person ever reads them.
The simulated panel only kills
A panel of simulated buyers can remove weak options and is never allowed to crown a winner, because stated preference is not purchase. One idea the panel rated poorly always goes to the real market, so the simulator can be caught being wrong.
Then reality decides
Survivors run against a measure agreed in advance and a condition that stops them. Where volume is too low for a classic test, we measure the things that happen often enough to read: chats started, replies, clicks.
What won gets carried forward
The traits behind a winner, the hook, the proof, the angle, update a written belief and seed the next round. That is the difference between a system that produces and one that compounds.
Everything we have written on this
Where this layer is put to work
- Demand creation: Tests angles and hooks on small bets before the calendar spends on them.
Questions
We get very little traffic. Can we test at all?
Not in the classic A/B sense, and pretending otherwise wastes months. We test on the things that happen often enough to read, such as replies, chats started and clicks, and use simulation to remove the weakest ideas before spending.
Are simulated buyers trustworthy?
Only for removing weak options, and we treat them as advisory until they have been checked against real outcomes. The first such check is scheduled and will be published whichever way it goes.
What stops the same idea being retried forever?
Every test is registered with an owner, a date to check it and a condition that ends it.
How honest is the simulation?
Advisory until it has been checked against real outcomes. The first check is scheduled, and we publish it either way.
What if a test is inconclusive?
It is reported as inconclusive. Thin evidence never moves a belief.