Tests read correctly:
not called the moment one bar looks taller
Most A/B tests get called the moment one variant looks ahead, often before there is enough data to say anything. We run a disciplined testing program. A ranked backlog of hypotheses. A proper significance check before any test gets called. And a log, so a team running many tests never loses track of what it actually learned.
Where a test gets called too soon
A dashboard that shows one variant ahead by a visible margin gets called a winner within a day or two, in most teams. That is long before the sample size or test duration actually supports the conclusion. A result that looks decisive on day two can flip by day seven as more data arrives. By then the “winning” variant is often already live everywhere.
The opposite problem is just as common, and harder to spot: a test that genuinely shows no real difference keeps running for months, because nobody formally calls it. It quietly eats traffic that could go toward a hypothesis with a real effect to find.
Running tests without a disciplined read is close to running no tests at all. Except it feels like progress, which is worse. It crowds out the fixes that would actually move the number.
How we actually run the program
We build a ranked backlog of hypotheses from your actual funnel data, not a generic list of “proven” tactics. We prioritize by expected impact and traffic available. The first test should be the one most likely to matter, not just the easiest to set up. One or two tests run at a time, on your current landing page builder or ad platform’s own testing tool where possible. Nothing new gets bolted onto your stack unnecessarily.
Every test gets a proper significance check: sample size, duration and variance, all factored in. We show an early-call warning clearly, instead of a confident-looking but premature winner. Where sample size allows, we break results down by device, channel or new versus returning visitor. An overall winner can still lose for a specific segment, a detail easy to miss in an aggregate read.
Every result, win, loss or genuine tie, gets logged with the reasoning behind the call. A program running tests for a year should not quietly repeat a question it already answered.
What you hand us to start
Access to the page or tool where tests will run. And whoever owns the final call on priority within the ranked backlog. We surface the data, but the business context on what matters most this quarter is yours. If you already have a testing tool installed and unused, tell us. We build on what exists rather than adding a new one.
What the quarterly number actually shows
Each test’s result against a proper significance check, with the sample size and duration it was based on logged for audit. Over a quarter, the number of real, confirmed wins against the number of tests run, and an honest count of how many came back inconclusive.
Price and timeline
| Option | Price | What it covers | Timeline |
|---|---|---|---|
| Launch or audit | from $500 | Single test setup on your current tool, one hypothesis | 1 week |
| Monthly management | from $1,200 / month | Ranked backlog, one to two tests running, monthly log | monthly, no lock-in |
| Full control, handover to your team | from $1,500 | Full testing framework and backlog handed to your team, training | 3 to 4 weeks |
Related
This program pairs directly with conversion rate optimisation and landing page design and optimisation. For the statistics engine behind the significance checks, see automated A/B test analysis. The full build is on the analytics service page. Real examples: the Bali lead-routing project, where a quiz funnel was measured step by step from entry to buyer.
Calling tests by eyeballing a dashboard? Get in touch and we will look at your current testing setup in the first call.
FAQ
How much does an A/B testing program cost?
From $1,200 a month for a ranked backlog and one to two tests running at a time, read correctly for significance, no lock-in. A single test setup without the ongoing program is available from $500.
How long does a single test need to run?
It depends on your traffic and current conversion rate, typically 2 to 4 weeks for a meaningful sample on a mid-traffic page. We tell you the expected duration before the test starts, not after.
What traffic do we need for testing to make sense?
A few hundred conversions a month on the page or step being tested is a reasonable minimum. Below that, tests take too long to read reliably, and we recommend bigger, obvious fixes instead.
What do you need from us?
Access to the page or platform where the test runs. And sign-off on which hypothesis to test first from the ranked backlog. We build the list, but the business context on priority is yours.
How do you report results?
Every test gets a written result: winner, confidence level, and what it means for the next step. We log it centrally, so the program compounds instead of repeating the same question six months later.