Integrations, Data & AI

Ship behind a flag
decide on real numbers, not on who argued longer

Shipping a change to everyone at once and hoping it helps is still how most teams ship. We build feature flag infrastructure so a change goes to a small slice of users first. It gets measured against a metric that matters, and rolls back in seconds if it doesn't help.

from$1,200
Timeline1 to 3 weeks
What is includedFeature flag system wired into your backend and/or frontendPercentage-based and user-attribute-based rollout rulesA/B test assignment with consistent bucketing per userIntegration with your analytics tool for measuring experiment resultsAn instant kill switch per flag, no deploy required to roll back
1-3 weekstypical time from kickoff to a working flag and experiment system
secondstypical time to roll back a bad change once it is behind a flag
onemetric agreed before launch, instead of an argument about results after

What it is

A feature flag wraps a piece of code in a condition. Show the new checkout flow to this percentage of users, or to users with this attribute, and the old flow to everyone else. A switch controls it, and that switch doesn’t need a deploy to change. Experiment infrastructure builds on top of this. It assigns users consistently to a test group, and measures a defined metric for each group. The result is a statistically grounded answer to whether the change actually helped.

When you need this (and when you don’t)

You need this once a change is risky enough that rolling it out to everyone at once is a real gamble. A new pricing page, a redesigned onboarding flow, a changed algorithm, all qualify, anywhere being wrong costs real revenue or real users. It’s also the right build once “ship it and see” has produced a few disagreements about whether a past change actually helped. That usually means nobody measured it properly at the time.

You don’t need this for low-risk changes, a copy tweak, a color change, a bug fix. The cost of being wrong there is trivial, and a flag just adds process without adding value. The tell that you need it: someone in the room says “I think this will help,” someone else disagrees, and neither has data to settle it.

How we build it

For most clients already on PostHog for product analytics, we use its built-in feature flags. That keeps assignment, tracking and results in one tool instead of stitching two systems together. Rollout rules support simple percentage-based assignment and attribute-based targeting alike. Think a specific plan tier or region, for cases where a blanket percentage isn’t the right test.

Every flag has an instant kill switch. Flipping it off takes effect immediately, with no deploy. That’s the entire point: a bad change should be reversible in seconds, not after an incident review and a hotfix. For experiments specifically, we help define the success metric and minimum sample size before launch, not after results come in. A metric chosen after the fact tends to be the one that makes the result look good, not the one that actually answers the question.

What to watch

The real risk with feature flags isn’t technical. It’s organizational drift. Flags meant to be temporary for a rollout stay in the codebase for years, and eventually nobody remembers which combination of flags is actually live for which users. The code itself turns into a maze of conditionals. We recommend a flag cleanup review on a schedule, and build flags to be removed once a rollout completes, not left in indefinitely by default.

On the experimentation side, calling a result significant before the sample size is reached is the most common statistical mistake teams make on their own. We build a minimum sample size check into the reporting, so a test doesn’t get called early just because the numbers looked good on day two.

Sample ratio mismatch is a subtler failure worth watching for. Say your assignment logic is even slightly biased, excluding users on an older app version from one group but not the other. The groups are then no longer comparable. The result isn’t trustworthy no matter how significant it looks. We check for this before reading any result as final. It’s a common, easy-to-miss way an otherwise correct-looking experiment produces a wrong answer, and nobody catches it until a decision is already made.

Price and timeline

Scope Price Timeline
Flags, rollout, kill switch from $1,200 1 to 2 weeks
Full experimentation with significance testing from $2,800 2 to 3 weeks

What this pairs with

Built as part of custom development and analytics. Builds directly on product analytics setup for measuring results. See the engagement mechanics in an AI coach and gamification fitness app. Tell us what change you want to test before committing to it: get in touch.

FAQ

How much does feature flag infrastructure cost?

A core setup with percentage rollout and a kill switch starts at $1,200. A fuller experimentation platform with statistical significance testing and analytics integration runs $2,500 to $4,500.

How long does it take?

1 to 3 weeks. It depends on whether flags need to work across both frontend and backend. It also depends on how many existing features need wrapping in a flag versus just new ones going forward.

Do we need PostHog or a dedicated tool?

PostHog's built-in feature flags cover most needs well and pair naturally with its analytics. A dedicated tool like LaunchDarkly makes sense at a scale, or a compliance requirement, that most of our clients haven't reached yet. We say so rather than oversell it.

What is the actual benefit over just deploying changes?

A bad deploy without a flag means a rollback, a new deploy, and downtime for everyone in between. A bad change behind a flag just means flipping a switch. No deploy, no downtime, and only the test group was ever affected.

Who decides what counts as a successful experiment?

You do, before the experiment starts. We help define a clear success metric and a minimum sample size upfront. That way the result can't be re-interpreted later to match whatever answer someone wanted.

Start here

Tell us the problem.
We bring the system.

A 30-minute call, then a written plan with numbers within 48 hours. No obligation. If we are not the right fit, we will say so and point you to someone who is.

LIKE WHAT YOU SEE?

This site is our work.
Want one like it?

Ten languages, no page builder, launched in 2026 by a team working since 2015. We can build the same quality into your site.

  • 10 languages
  • Since 2015
Get a site like this →