Product management · Experimentation

Product experiments need a decision contract

A practical product-management framework for defining hypotheses, success metrics, guardrails, data-quality checks and ship-or-stop rules before an experiment starts.

TOPIC HUBProduct Engineering
Original conceptual illustration of two experiment cohorts flowing through a decision contract into success, iterate and stop outcomes; no fabricated results.
An editorial interpretation of the topic, followed by a practical execution diagram.

A product experiment often begins with a feature and ends with a debate. The team ships two variants, watches a dashboard, and only then asks what movement is large enough to matter, which users should count, how long the result should persist, or how much reliability it is willing to trade for activation. The data has arrived, but the decision was never designed.

A decision contract fixes that sequence. It is a short agreement written before exposure that states the user problem, the riskiest assumption, the intended mechanism, the population, the outcome metric, guardrails, data-quality checks and the actions attached to each credible result. It does not guarantee a positive result. It makes a negative or inconclusive result usable instead of politically negotiable.

The GOV.UK Service Manual frames alpha as the place to test the riskiest assumptions with work that is only complex enough to learn, and to decide which ideas deserve beta. Microsoft Research's experimentation guidance adds the measurement discipline: validate assignment, protect critical metrics, ramp safely and verify that telemetry itself did not change. This article turns those ideas into one working product-management document.

Begin with the decision, not the variant

Write the decision the team expects to make: launch broadly, continue learning, redesign the intervention, or stop. Then write the uncertainty preventing that decision. “Should we launch the new onboarding?” is too broad. “Will reducing the account-setup sequence help eligible first-time administrators reach the first successful integration without increasing configuration errors?” names a population, a behavior and a cost.

Separate the problem from the solution. The user problem might be that administrators cannot understand which credential belongs in which field. A shorter form is only one proposed mechanism. A prototype interview, comprehension test or instrumented concierge flow may test the risky assumption more cheaply than a production A/B test. Controlled experiments estimate causal effects when random assignment and reliable telemetry are available; they are not the default answer to every discovery question.

State what would change your mind. If no plausible evidence would make the team stop the feature, the exercise is rollout monitoring, not hypothesis testing. That may still be valid, but label it honestly and design safety checks rather than pretending the launch is an experiment.

Write the causal chain

A useful contract has one sentence in this form: for a defined population, intervention X should change behavior Y through mechanism Z, leading to product outcome O within window W. Every link needs evidence. Exposure logs whether the user actually became eligible. The local behavior metric checks the mechanism. The outcome metric tests whether the behavior mattered. Guardrails check whether value was purchased by harming something else.

Consider a hypothetical B2B SaaS onboarding change. The population is new account administrators who have not connected an external system. The intervention replaces a six-step setup with a guided three-step path. The mechanism is reduced uncertainty, not merely fewer clicks. The local behavior is completing credential validation. The outcome is reaching the first successful sync within seven days. Guardrails include configuration-error rate, support contacts, page latency and later disconnects.

This example is illustrative, not a reported case study. Its purpose is to show why “completion rate increased” is insufficient. A user can complete a shorter flow and still connect the wrong account. The causal chain tells the team where the apparent gain can break.

Define evidence before seeing it

Choose one primary outcome for the decision and specify its direction, unit of analysis, denominator and observation window. “Activation” is not a metric until the event sequence and eligible population are defined. A rate needs both numerator and denominator under version control. Microsoft Research warns that denominator movement can make rate metrics contradict one another, so expose the denominator as its own diagnostic metric.

Write the smallest effect that would change the product decision. Experimentation teams call this the Minimum Detectable Effect, or MDE, when planning statistical power. The STEDII metric guidance describes MDE as the smallest movement a test can detect with high probability, commonly planned around 80% power. Product management still has a second question: even if a tiny movement is detectable, is it worth engineering, operating and support cost?

Use guardrails differently from success metrics. A high-level retention or reliability metric may be difficult to improve in a short test but easy to damage, which makes it a useful guardrail. For every guardrail, record the unacceptable regression and the action it triggers. “Watch errors” is not a rule. “Pause exposure if configuration failures exceed the agreed bound after data-quality checks pass” is operational.

The decision contract connects the product hypothesis to exposure, trustworthy evidence, guardrails and a pre-agreed outcome.
The decision contract connects the product hypothesis to exposure, trustworthy evidence, guardrails and a pre-agreed outcome. Open for a larger view

Protect the validity of the test

Before interpreting user behavior, verify that the experiment produced trustworthy groups. Sample Ratio Mismatch, or SRM, means the observed treatment-control allocation does not fit the expected ratio. Microsoft's during-experiment guidance says SRM typically invalidates the test and recommends early alerts. A surprising conversion lift cannot rescue broken assignment.

Specify the assignment unit: user, account, workspace, device or session. Choose the unit that prevents treatment leaking into control. In a collaborative SaaS product, assigning individuals inside the same account may expose both experiences to one organization and contaminate the comparison. Record exclusions before launch and keep them symmetric. Filtering treatment by a behavior caused by treatment creates a biased population.

Validate event parity with an A/A or pre-flight check where appropriate. Confirm that exposure, eligibility, primary outcome, guardrails and denominator events arrive for both variants. Microsoft Research documents cases where improved logging in one variant moved metrics even when user behavior was not the intended cause. Track telemetry completeness as a data-quality metric, not as a footnote after results look strange.

Connect guardrails to automatic actions

The diagram below turns the contract into a gated flow. Eligibility and randomization create the cohorts. Data-quality checks decide whether the evidence can be interpreted. The outcome and guardrails then lead to ship, iterate, stop or inconclusive. A red guardrail is not averaged away by a green primary metric.

Ramp by risk, not optimism. Microsoft's pre-experiment guidance recommends progressing through populations and gradually increasing exposure to balance certainty against user harm. Start with internal or low-risk traffic where appropriate, then a small percentage, and expand only after assignment, telemetry and guardrails are healthy. The exact percentages belong to the product's risk model, not a copied universal sequence.

Pre-authorize automatic shutdown for egregious harm such as broken navigation, severe error growth or a critical performance regression. Microsoft reports that automated shutdown avoids waiting for a feature team to react to a clearly damaging test. Keep the trigger conservative, observable and reversible, and page an owner with the evidence that fired it.

Use a four-way decision matrix

Ship when the outcome clears the practical threshold, guardrails remain acceptable, assignment and telemetry are trustworthy, and the mechanism is plausible. Do not convert “statistically significant” into “valuable” without checking effect size, operational cost and affected population.

Iterate when the mechanism moved but the outcome did not, or when a specific friction blocked the intended behavior. The next experiment should address that diagnosed break rather than rerun the same idea with new colors. Stop when the outcome misses the practical threshold with enough evidence, a guardrail fails, or the underlying assumption is contradicted.

Mark the result inconclusive when the test lacks power, suffers SRM, has broken telemetry, changes its population midstream or cannot separate the treatment from infrastructure effects. Inconclusive is a data-quality or design result, not permission to choose the preferred variant. The contract should say whether the team will repair and rerun, switch to qualitative research, or archive the idea.

Review segments without manufacturing a winner

A treatment may affect markets, devices or prior-activity cohorts differently. Predefine a few stable segments tied to the hypothesis. Microsoft's guidance favors segments whose membership is not changed by treatment, such as market, browser, app version or activity measured before the experiment. Post-hoc slicing through dozens of groups makes a lucky movement easy to find and hard to reproduce.

For every segment claim, verify balance and sample size. A global neutral result with one positive subgroup can be a real opportunity, a data problem or noise. Treat it as a new hypothesis unless the segment and expected direction were part of the original contract. Do not silently redefine the target population after seeing the scorecard.

Keep the contract short and executable

A useful one-page contract contains: decision owner and review date; user problem and riskiest assumption; population and exclusions; intervention, control and expected mechanism; assignment and exposure units; primary outcome, practical threshold, MDE and window; guardrails and automatic stop rules; telemetry and SRM checks; ramp plan; and the ship, iterate, stop and inconclusive actions.

Review it with Product, Design, Engineering, Data and Operations before exposure. The Product Manager owns clarity of the decision, not unilateral control of the numbers. Engineering confirms assignment and rollback feasibility. Data confirms metric definitions and power. Design checks whether the treatment actually represents the hypothesis. Operations names failure modes that a conversion metric will miss.

The practical rule is simple: do not ask a dashboard to invent a decision after the result arrives. Decide what evidence would change the product, protect the user from what must not regress, and give every credible outcome a pre-agreed next action. The experiment then becomes a learning instrument rather than a vote on the feature.

Official references

These references document the tools discussed. Examples and design decisions are illustrative and should be adapted to the project and its versions.

Prepared by: Noor Yasser

FROM DECISION TO DELIVERY

Working through a similar engineering challenge?

I help teams turn architecture decisions into a clear scope and dependable, reviewable implementation.

Book a 30-minute callRelated serviceSaaS & digital product developmentRelevant projectMember Plus