A service can report 99.99% uptime while customers repeatedly fail to finish the action that creates value. The API answered, the queue ran and the database stayed available, but the checkout confirmation never appeared, a retry created a duplicate, or the result arrived too late to be useful. Infrastructure health is necessary; it is not the product outcome.
Product Engineering needs a reliability boundary that follows the user's intent across the interface and the backend. A user-journey SLO does that. It defines which attempts count, what a good outcome means, how fast it must happen and how much failure the product can tolerate before the release policy changes.
Google's SLO definition separates the measured service-level indicator, or SLI, from the target SLO. The useful product move is to apply that separation to a journey such as placing an order, publishing a campaign, importing a file or receiving an AI answer—not only to an HTTP endpoint.
Choose one journey with a verifiable outcome
Start with a journey that is valuable, frequent enough to measure and owned by a team that can change it. “The application works” is not a journey. “A merchant submits a valid campaign and sees it scheduled exactly once within 30 seconds” is.
Write four boundaries. The start event marks a real attempt after prerequisites are satisfied. The terminal success event proves the promised outcome, not merely a button click. Expected business rejections are classified separately from system failures. The measurement window states when an attempt becomes late or abandoned.
For checkout, a valid attempt may begin when the server accepts a checkout command with a stable idempotency key. Success may require an order identifier, a persisted payment state and a confirmation visible to the customer. A declined card is not platform unavailability if the system explains it and enables recovery; an unexplained timeout is.
Do not begin with every journey. Pick one whose failure changes revenue, trust or support load, establish the contract, then reuse the method.
Define the SLI from valid attempts and good outcomes
The simplest journey availability SLI is a ratio:
journey_availability = good_terminal_outcomes / valid_journey_attempts
good = completed exactly once
AND result is visible
AND correctness checks pass
AND duration <= promised thresholdThe denominator is a product decision. Exclude synthetic probes, load tests and requests rejected before the product accepts responsibility. Do not exclude real failures because they are inconvenient. Record each exclusion reason and review its volume; a growing “invalid” bucket can hide a broken prerequisite or client integration.
Use buckets rather than averages for latency. The Google SRE example SLO uses proportions of requests below explicit thresholds. A product journey can similarly require, for example, 95% under 3 seconds and 99.5% under 10 seconds. An average of 1.8 seconds can conceal a painful tail affecting a meaningful customer segment.
Correctness belongs in the indicator. Exactly-once order creation, complete imported rows, correct entitlement state and a cited AI answer are different predicates. If correctness cannot be checked, the team has defined a completion event, not a reliable outcome.
Join the browser and backend with one journey identity
Create a journey_id when the product accepts the attempt, and propagate it through the browser event, API command, queue message, database write and terminal response. Keep it separate from trace_id: the journey can span retries, async jobs and more than one trace.
OpenTelemetry traces describe the path of a request across application components, while context propagation allows signals from distributed services to be correlated. Add low-cardinality journey attributes such as journey_name, release_version, result_class and tenant_tier. Keep customer identifiers and free-form input out of metric labels and Baggage; use controlled references in protected logs when investigation requires them.
Emit state transitions rather than a single success counter: accepted, processing, waiting_on_dependency, completed, rejected_expected, failed_system and expired. Include occurred_at and a monotonic sequence or version. The terminal projector should count a journey once even when retries repeat events.
The frontend adds evidence the backend cannot see. Record whether the confirmation was rendered, whether the page became interactive and whether the user recovered from an error. The backend remains authoritative for durable state; the browser remains authoritative for what the customer experienced. The SLO can require both without pretending they are the same signal.
Put experience latency beside system latency
Backend duration answers how long the server worked. Journey duration answers how long the user waited for value. Measure from accepted intent to verified outcome, and split waiting into server work, queue delay, dependency time and presentation delay.
For web journeys, Core Web Vitals provide standardized experience signals: LCP for loading, INP for responsiveness and CLS for visual stability. They are not a substitute for a journey SLO. Join them by release, page family and journey—not by raw user identity—so a conversion regression can be investigated alongside field performance.
Use distributions and segment responsibly. Device class, country or tenant tier may reveal a concentrated failure that the global SLO hides. Avoid turning every dimension into a permanent metric label; high cardinality raises cost and operational risk. Keep broad dimensions in metrics and investigate detailed cohorts in traces or the warehouse.
The web.dev guide for combining Web Vitals with GA4 and BigQuery is a useful example of querying field experience beside business dimensions. Correlation starts an investigation; it does not prove that latency caused the business outcome.
Turn the objective into an error budget
An SLO without a consequence becomes a dashboard decoration. If the target is 99.5% over 28 days, the remaining 0.5% is the error budget. Every bad valid attempt spends it. Track remaining budget and burn rate, not only current success percentage.
The Google SRE Workbook treats SLOs as a basis for data-driven reliability decisions. Its alerting guidance focuses response on significant budget consumption rather than every transient error. Use at least a fast-burn window for sudden regressions and a slow-burn window for gradual deterioration.
Write the policy before a breach. When the budget is healthy, the team may increase rollout exposure. When burn accelerates, pause the cohort expansion and inspect the affected journey. When the budget is exhausted, stop unrelated risky changes, restore reliability and require evidence before resuming. Google's example error-budget policy presents the policy as a way to protect users and rebalance work, not punish a team.
Keep exceptions explicit: security fixes, a dependency-wide incident, an excluded test population or a verified measurement defect. An exception needs an owner, evidence and an expiry date.
Make release gates depend on outcome evidence
Attach release_version, flag_variant and rollout_cohort to the journey record. Begin with internal traffic or a small customer cohort. Compare the new version with a stable baseline on journey availability, tail duration, duplicate outcomes, recovery rate and support signals.
Do not promote on “no alerts” alone. Require a minimum number of valid attempts and enough observation time to cover asynchronous completion. A canary with three successful checkouts proves almost nothing. For low-volume journeys, combine production evidence with deterministic contract tests, replayed fixtures and longer observation.
Automate only decisions the metric can support. A clear fast-burn breach can trigger rollback or freeze expansion. A small ambiguous difference should create a review, not oscillate traffic. Add a cooldown after rollback and ensure the previous version, schema and queue consumers remain compatible.
Record the decision: exposure before and after, SLO state, evidence window, operator and reason. That history turns release automation into an auditable product system.
Design alerts around customer harm
Page an engineer when a critical journey is consuming budget quickly and action can reduce harm. Create a ticket or daily review for slow burn, growing exclusions, delayed async work and a narrow segment regression. Keep diagnostic infrastructure alerts, but route them as supporting evidence unless they threaten a journey or a shared hard limit.
An alert should include journey name, affected cohort, release version, burn windows, recent changes and trace links sampled from bad outcomes. Do not make the responder reconstruct the product impact from CPU charts.
Guard against silent telemetry failure. Monitor the ratio between accepted commands, emitted journey starts and terminal events. A perfect SLO with a collapsing denominator is not reliability; it is missing measurement.
Review the contract with Product, Design and Engineering
Product defines the promised outcome and acceptable failure. Design defines what success and recovery look like to the customer. Engineering defines durable state, correctness predicates and instrumentation. Operations contributes alertability, capacity and incident response. Finance or compliance may define reconciliation or evidence requirements.
Review the SLO when the journey, audience or promise changes—not whenever a bad week makes the target uncomfortable. Keep the old definition queryable so trend changes are not confused with contract changes.
Every review should ask: Are valid customers excluded? Can success happen without value? Can one attempt be counted twice? Does the threshold match the customer promise? Can the team act on a breach? Is the telemetry itself observable?
A practical implementation sequence
First, name one critical journey and draw its accepted and terminal states. Second, define valid attempts, good outcomes, expected rejections and expiry. Third, add journey_id and idempotent state events across frontend and backend. Fourth, build the SLI query and compare it with manually inspected samples. Fifth, choose an initial target from actual performance and product expectations, then write the error-budget policy. Sixth, connect the SLO to a staged rollout without enabling automatic rollback until false positives and data delay are understood.
Run a shadow period. Reconcile terminal counts with the authoritative business records. Sample good and bad journeys into traces. Test telemetry loss, duplicate events, delayed jobs and rollback compatibility. Publish the definition beside the dashboard so nobody has to infer what “99.5% reliable” means.
The product engineering decision
Product reliability is not the percentage of servers that answered. It is the proportion of valid user intentions that became correct, visible and timely outcomes.
A user-journey SLO makes that promise executable. It gives Design a recovery boundary, Engineering a measurement contract, Product a risk budget and Operations a signal tied to customer harm. When the same contract controls rollout, reliability stops being a post-release report and becomes part of how the product is built.
Official references
These references document the tools discussed. Examples and design decisions are illustrative and should be adapted to the project and its versions.
- Google SRE — Service Level Objectives
- Google SRE Workbook — Implementing SLOs
- Google SRE Workbook — Alerting on SLOs
- Google SRE Workbook — Example Error Budget Policy
- OpenTelemetry — Traces
- OpenTelemetry — Context propagation
- web.dev — Web Vitals
- web.dev — Measure and debug performance with GA4 and BigQuery
Prepared by: Noor Yasser
Working through a similar engineering challenge?
I help teams turn architecture decisions into a clear scope and dependable, reviewable implementation.




