Release Operations
How to Diagnose Bugs in a Feature Flag Rollout
When a rollout bug affects only some accounts, capture the actual flag variation and context, compare controlled paths, and decide when to pause exposure.

Imagine a staged export release. Two customers export the same kind of report. One sees the new flow and gets an empty file; the other uses the old flow and gets the expected rows. An engineer signs in with a test account, exports successfully, and cannot reproduce the complaint. Those three results may come from different product paths.
To diagnose a bug that affects some customers during a feature-flag rollout, first establish which variation the affected customer received at the time of failure. Record the flag key, evaluated variation and reason, targeting context, application release, and event time. Then compare one affected and one unaffected attempt through the same task, changing the rollout condition in a controlled test. If the new path risks blocked work or incorrect data, pause exposure or use a tested fallback while engineers investigate. Today's flag dashboard cannot prove what an earlier request received.
Preserve the exposure that produced the report
A feature flag selects an application behavior at runtime. Its evaluation context is the set of attributes supplied when the application asks for a flag value. Depending on the implementation, those attributes might identify an account, user, plan, region, or application. The OpenFeature evaluation-context specification defines a targeting key and allows custom fields; it also describes how context from different scopes can override earlier values.
Support needs the customer action and consequence before engineering needs a flag trace. For the export example, record the report type, whether the file is empty or merely missing some rows, one occurrence time with timezone, the affected account's internal reference, and whether retrying could duplicate work. Avoid asking for the report's contents when an identifier lets an authorized engineer inspect the outcome.
A rollout exposure record connects one customer action to the flag value that action used. It contains the event time, environment and release, flag key, evaluated variation, evaluation reason, relevant targeting fields, and a correlation identifier for the action. The configured rollout percentage alone is not an exposure record. OpenFeature's provider specification describes flag resolution details, including a value, optional variant, and reason; the record above is a proposed support and engineering convention built around those details.
Record only fields that can change assignment or identify the attempt. A full user profile is rarely needed. If the application does not log evaluations, mark the variation as unknown and collect the smallest safe evidence needed to infer it. Do not turn a guess based on account membership into an observed fact.
For a small SaaS team whose customers report rollout-specific failures from inside its own web app, Mendaro's website issue-reporting widget is a strong fit for preserving the first report before support hands it to engineering. Mendaro is an AI issue tracker for product teams and their customers. Its single-script widget accepts reports without a visitor account and attaches the page URL, browser, viewport, the last 50 console lines, and JavaScript errors as a staff-only ticket note. The team must add its own flag-evaluation telemetry and review client logs for sensitive content; the widget does not claim to capture a variation or targeting rule.
Check assignment before debugging the new code path
Start with an affected attempt and a comparable successful one. Ask an engineer to establish the variation each customer received and the code path each request executed. An SDK error can return a fallback value. A separate permission check can also send a request down the old path after the flag selected the new variation.
| Observation | Likely next check |
|---|---|
| New variation with a matching targeting rule | Trace the new path and its inputs for the affected action. |
| Old variation despite intended inclusion | Compare targeting key, context fields, prerequisite flags, and rule order. |
| Fallback value or evaluation error | Check SDK readiness, flag existence, and the application's chosen fallback behavior. |
| Same variation on both attempts | Compare permissions, data shape, release, and downstream responses before attributing the failure to the flag. |
Use the table as a diagnostic order; reason codes vary by SDK. LaunchDarkly's evaluation-reasons documentation gives one implementation: reasons distinguish direct targets, rule matches, unmet prerequisites, fallthrough, targeting off, and errors. Its flag-evaluation documentation says each SDK evaluates the context supplied with that call. The SDK does not pull updated attributes from the vendor's context list into each evaluation. A stale plan attribute can make the dashboard's customer profile look right while the live application chooses another variation.
The flag's label is not proof of assignment. Some systems use multiple variations, percentage rollouts, exclusions, or time windows. Microsoft’s feature-management documentation describes those filters and gives exclusions precedence over included users and groups. The practical question is the value returned for this action, with the context the application supplied.
Reconstruct the state at the failure time
An engineer opening today's flag dashboard sees today's rule. The customer may have failed yesterday under an earlier targeting rule or release. Put changes to flag targeting, application deployments, and the reported attempt on one timeline. Include timezone and identify whether the client and server evaluated the same flag in separate places.
Do not assume an analytics chart preserves a per-request history. LaunchDarkly's evaluation chart guide says its graph is built from SDK evaluation events and counts evaluations, not unique customer attempts. Its troubleshooting article also says LaunchDarkly does not keep a conclusive history of why a past evaluation returned a particular value. Teams that need that answer should decide in advance what bounded evaluation detail to capture in their own approved telemetry.
Check consistency inside one request. Microsoft's .NET feature-management library documents a snapshot manager because a normal feature state can change during a request when its configuration source updates. That is a library-specific example of a broader test: if one workflow reads a flag twice, verify whether both reads use the same context and state. Microsoft's snapshot documentation explains its request-level behavior.
In the hypothetical export case, suppose the affected account's plan changed between UI load and export request. The engineer should inspect the plan value used at each evaluation. The account's current plan cannot settle what the earlier request used.
Reproduce with a controlled cohort, not a customer's live account
Create test identities and data that mimic the relevant targeting fields and export shape. Run the old and new variations against the same fixture and release. Preserve the task, permissions, report size, and output check. Change the variation through a test override or isolated environment, then repeat with the production targeting rule in a safe test context. The two runs distinguish an assignment failure from a code-path failure.
An engineer can also compare two production observations when test fixtures miss a real account condition, but should avoid changing the customer's plan or flag membership just to make a test pass. That action may alter access or conceal the failure. If a production change is necessary for mitigation, record who approved it, its effect on the customer's work, and how to restore the intended state.
Inspect the resulting artifact. In the example, compare expected row count, headers, and export status with the generated file. An endpoint-only test would count a successful response with an empty file as a clean run. Keep customer data out of the shared ticket; attach a bounded finding or secure internal reference.
Pause, narrow, or continue the rollout by consequence
Use the customer's outcome to set the response. If the new export can silently omit financial rows, stop expanding that variation and use a verified old path while the owner checks already-exposed exports. If the only difference is a cosmetic alignment issue, a smaller cohort and scheduled fix may be proportionate. For a blocked but reversible task, offer a tested alternate route and a named update time.
A rollout response rule should combine consequence, exposure, and reversibility: pause or roll back a variation when it can corrupt or hide customer work; narrow exposure while investigating a reproducible blocked task; continue with an owned fix when the defect is low impact and the fallback itself has cost. Verify that changing the flag actually restores the task before calling it mitigation. The rule is an operating proposal, not a claim that flipping every flag is safe. Some flags govern data writes or migrations that an old interface cannot undo.
After a change, verify the same customer journey using a safe account with the affected targeting context. Check both the value selected and the output. Then review the previously exposed cohort for incomplete work. An error count that falls after disabling a variation is useful, but it cannot establish that every affected export has been repaired.
For the next rollout, give support a short intake card and engineering a matching trace: customer task and consequence; occurrence time; account reference; release; flag key, variation, and reason if observed; correlation identifier; safe workaround; investigation owner. Decide where those fields live before the first customer report. A team that can join one complaint to one actual evaluation can test the right product path without asking the customer to diagnose its rollout.
Sources and further reading
- OpenFeature: Evaluation Context
- OpenFeature: Provider and Resolution Details
- LaunchDarkly: Evaluation Reasons
- LaunchDarkly: Flag Variation Evaluation
- LaunchDarkly: Flag Evaluations
- LaunchDarkly: Troubleshoot Evaluation Reasons
- Microsoft Learn: .NET Feature Flag Management
- Mendaro Website Widget, checked September 22, 2026


