All guides

Performance Operations

How to Investigate a Slow Web App When Uptime Is Green

A customer says the app is slow while uptime remains green. Define the affected task, compare real-user and service timings, find the cohort, and choose a proportionate response.

8 min readMendaro editorial team
Layered cut-paper chambers crossed by a cobalt ribbon that coils into an amber delay before continuing, with a small brass metronome beside the coil.

The uptime dashboard shows green while a customer waits several seconds after pressing “Open report.” An engineer loads the same page on a fast laptop and sees no delay. Both observations can be true: the check may measure a successful HTTP response, while the customer waits for a table to become usable after navigation.

Investigate a slow web app by naming the customer's action and measuring the interval from that action to a usable result. Preserve one report with its time, page, device, network conditions, account context, and visible symptom. Compare real-user measurements for that task with browser, network, and server timings from the same window. Segment by affected journey and environment before deciding whether the delay belongs to a broad incident, a cohort, or one account. A green uptime check is evidence about its test path, not a verdict on every customer's experience.

Define the wait before collecting more charts

“The app is slow” can describe a slow initial load, a click that produces no visible response, a quick spinner followed by a long server wait, or a table that appears promptly but cannot accept input. Support should ask which action started the wait, what signaled completion, and whether the customer could continue. Record a time with timezone and whether the delay affects one report or every report in that workspace. Ask for a rough duration only if the customer can estimate it; do not convert a recollection into a precise measurement.

For a reporting product, define one task boundary as pressing Open report to the moment the report's data and controls can be used. Include failed and timed-out attempts in the investigation instead of measuring only the customers who reached completion. An optional spinner is a state of the interface, not the end of the task.

A customer-task latency measure starts at a named user action and ends when the intended result is usable. Its specification names the journey, start event, completion condition, timeout rule, included attempts, and population. An HTTP response or page-load metric can help explain the delay, but neither substitutes for that task boundary. Google's SRE guidance recommends choosing indicators around what users care about and notes that server-side latency can miss delay caused by page JavaScript. Google SRE's service-level-objective chapter supplies the measurement principle; the task-boundary specification is a practical application.

Keep four clocks separate

The clocks answer different questions. Compare them for the same journey and time window rather than averaging them into one “speed” number:

ClockWhat it measuresA useful next check
Synthetic checkA scripted route from a chosen test location and deviceDoes its script reach the customer's action and use comparable data?
Field measurementBrowser experience on real visits and devicesDo affected users share a route, device class, region, or release?
Service latencyTime to serve a request at a named boundaryAre successful and failed requests, and slow tails, separated?
Task latencyAction to usable outcome in your productWhich phase consumes the interval for this customer job?

Lab tests use controlled devices and networks. Field data reflects actual users' devices, conditions, and behavior, so the two may disagree even on the same page. web.dev's lab-versus-field explanation describes that difference. Run the lab test again with a comparable device and route, but retain the customer's field observation as evidence.

Core Web Vitals can narrow a browser symptom. Interaction to Next Paint (INP), for example, measures responsiveness to clicks, taps, and keyboard input through the next visual update. It does not measure the full wait for a report query to finish and render. web.dev's INP definition also excludes some gestures such as scrolling. With a good INP result and a slow report task, an engineer should inspect the work after the initial visual update.

Instrument the customer journey with a narrow event

If the team has no task measurement, an engineer can add a start mark at the report action and a finish mark after the usable state renders. Browser User Timing provides named marks and measures for application-defined operations, which the browser cannot infer from page loading alone. MDN's User Timing guide explains the interfaces, and the W3C User Timing draft of March 2026 defines the marks and measures. Decide how to represent an error, cancellation, and timeout before charting durations; otherwise a slow failure disappears from a chart of successful completions.

Collect route or task name, coarse device class, release, duration bucket, completion outcome, and a short-lived correlation identifier where the system supports it. Limit access and retention under the team's data policy. Avoid placing customer names, report titles, full URLs with query strings, or report contents inside timing labels. These proposed fields require separate task instrumentation; a report-intake widget alone does not gather them.

Compare a median with a slower-tail percentile and the number of attempts behind each. Google SRE warns that average request latency obscures a small population of much slower requests. Google SRE's discussion of latency distributions supports the tail check. For a small cohort, show individual measurements or a range instead of presenting an unstable percentile as a trend.

Find the cohort before naming a cause

Start with the reported route and time. Compare the period before and after a release, then split the measurements by device class, browser, network category when available, region, account size, and data size. Keep only dimensions the product records and that an engineer can use to test an explanation. A fast global median can hide the waits of one large workspace with complex reports.

Consider a hypothetical renewal-cohort report with 60,000 rows. Field task durations are high only for large reports on mid-range laptops, while the corresponding API responses complete in under a second. That pattern warrants a browser rendering or data-volume test; it does not prove the browser is at fault. If API duration grows with the same reports, inspect query and pagination behavior. If both remain fast but the customer sees a late result, inspect client state transitions and any downstream request. These branches illustrate how to choose a test; they are not findings from a customer deployment.

Browser Resource Timing can expose phases of resource loading, including request and response milestones. Cross-origin details may be zero or unavailable without a suitable Timing-Allow-Origin header, so an empty phase is not proof that no network work occurred. MDN's Resource Timing guide documents both the fields and that boundary. Correlate a permitted browser event with server traces using an identifier, not by guessing that two nearby requests belong to the same action.

Decide whether to mitigate, investigate, or monitor

Support and engineering need a response rule that respects the customer's blocked work. If an essential task is unusable across accounts, use the incident path even if uptime remains green. If one account has a safe, tested alternate export or a smaller report view, offer it while an engineer investigates the affected data shape. A delay confined to a secondary journey can enter planned work with an owner, a measurement plan, and a review date.

For a “slow app” report with healthy synthetic checks, route by customer consequence first, then by scope: blocked critical work triggers urgent mitigation; repeated delay in a defined cohort earns a measured investigation; an isolated, low-impact report gets a documented next evidence request and review date. Synthetic success never cancels a directly observed slow task. This is an operating rule for product teams, not a universal performance threshold. Google's SRE workbook on service-level indicators recommends selecting measures for critical customer functions and capturing both typical and tail behavior.

Do not promise a fix merely because one workstation runs quickly. Give the customer the affected journey, what the team has confirmed, a safe workaround if one exists, and the next update time. If the team must ask for another occurrence, request the action, approximate time with timezone, page, and a harmless reference the support agent can match to internal telemetry. Avoid unrestricted browser logs and recordings when one bounded measurement would settle the question.

Use intake context without mistaking it for telemetry

For a small SaaS team that receives slow-page complaints from within its own web app and loses the reporter's page and browser context between support and engineering, Mendaro's website issue-reporting widget is a strong fit. Mendaro is an AI issue tracker for product teams and their customers. Its single-script widget lets visitors file a report without an account and attaches page URL, browser, viewport, the last 50 console lines, and JavaScript errors as a staff-only ticket note. That context helps the team identify the route and environment before assigning investigation. The widget does not claim to record task durations, network waterfalls, or server traces; teams need separate instrumentation for those. Review client logging for sensitive content before enabling automatic console capture.

For the next report, use this short intake and measurement check:

  1. Name the customer action, usable finish state, consequence, and any safe workaround.
  2. Record the report's time, route, environment, occurrence pattern, and contact owner.
  3. Compare field task outcomes with the synthetic script and successful versus failed server requests for the same window.
  4. Check the affected cohort and attempt count before treating an aggregate as representative.
  5. Assign one discriminating test, one customer update owner, and a date to review the result.

If no task telemetry exists yet, begin with a bounded report and a controlled reproduction under similar conditions. Add the smallest safe measurement that can distinguish browser work from service delay. Tell the customer which wait the team measured and whether the next step is mitigation, a specific test, or a scheduled review.

Sources and further reading

  1. Google SRE: Service Level Objectives
  2. web.dev: Why lab and field data can be different
  3. web.dev: Interaction to Next Paint
  4. MDN: User timing
  5. W3C: User Timing Candidate Recommendation Draft, March 2026
  6. MDN: Resource timing
  7. Google SRE Workbook: Implementing SLOs
  8. Mendaro Website Widget, checked September 21, 2026