19 September 2025 · Guest desk

When stability alerts should wait

Late-night office with computer screens

This note is from a ride-hailing reliability lead who sat Org Watch with us in 2025. Names omitted; the paging rules are theirs, lightly edited.

We used to page on any 0.3-point crash-free dip after 22:00. Night shift learned to snooze. The one evening a payment library actually wedged the driver app, the on-call assumed it was another festival mix shift and went back to sleep. That is alert fatigue with a body count measured in stranded drivers, not in Jira tickets.

Wait if all three are true

  • The stack is unsymbolicated or the mapping file is younger than the binary.
  • Affected users are under a threshold you set per journey — ours is 80 concurrent drivers in a city, not a global percentage.
  • The dip tracks a known traffic event (payday, concert let-out, heavy rain) that historically moves the chart without a new cluster.

Waiting is not ignoring. The night board still shows the cluster. Chat still gets a quiet note. Nobody’s phone rings.

Do not wait if

A named journey is dead with a readable stack, even if the global crash-free line looks polite. Checkout, login, and “start trip” are journeys. Settings fonts are not. If you cannot name the journey, you are not ready to page or to wait; you are ready to open the export.

Devqueuegrid did not invent these thresholds. They only made us write them down so product could see why we sometimes refuse to escalate a pretty chart. If your company still equates care with volume of pages, no syllabus will help until that politics moves.

See also crash-free rate is not a health check and the staged-rollout case study.