19 September 2025 · Guest desk
When stability alerts should wait
This note is from a ride-hailing reliability lead who sat Org Watch with us in 2025. Names omitted; the paging rules are theirs, lightly edited.
We used to page on any 0.3-point crash-free dip after 22:00. Night shift learned to snooze. The one evening a payment library actually wedged the driver app, the on-call assumed it was another festival mix shift and went back to sleep. That is alert fatigue with a body count measured in stranded drivers, not in Jira tickets.
Wait if all three are true
- The stack is unsymbolicated or the mapping file is younger than the binary.
- Affected users are under a threshold you set per journey — ours is 80 concurrent drivers in a city, not a global percentage.
- The dip tracks a known traffic event (payday, concert let-out, heavy rain) that historically moves the chart without a new cluster.
Waiting is not ignoring. The night board still shows the cluster. Chat still gets a quiet note. Nobody’s phone rings.
Do not wait if
A named journey is dead with a readable stack, even if the global crash-free line looks polite. Checkout, login, and “start trip” are journeys. Settings fonts are not. If you cannot name the journey, you are not ready to page or to wait; you are ready to open the export.
Devqueuegrid did not invent these thresholds. They only made us write them down so product could see why we sometimes refuse to escalate a pretty chart. If your company still equates care with volume of pages, no syllabus will help until that politics moves.
See also crash-free rate is not a health check and the staged-rollout case study.